Three-dimensional scene generation method and electronic equipment

By semantic segmentation and point cloud segmentation of the initial point cloud, each segmentation result attribute characteristics are assigned to generate interactive three-dimensional scenes, which solves the problem that cannot be independently described and interacted in the prior art and improves the generation efficiency.

CN120374868APending Publication Date: 2025-07-25LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510539560.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The three-dimensional scene generated by the prior art cannot independently describe each object in the scene, the user cannot interact with the objects in the three-dimensional scene, and subsequent processing is inefficient.

Method used

By semantic segmentation and point cloud segmentation of the initial point cloud, each segmentation result attribute characteristics are assigned to generate an interactive three-dimensional scene.

Benefits of technology

The independent description and interaction ability of each object in a three-dimensional scene is realized, and the generation efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374868A_ABST
    Figure CN120374868A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional scene generation method and electronic equipment, and the method comprises the steps: carrying out the semantic segmentation of an initial point cloud of a to-be-processed scene, and obtaining the semantic feature of each point in the initial point cloud; based on the semantic features, point cloud segmentation is carried out on the initial point cloud to obtain at least one segmentation result, the scene to be processed comprises at least one object, and each segmentation result corresponds to one object; corresponding attribute features are given to each segmentation result, at least one segmentation result after assignment is obtained, and the attribute features represent features of an object corresponding to each segmentation result; and generating a three-dimensional scene based on the at least one assigned segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of three-dimensional scene generation, and in particular, to a method for generating a three-dimensional scene and an electronic device. Background Art

[0002] When generating a three-dimensional scene, the data of the real scene can be scanned to generate a corresponding virtual three-dimensional scene. In the related art, the generated three-dimensional scene cannot independently describe each object in the scene, and the user cannot interact with the objects in the three-dimensional scene. Summary of the Invention

[0003] The present disclosure provides a method for generating a three-dimensional scene and a display device.

[0004] According to one aspect of the present disclosure, there is provided a method for generating a three-dimensional scene, including: performing semantic segmentation on the initial point cloud of the scene to be processed to obtain the semantic features of each point in the initial point cloud; based on the semantic features, performing point cloud segmentation on the initial point cloud to obtain at least one segmentation result, the scene to be processed includes at least one object, and each segmentation result corresponds to one object; assigning corresponding attribute features to each segmentation result to obtain at least one segmented result after assignment, the attribute features characterize the features of the object corresponding to each segmentation result; generating a three-dimensional scene based on at least one segmented result after assignment.

[0005] According to an embodiment of the present disclosure, the segmentation result is a Mesh model. Based on the semantic features, performing point cloud segmentation on the initial point cloud to obtain at least one segmentation result includes: based on the semantic categories represented by the semantic features, segmenting the initial point cloud into at least one target point cloud block, and each target point cloud block corresponds to one object in at least one object; generating a Mesh model corresponding to each target point cloud block to obtain at least one Mesh model.

[0006] According to an embodiment of the present disclosure, based on the semantic categories represented by the semantic features, segmenting the initial point cloud into at least one target point cloud block includes: based on the semantic features, segmenting the initial point cloud into at least one initial point cloud block, and different initial point cloud blocks correspond to different semantic categories; according to the spatial distance between adjacent points in each initial point cloud block, segmenting each initial point cloud block into at least one target point cloud block.

[0007] According to an embodiment of the present disclosure, generating a Mesh model corresponding to each target point cloud block to obtain at least one Mesh model includes: determining a Mesh model corresponding to each target point cloud block according to the continuous three-dimensional surface generated by each target point cloud block.

[0008] According to an embodiment of the present disclosure, the attribute features are determined through the following steps: determining a plurality of initial points corresponding to each segmentation result, where the plurality of initial points are the points in the initial point cloud that constitute each segmentation result; determining the attribute features corresponding to each segmentation result according to the initial point features of each initial point, and the initial point features can characterize the geometric features, appearance features, and semantic features of each initial point.

[0009] According to an embodiment of the present disclosure, feature extraction is performed on the initial point cloud to obtain the initial point features of each point in the initial point cloud.

[0010] According to an embodiment of the present disclosure, the attribute features can characterize the semantic category of the object corresponding to the segmentation result.

[0011] According to an embodiment of the present disclosure, generating a three-dimensional scene based on at least one segmented result after assignment includes: determining the position of each segmentation result in the virtual three-dimensional coordinate system according to the three-dimensional coordinates of the multiple points in each segmentation result in the initial point cloud; placing each segmentation result at the corresponding position in the virtual three-dimensional coordinate system to obtain a three-dimensional scene.

[0012] According to an embodiment of the present disclosure, acquiring the scan data of the scene to be processed, where the scan data includes laser scan data and / or image data with multiple different shooting angles; determining the initial point cloud according to the scan data.

[0013] Another aspect of the present disclosure provides an electronic device, including: a memory in which the initial point cloud of the scene to be processed is stored; a processor configured to perform semantic segmentation on the initial point cloud to obtain the semantic features of each point in the initial point cloud; based on the semantic features, performing point cloud segmentation on the initial point cloud to obtain at least one segmentation result, where the scene to be processed includes at least one object, and each segmentation result corresponds to an object; assigning corresponding attribute features to each segmentation result to obtain at least one segmented result after assignment, where the attribute features characterize the features of the object corresponding to each segmentation result; generating a three-dimensional scene based on at least one segmented result after assignment.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0016] Figure 1 is a flowchart of a three-dimensional scene generation method according to an embodiment of the present disclosure;

[0017] Figure 2It is a schematic diagram of a scene to be processed according to an embodiment of the present disclosure;

[0018] Figure 3 It is a schematic diagram of an initial point cloud segmentation process according to an embodiment of the present disclosure;

[0019] Figure 4 It is a schematic diagram of an initial point cloud segmentation process according to another embodiment of the present disclosure; and

[0020] Figure 5 It is a schematic block diagram of an exemplary electronic device for implementing the embodiments of the present disclosure. Detailed implementation manners

[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0022] In the technical solution of the present disclosure, the processing of data involved (such as including but not limited to user personal information) in aspects such as collection, storage, use, processing, transmission, provision, disclosure, and application complies with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.

[0023] When generating a three-dimensional scene, a corresponding virtual three-dimensional scene can be generated by scanning the data of the real scene. In application scenarios such as three-dimensional games and high-precision simulators, the generated three-dimensional scene needs to be interactive. For example, a user needs to interact with an object in a three-dimensional game.

[0024] Three-dimensional scanning devices and software in the related art can generate a three-dimensional scene based on the scanned data, but the generated three-dimensional scene is a single model, which cannot independently describe each object in the scene and achieve interaction. Subsequent fine post-processing processes such as manual annotation and target segmentation are required to meet the interaction requirements, and the efficiency is also relatively low.

[0025] Figure 1 It is a flowchart of a three-dimensional scene generation method according to an embodiment of the present disclosure.

[0026] As Figure 1 shown, the three-dimensional scene generation method of this embodiment includes operations S110 - S140.

[0027] In operation S110, semantic segmentation is performed on the initial point cloud of the scene to be processed to obtain the semantic features of each point in the initial point cloud.

[0028] In an embodiment of the present disclosure, the scene to be processed is the real scene corresponding to the three-dimensional scene. For example, the scene to be processed is an office scene, which includes desks and chairs. The three-dimensional scene generated based on the scene to be processed is the virtual three-dimensional scene of the office scene, and the virtual three-dimensional scene includes virtual desks and virtual chairs.

[0029] The initial point cloud is a set of discrete three-dimensional space data directly obtained through sensors or three-dimensional reconstruction techniques without any manual or algorithmic processing. The initial point cloud records the geometric coordinates of the surfaces of objects in the scene to be processed in the form of a point set and can include additional attributes (such as color, reflection intensity, timestamp, etc.). The initial point cloud is the most primitive sampling representation of the three-dimensional environment in the digital space.

[0030] In an embodiment of the present disclosure, semantic segmentation is performed on the initial point cloud of the scene to be processed to obtain the semantic features of each point in the initial point cloud. After performing semantic segmentation on the initial point cloud of the scene to be processed to obtain the semantic features of each point, semantic class labels (such as desks, chairs, computers, etc.) can be assigned to each point in the initial point cloud based on the semantic features, thereby converting the original geometric data into structured information with semantic meaning.

[0031] In an embodiment of the present disclosure, the initial point cloud of the scene to be processed can be semantically segmented by deep learning methods or traditional methods. For example, by using deep learning methods to perform semantic segmentation on the initial point cloud, the initial point cloud can be input into a pre-trained three-dimensional semantic segmentation algorithm, and the algorithm will classify each point to identify its belonging to different semantic classes. The three-dimensional semantic segmentation algorithm can be the PointNet++ algorithm or the OmniSeg3D algorithm, etc. For example, by using traditional methods to perform semantic segmentation on the initial point cloud, methods such as graph cut-based segmentation methods or random forest classification can be used to perform semantic segmentation on the initial point cloud of the scene to be processed.

[0032] In an embodiment of the present disclosure, the semantic features can represent the semantic class of each point. After obtaining the semantic features of each point in the initial point cloud, the semantic class of each point can be assigned to each point in the form of a label. For example, after assigning the corresponding labels to each point in the initial point cloud, a semantic point cloud can be obtained. Each point in the semantic point cloud has a one-to-one mapping relationship with the corresponding point in the initial point cloud, and the points in the semantic point cloud have additional semantic label information compared to the points in the initial point cloud.

[0033] In operation S120, based on the semantic features, point cloud segmentation is performed on the initial point cloud to obtain at least one segmentation result. The scene to be processed includes at least one object, and each segmentation result corresponds to one object.

[0034] In an embodiment of the present disclosure, the scene to be processed includes at least one object. Here, an object refers to an object in the scene to be processed. For example, an object can be a chair, a table, an air conditioner, etc.

[0035] In an embodiment of the present disclosure, based on semantic features, point cloud segmentation is performed on the initial point cloud to obtain at least one segmentation result, and each segmentation result corresponds to an object. The segmentation result is composed of multiple points in the initial point cloud, and the semantic categories represented by the semantic features corresponding to the multiple points in each segmentation result are the same.

[0036] For example, the scene to be processed includes a chair and a table. Point cloud segmentation is performed on the initial point cloud to obtain a first segmentation result and a second segmentation result. The semantic categories corresponding to the multiple points in the first segmentation result are all chairs, and the first segmentation result corresponds to the chair in the scene to be processed. The semantic categories corresponding to the multiple points in the second segmentation result are all tables, and the second segmentation result corresponds to the table in the scene to be processed.

[0037] In operation S130, corresponding attribute features are assigned to each segmentation result to obtain at least one segmented result after assignment, and the attribute features characterize the features of the object corresponding to each segmentation result.

[0038] In an embodiment of the present disclosure, the attribute features characterize the features of the object corresponding to each segmentation result. For example, the attribute features may include features such as the semantic category, texture map, material, physical properties, usability, etc. of the object corresponding to each segmentation result. The attribute features can be obtained by feature extraction from the initial point cloud. For example, the initial point cloud can be input into a pre-trained feature extraction algorithm to obtain the attribute features corresponding to each segmentation result.

[0039] For example, the first segmentation result corresponds to the chair in the scene to be processed. The attribute features corresponding to the first segmentation result may include that the semantic category of the first segmentation result is a chair, the material of the chair is wood, the texture of the chair is wood grain, etc. Assigning the attribute features to the first segmentation result to obtain the first segmentation result after assignment.

[0040] In operation S140, a three-dimensional scene is generated based on at least one segmented result after assignment.

[0041] In an embodiment of the present disclosure, a three-dimensional scene can be generated based on the positions of the objects corresponding to the segmented results after assignment in the scene to be processed. For example, the scene to be processed is an office scene, and the objects corresponding to at least one segmented result after assignment include a chair, a table, and a computer. According to the positions of a chair, a table, and a computer in the scene to be processed, at least one segmented result after assignment can be placed at the corresponding positions in the virtual three-dimensional scene, thereby generating a three-dimensional scene.

[0042] Through the embodiments of the present disclosure, point cloud segmentation is performed on the initial point cloud based on semantic features, so that each obtained segmentation result corresponds to an object. In the three-dimensional scene generated based on the segmented results after assignment, the user can interact with each segmented result in the three-dimensional scene. At the same time, each segmented result is assigned the features of the corresponding object, so that each segmented result after assignment can reflect the true attributes of the corresponding object in the scene to be processed, improving the authenticity of the three-dimensional scene.

[0043] Figure 2 It is a schematic diagram of the scene to be processed according to the embodiments of the present disclosure.

[0044] As Figure 2 shown, in the scene 200 to be processed, the cube 201 is the first object, the cube 202 is the second object, the cube 203 is the third object, and the cube 204 is the fourth object. For example, the cube 201 can be the first wooden chair, the cube 202 can be the second wooden chair, the cube 203 can be the first wooden table, and the cube 204 can be the second wooden table.

[0045] In the embodiments of the present disclosure, the segmentation result is a Mesh model. Based on semantic features, point cloud segmentation is performed on the initial point cloud to obtain at least one segmentation result, including: dividing the initial point cloud into at least one target point cloud block based on the semantic category represented by the semantic features, and each target point cloud block corresponds to one object in at least one object. Generate a Mesh model corresponding to each target point cloud block to obtain at least one Mesh model.

[0046] Figure 3 It is a schematic diagram of the initial point cloud segmentation process according to the embodiments of the present disclosure.

[0047] As Figure 3 shown, the initial point cloud 300 includes a first target point cloud block 301, a second target point cloud block 302, a third target point cloud block 303, and a fourth target point cloud block 304.

[0048] In the embodiments of the present disclosure, the first target point cloud block 301 may correspond to the cube 201, that is, the first target point cloud block corresponds to the first wooden chair in the scene to be processed, and the multiple points in the first target point cloud block are the multiple points in the initial point cloud corresponding to the first wooden chair. The second target point cloud block 302 may correspond to the cube 202, that is, the second target point cloud block corresponds to the second wooden chair in the scene to be processed, and the multiple points in the second target point cloud block are the multiple points in the initial point cloud corresponding to the second wooden chair. The third target point cloud block 303 may correspond to the cube 203, that is, the third target point cloud block corresponds to the first wooden table in the scene to be processed, and the multiple points in the third target point cloud block are the multiple points in the initial point cloud corresponding to the first wooden table. The fourth target point cloud block 304 may correspond to the cube 204, that is, the fourth target point cloud block corresponds to the second wooden table in the scene to be processed, and the multiple points in the fourth target point cloud block are the multiple points in the initial point cloud corresponding to the second wooden table.

[0049] In the embodiments of the present disclosure, a Mesh model is a data structure for representing the surface shape of a three-dimensional object. The Mesh model is composed of vertices, edges, and faces. Converting the target point cloud block into the corresponding Mesh model can convert the discrete points in the target point cloud block into a continuous surface, providing structured and editable set data for subsequent steps.

[0050] Figure 4 It is a schematic diagram of the initial point cloud segmentation process according to another embodiment of the present disclosure.

[0051] In the embodiments of the present disclosure, based on the semantic categories represented by the semantic features, the initial point cloud is segmented into at least one target point cloud block, including: based on the semantic features, the initial point cloud is segmented into at least one initial point cloud block, and different initial point cloud blocks correspond to different semantic categories. According to the spatial distance between adjacent points in each initial point cloud block, each initial point cloud block is segmented into at least one target point cloud block.

[0052] As Figure 4 shown, in the embodiments of the present disclosure, based on the semantic features, the initial point cloud 400 is segmented into a first initial point cloud block 410 and a second initial point cloud block 420, and different initial point cloud blocks correspond to different semantic categories. For example, the first initial point cloud block corresponds to a wooden chair, and the second initial point cloud block corresponds to a wooden table.

[0053] In the embodiments of the present disclosure, based on the semantic categories represented by the semantic features, the points with the same semantic category are divided into the same initial point cloud block. For example, all the points whose semantic categories represented by the semantic features are wooden chairs are divided into the first initial point cloud block 410, and all the points whose semantic categories represented by the semantic features are wooden tables are divided into the second initial point cloud block 420.

[0054] In the embodiments of the present disclosure, a region growing algorithm can be adopted to segment the initial point cloud blocks according to the spatial distances between adjacent points in each initial point cloud block. After determining the initial point cloud blocks, a random point and its neighbor points are selected from the initial point cloud blocks. If the spatial distance between the two points is less than the target division threshold corresponding to the semantic category of the initial point cloud block, the group of adjacent points is merged into the same target set, and it is considered that the group of adjacent points belongs to the same object. By iteratively processing all the points in the initial point cloud block in this way, each initial point cloud block is segmented into at least one target point cloud block, so as to separate multiple objects of the same category (such as adjacent chairs) in the scene. It should be noted that the target division threshold is a preset threshold and is related to the semantic category corresponding to the initial point cloud block.

[0055] For example, the object corresponding to the first initial point cloud block is a wooden chair. According to the spatial distances between adjacent points in the first initial point cloud block, the first initial point cloud block can be divided into a first target point cloud block and a second target point cloud block. The first target point cloud block corresponds to the first wooden chair, and the second target point cloud block corresponds to the second wooden chair. Similarly, the second initial point cloud block is segmented to obtain a third target point cloud block and a fourth target point cloud block. The third target point cloud block corresponds to the first wooden table, and the fourth target point cloud block corresponds to the second wooden table.

[0056] Through the embodiments of the present disclosure, the accuracy of the obtained target point cloud blocks can be improved.

[0057] In some embodiments of the present disclosure, generating a Mesh model corresponding to each target point cloud block to obtain at least one Mesh model includes: determining the Mesh model corresponding to each target point cloud block according to the continuous three-dimensional surface generated from each target point cloud block.

[0058] Since the point cloud only contains discrete coordinates and cannot directly express the continuity of the object surface (such as walls, desktops, etc.), the object corresponding to each target point cloud block can be characterized by generating the Mesh model corresponding to each target point cloud block.

[0059] In the embodiments of the present disclosure, the Mesh model can be generated by performing steps such as point cloud preprocessing, surface reconstruction, and post-processing optimization on the target point cloud block. Point cloud preprocessing is used to remove the noise in the point cloud and optimize the point cloud. Poisson reconstruction, Delaunay triangulation, and marching cubes algorithm can be used for surface reconstruction. Finally, steps such as hole filling and smoothing processing can be used for post-processing optimization.

[0060] For example, for each target point cloud block, the Poisson reconstruction algorithm can be used. Based on the distribution of the point cloud in the target point cloud block, a continuous Mesh surface is constructed by solving the Poisson equation, so as to obtain the Mesh models of each object such as tables, chairs, and computers in the scene.

[0061] Through the embodiments of the present disclosure, the accuracy of the determined power supply Mesh model can be improved.

[0062] In the embodiments of the present disclosure, the attribute features can characterize the semantic categories of the objects corresponding to the segmentation results.

[0063] For example, when the object corresponding to the segmentation result is a chair, the attribute features assigned to the segmentation result are the semantic categories of the object corresponding to the segmentation result, that is, the semantic category of the chair is assigned to the segmentation result.

[0064] In the embodiments of the present disclosure, in addition to semantic features, the attribute features can also include features such as texture maps, materials, physical properties, usability, etc. of the objects corresponding to the segmentation results. The attribute features can be stored through attribute point clouds. For example, each point in the attribute point cloud has a one-to-one mapping relationship with the corresponding point in the initial point cloud, and the points in the attribute point cloud have additional information about the attribute features compared to the points in the initial point cloud.

[0065] In the embodiments of the present disclosure, it further includes: extracting features from the initial point cloud to obtain the initial point features of each point in the initial point cloud.

[0066] In the embodiments of the present disclosure, the initial point features can include features such as normal vector features, curvature features, 3DGS spherical harmonics, three-dimensional usability, etc. For different features, different algorithms can be used to extract features from the initial point cloud.

[0067] For example, the normal vector features of each point in the initial point cloud can be determined through principal component analysis or deep learning-based methods. The curvature features of each point in the initial point cloud can be determined through principal curvature estimation or deep learning prediction methods. The 3DGS spherical harmonics of each point in the initial point cloud can be determined through 3D Gaussian parameterization or differentiable rendering optimization methods.

[0068] In the embodiments of the present disclosure, the attribute features are determined through the following steps: determining a plurality of initial points corresponding to each segmentation result, where the plurality of initial points are the points in the initial point cloud that make up each segmentation result. According to the initial point features of each initial point, the attribute features corresponding to each segmentation result are determined, and the initial point features can characterize the geometric features, appearance features, and semantic features of each initial point.

[0069] In the embodiments of the present disclosure, after determining the initial point features of the segmentation results, the attribute features of the segmentation results can be determined based on the initial point features. The initial point features can include features such as normal vector features, curvature features, 3DGS spherical harmonics, three-dimensional usability, etc., and the attribute features can include features such as semantic features, texture maps, materials, physical properties, usability, etc.

[0070] For example, for the texture mapping feature, projection texture coordinates can be generated based on the normal vector direction characterized by the normal vector feature. Also, a curvature map characterized by the curvature feature can be used as a mask to overlay detailed textures (scratches, wear) in high-curvature regions, thereby determining the texture mapping feature of the segmentation result. For the physical property feature, the friction coefficient of different regions can be determined based on the curvature feature, and the simplified collision body type can be selected according to the evaluation result of three-dimensional usability, thereby determining the physical property of the segmentation result.

[0071] Through the embodiments of the present disclosure, the geometric and semantic features of the point cloud can be efficiently mapped to the multi-dimensional attributes of the segmentation result, realizing the transformation from the initial point features of the point cloud to the attribute features of the segmentation result.

[0072] In the embodiments of the present disclosure, generating a three-dimensional scene based on at least one assigned segmentation result includes: determining the position of each segmentation result in the virtual three-dimensional coordinate system according to the three-dimensional coordinates of multiple points in the initial point cloud in each segmentation result. Placing each segmentation result at the corresponding position in the virtual three-dimensional coordinate system to obtain a three-dimensional scene.

[0073] In the embodiments of the present disclosure, each segmentation result is composed of multiple points. According to the three-dimensional coordinates of the multiple points constituting the segmentation result in the initial point cloud, the position of each segmentation result in the virtual three-dimensional coordinate system can be determined. The three-dimensional scene is generated in the virtual three-dimensional coordinate system. Placing each segmentation result at the corresponding pose in the virtual three-dimensional coordinate system can obtain a three-dimensional scene. Since each segmentation result corresponds to an object, the obtained three-dimensional scene is interactive.

[0074] For example, at least one assigned segmentation result includes a table, a chair, a computer, etc. The user can manipulate the virtual table, virtual chair, virtual computer, etc. in the generated three-dimensional scene, and thus interaction can be realized.

[0075] In the embodiments of the present disclosure, it further includes: obtaining the scan data of the scene to be processed, where the scan data includes laser scan data and / or image data with multiple different shooting angles. Determining the initial point cloud according to the scan data.

[0076] In the embodiments of the present disclosure, the scan data can be laser scan data obtained by a laser scanner, or image data with multiple different shooting angles collected by a camera. The scan data can be imported into three-dimensional scene reconstruction processing software, and the initial point cloud can be determined using an algorithm.

[0077] For example, for the laser scan data, the point cloud data in the laser scan data can be converted and added to a unified three-dimensional space coordinate system to obtain the initial point cloud of the scene to be processed. For the image data, the image data can be imported into multi-view Figure 3The Structure-from-Motion and Multi-View Stereo (COLMAP) reconstruction tool uses the Structure-from-Motion Algorithm (SfM) to calculate the camera's motion trajectory and the 3D point cloud information of the scene by analyzing the feature point matching relationships from different perspectives in the image data, and then obtains the initial point cloud.

[0078] Through the embodiments of the present disclosure, the initial point cloud of the scene to be processed can be accurately determined.

[0079] In the embodiments of the present disclosure, the electronic device includes: a memory that stores the initial point cloud of the scene to be processed; a processor that is configured to perform semantic segmentation on the initial point cloud to obtain the semantic features of each point in the initial point cloud; based on the semantic features, perform point cloud segmentation on the initial point cloud to obtain at least one segmentation result, where the scene to be processed includes at least one object, and each segmentation result corresponds to an object; assign corresponding attribute features to each segmentation result to obtain at least one assigned segmentation result, where the attribute features characterize the features of the object corresponding to each segmentation result; and generate a 3D scene based on at least one assigned segmentation result.

[0080] In the embodiments of the present disclosure, the processor executes the 3D scene generation method described above based on the initial point cloud of the scene to be processed stored in the memory.

[0081] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0082] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement the method of the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0083] As Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0084] Multiple components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0085] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the three-dimensional scene generation method. For example, in some embodiments, the three-dimensional scene generation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the three-dimensional scene generation method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the three-dimensional scene generation method by any other appropriate means (e.g., by means of firmware).

[0086] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0087] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0088] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0089] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: an electronic device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0090] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser, through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0091] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. Among them, the server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0092] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0093] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for generating a three-dimensional scene, comprising: Performing semantic segmentation on the initial point cloud of the scene to be processed to obtain the semantic features of each point in the initial point cloud; Based on the semantic features, performing point cloud segmentation on the initial point cloud to obtain at least one segmentation result, where the scene to be processed includes at least one object, and each segmentation result corresponds to one object; Assigning corresponding attribute features to each of the segmentation results to obtain at least one segmented result after assignment, where the attribute features characterize the features of the object corresponding to each segmentation result; Generating a three-dimensional scene based on the at least one segmented result after assignment.

2. The method according to claim 1, where the segmentation result is a Mesh model, and the performing point cloud segmentation on the initial point cloud based on the semantic features to obtain at least one segmentation result includes: Based on the semantic categories represented by the semantic features, segmenting the initial point cloud into at least one target point cloud block, and each target point cloud block corresponds to one of the at least one object; Generating a Mesh model corresponding to each of the target point cloud blocks to obtain at least one Mesh model.

3. The method according to claim 2, where the segmenting the initial point cloud into at least one target point cloud block based on the semantic categories represented by the semantic features includes: Based on the semantic features, segmenting the initial point cloud into at least one initial point cloud block, and different initial point cloud blocks correspond to different semantic categories; According to the spatial distance between adjacent points in each initial point cloud block, segmenting each initial point cloud block into at least one target point cloud block.

4. The method according to claim 2, where the generating a Mesh model corresponding to each of the target point cloud blocks to obtain at least one Mesh model includes: Determining the Mesh model corresponding to each of the target point cloud blocks according to the continuous three-dimensional surface generated by each of the target point cloud blocks.

5. The method according to claim 1, where the attribute features are determined through the following steps: Determining a plurality of initial points corresponding to each of the segmentation results, and the plurality of initial points are the points in the initial point cloud that make up each segmentation result; According to the initial point features of each initial point, determining the attribute features corresponding to each segmentation result, and the initial point features can characterize the geometric features, appearance features, and semantic features of each initial point.

6. The method according to claim 5 further includes: Performing feature extraction on the initial point cloud to obtain the initial point features of each point in the initial point cloud.

7. The method according to claim 1, where the attribute features can characterize the semantic category of the object corresponding to the segmentation result.

8. The method according to claim 1, where the generating a three-dimensional scene based on the at least one segmented result after assignment includes: According to the three-dimensional coordinates of multiple points in each segmentation result in the initial point cloud, determining the position of each segmentation result in the virtual three-dimensional coordinate system; Placing each segmentation result at the corresponding position in the virtual three-dimensional coordinate system to obtain the three-dimensional scene.

9. The method according to claim 1 further includes: Obtain the scan data of the to-be-processed scene, where the scan data includes laser scan data and / or image data with multiple different shooting angles; Determine the initial point cloud according to the scan data.

10. An electronic device, comprising: A memory in which the initial point cloud of the to-be-processed scene is stored; A processor configured to perform semantic segmentation on the initial point cloud to obtain the semantic features of each point in the initial point cloud; based on the semantic features, perform point cloud segmentation on the initial point cloud to obtain at least one segmentation result, where the to-be-processed scene includes at least one object, and each segmentation result corresponds to an object; assign corresponding attribute features to each segmentation result to obtain at least one segmented result after assignment, where the attribute features characterize the features of the object corresponding to each segmentation result; generate a three-dimensional scene based on the at least one segmented result after assignment.