A complex scenario modeling method, device, server, and readable storage medium
By acquiring depth images and camera poses, combining scene recognition models and dense point cloud reconstruction algorithms, complex scene models are generated, and the problem of unreasonable connection between simple scene models is solved, achieving efficient three-dimensional reconstruction rendering effect and good roaming experience.
Patent Information
- Application Number
- CN202110617217.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-06-03
AI Technical Summary
The rendering effect of the three-dimensional model of complex scenes in the prior art is poor, and the connection between simple scene models is unreasonable, resulting in poor user roaming in the three-dimensional model.
By acquiring depth images and camera poses, combining pre-trained scene recognition models and dense point cloud reconstruction algorithms, the target scene model is generated, including point cloud fusion for simple and special scenes, and the point cloud is optimized using TSDF and MVS algorithms to achieve complete connection between scenes.
Real-time automatic modeling is realized in complex scenarios, and the generated target scene model does not require a large amount of computing resources for rendering, improving the three-dimensional reconstruction rendering effect and roaming experience.
Smart Images

Figure CN113327319B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of 3D reconstruction, and particularly relates to a complex scene modeling method, device, server, and readable storage medium. Background Art
[0002] The principle of 3D reconstruction is to reconstruct a 3D model based on images taken by a camera at different shooting positions in a scene. In practical applications, the scene is generally complex, including simple scenes and other scenes. In the prior art, the 3D reconstruction of complex scenes is to obtain simple scene models by separately shooting multiple simple scenes in the complex scene with a camera, and then render multiple simple scene models in the later stage to obtain a complex scene model. However, due to the lack of reasonable transitions at the joints between the modules corresponding to multiple simple scenes in the 3D model of the complex scene, the rendering effect of the 3D model of the complex scene is poor, resulting in a poor viewing effect for users during the roaming process in the 3D model of the complex scene from the first perspective. Summary of the Invention
[0003] Embodiments of this application provide a complex scene modeling method, device, server, and readable storage medium, which can solve the problem of poor model rendering effect during the 3D reconstruction of complex scenes.
[0004] In a first aspect, embodiments of this application provide a complex scene modeling method, including:
[0005] Obtain a first image to be processed and a first camera pose corresponding to the first image to be processed, where the image to be processed is a depth image taken by a camera at different shooting positions in a scene to be recognized;
[0006] Generate a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized, where the scene type of the scene to be recognized includes simple scenes and special scenes.
[0007] In a possible implementation manner of the first aspect, before generating a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized, it further includes:
[0008] Determine the scene type of the scene to be recognized according to a pre-trained scene recognition model.
[0009] In a possible implementation manner of the first aspect, the pre-trained scene recognition model includes a feature extraction layer, a feature selection layer, and a classification layer;
[0010] Determining the scene type of the scene to be recognized according to a pre-trained scene recognition model includes:
[0011] Import the image to be processed into the feature extraction layer, and output significant features and supplementary features;
[0012] Import the significant features into a feature selector to output target significant features;
[0013] Extract the local representation information of the target significant features and the global representation information in the supplementary features respectively;
[0014] Concatenate the local representation information and the global representation information and import them into the classification layer to output the scene type of the scene to be recognized.
[0015] In a possible implementation manner of the first aspect, generating a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized includes:
[0016] Construct a first dense point cloud according to the first image to be processed, the second camera pose corresponding to the first image to be processed, and a preset first scene reconstruction algorithm;
[0017] If the scene type of the scene to be recognized is the first scene, form a target scene model according to the first dense point cloud;
[0018] If the scene type of the scene to be recognized is the second scene, construct a second dense point cloud, and generate a target scene model according to the first dense point cloud and the second dense point cloud.
[0019] In a possible implementation manner of the first aspect, if the scene type of the scene to be recognized is the second scene, constructing a second dense point cloud and forming a target scene model according to the first dense point cloud and the second dense point cloud includes:
[0020] If the scene type of the scene to be recognized is the second scene, generate a shooting reminder instruction, send the shooting reminder instruction to the camera to instruct the camera to display a predicted shooting position to the user, and the predicted shooting position is the position indicating the user to shoot in the second scene;
[0021] Obtain a second image to be processed and the second camera pose corresponding to the second image to be processed;
[0022] Generate a second dense point cloud according to the second image to be processed, the second camera pose corresponding to the second image to be processed, and a preset second scene reconstruction algorithm;
[0023] Register the first dense point cloud and the second dense point cloud;
[0024] Form a target scene model based on the registered first dense point cloud and the second dense point cloud.
[0025] In a possible implementation of the first aspect, generating the second dense point cloud according to the second image to be processed, the second camera pose corresponding to the second image to be processed, and a preset second scene reconstruction algorithm includes:
[0026] Generate a point cloud to be processed according to the image to be processed and the camera pose corresponding to the image to be processed;
[0027] Fuse the point cloud to be processed based on the TSDF algorithm to obtain a fused point cloud;
[0028] Perform statistical filtering on the fused point cloud to obtain an optimized fused point cloud;
[0029] Densify the optimized fused point cloud based on the MVS algorithm to obtain the second dense point cloud.
[0030] In a possible implementation of the first aspect, if the scene type of the scene to be recognized is the second scene, generate a shooting reminder instruction, and send the shooting reminder instruction to the camera to indicate the camera to display the predicted shooting position to the user, where the predicted shooting position is the position that prompts the user to shoot in the second scene, including:
[0031] Obtain the predicted shooting position according to a preset position prediction algorithm and the first camera pose;
[0032] Generate a shooting reminder instruction according to the predicted shooting position.
[0033] In a second aspect, an embodiment of the present application provides a device, including:
[0034] An acquisition module, configured to acquire a first image to be processed and a first camera pose corresponding to the first image to be processed, where the image to be processed is an image captured by the camera in response to a shooting reminder instruction of the user in the scene to be recognized;
[0035] A generation module, configured to generate a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized.
[0036] In a possible implementation manner, the device further includes:
[0037] An identification module, configured to determine the scene type of the scene to be recognized according to a pre-trained scene recognition model.
[0038] In a possible implementation manner, the pre-trained scene recognition model includes a feature extraction layer, a feature selection layer, and a classification layer;
[0039] The recognition module includes:
[0040] A first processing sub-module, configured to import the image to be processed into the feature extraction layer and output significant features and supplementary features;
[0041] A second processing sub-module, configured to import the significant features into a feature selector and output target significant features;
[0042] An extraction sub-module, configured to extract local representation information of the target significant features and global representation information in the supplementary features respectively;
[0043] A classification sub-module, configured to splice the local representation information and the global representation information and import the result into the classification layer, and output the scene type of the scene to be recognized.
[0044] In a possible implementation manner, the generation module includes:
[0045] A construction sub-module, configured to construct a first dense point cloud according to the first image to be processed, the second camera pose corresponding to the first image to be processed, and a preset first scene reconstruction algorithm;
[0046] A first generation sub-module, configured to, if the scene type of the scene to be recognized is a first scene, form a target scene model according to the first dense point cloud;
[0047] A second generation sub-module, configured to, if the scene type of the scene to be recognized is a second scene, construct a second dense point cloud and generate a target scene model according to the first dense point cloud and the second dense point cloud.
[0048] In a possible implementation manner, the first generation sub-module includes:
[0049] A generation unit, configured to, if the scene type of the scene to be recognized is a second scene, generate a shooting reminder instruction, send the shooting reminder instruction to the camera to instruct the camera to display a predicted shooting position to the user, and the predicted shooting position is a position for instructing the user to shoot in the second scene;
[0050] An acquisition unit, configured to acquire a second image to be processed and the second camera pose corresponding to the second image to be processed;
[0051] A second generation unit, configured to generate a second dense point cloud according to the second image to be processed, the second camera pose corresponding to the second image to be processed, and a preset second scene reconstruction algorithm;
[0052] A registration unit for registering the first dense point cloud and the second dense point cloud;
[0053] A third generation unit for forming a target scene model based on the registered first dense point cloud and the second dense point cloud.
[0054] In a possible implementation manner, the second generation sub-module includes:
[0055] A fourth generation unit for generating a point cloud to be processed according to the image to be processed and the camera pose corresponding to the image to be processed;
[0056] A fusion unit for fusing the point cloud to be processed based on the TSDF algorithm to obtain a fused point cloud;
[0057] An optimization unit for performing statistical filtering on the fused point cloud to obtain an optimized fused point cloud;
[0058] A dense processing unit for performing dense processing on the optimized fused point cloud based on the MVS algorithm to obtain a second dense point cloud.
[0059] In a possible implementation manner, the generation unit includes:
[0060] A prediction sub-unit for obtaining a predicted shooting point according to a preset point prediction algorithm and the first camera pose;
[0061] A generation sub-unit for generating a shooting reminder instruction according to the predicted shooting point.
[0062] In a third aspect, an embodiment of the present application provides a server, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in the first aspect above is implemented.
[0063] In a fourth aspect, an embodiment of the present application provides a readable storage medium. When the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0064] The beneficial effects of the embodiments of the present application compared with the prior art are:
[0065] In an embodiment of the present application, a first image to be processed and a first camera pose corresponding to the first image to be processed are obtained. The image to be processed is a depth image obtained by the camera at different shooting positions in the scene to be recognized. Based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized, a target scene model is generated. It can be seen that the present application can perform real-time automatic modeling in a complex scene. Since the connection between scenes in the directly generated target scene model is complete, relatively little computing resources are required for rendering in the later stage, and the rendering effect of 3D reconstruction is also improved, resulting in a better viewing effect when the user roams in the 3D model of the complex scene from the first perspective. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0067] Figure 1 is a schematic flowchart of a complex scene modeling method provided by an embodiment of the present application;
[0068] Figure 2 is a complex scene modeling method provided by an embodiment of the present application Figure 1 specific implementation flowchart of step S104 therein;
[0069] Figure 3 is a complex scene modeling method provided by an embodiment of the present application Figure 2 specific flowchart of step S206 therein;
[0070] Figure 4 is a complex scene modeling method provided by an embodiment of the present application Figure 3 specific implementation flowchart of step S302 therein;
[0071] Figure 5 is a complex scene modeling method provided by an embodiment of the present application Figure 3 specific implementation flowchart of step S306 therein;
[0072] Figure 6 is a schematic structural diagram of a complex scene modeling device provided by an embodiment of the present application;
[0073] Figure 7 is a schematic structural diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] In the following description, specific details such as specific system architectures and technologies are presented for purposes of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.
[0075] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0076] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0077] As used in the specification of the present application and the appended claims, the term "if" can be interpreted, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0078] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0079] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0080] The technical solutions provided by the embodiments of the present application will be introduced below through specific examples.
[0081] SeeFigure 1 , which is a schematic flowchart of the complex scene modeling method provided by the embodiments of the present application. As an example rather than a limitation, this method can be applied to a server that is connected to a camera. The server can be a computing device such as a cloud server. This method can include the following steps:
[0082] Step S102: Obtain a first image to be processed and a first camera pose corresponding to the first image to be processed.
[0083] Among them, the image to be processed is a depth image obtained by the camera at different shooting positions in the scene to be recognized. The first camera pose refers to the IMU data collected by the IMU control unit of the camera when the camera captures the image to be processed. Preferably, the camera in the embodiments of the present application can refer to an eight-eye camera, that is, the eight-eye camera consists of two groups, each group having four fisheye lenses. The four lenses respectively collect four groups of lens images and stitch them into a 360° panoramic image.
[0084] Step S104: Generate a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized.
[0085] Among them, the scene types of the scene to be recognized include simple scenes and special scenes. It should be noted that the complex scene in the embodiments of the present application refers to a scene with inconsistent semantic information. The complex scene can include simple scenes and special scenes. Exemplarily, the complex scene can be a multi-floor indoor scene, where the simple scenes are each floor and the special scenes are the stairs between each floor. Of course, the embodiments of the present application do not limit the specific types of complex scenes.
[0086] Preferably, before generating a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized, it further includes:
[0087] Determine the scene type of the scene to be recognized according to a pre-trained scene recognition model.
[0088] Among them, the pre-trained scene recognition model can be pre-trained using an open-source dataset as a training set. The pre-trained scene recognition model includes a feature extraction layer, a feature selection layer, and a classification layer.
[0089] Specifically, determining the scene type of the scene to be recognized according to the pre-trained scene recognition model includes the following four steps:
[0090] The first step: Import the image to be processed into the feature extraction layer and output significant features and supplementary features.
[0091] Among them, the feature extraction layer is mainly composed of a CNN convolutional neural network. The feature extraction layer includes a significant feature extraction sub-layer and a supplementary feature extraction sub-layer. The significant feature refers to the significant target object in the scene to be recognized, and the supplementary feature refers to some line contour features in the scene to be recognized.
[0092] In a specific application, in the significant feature extraction sub-layer, the selective search algorithm is used to determine the significant features and supplementary features for the image to be processed.
[0093] According to the support vector machine, calculate the counting ratio of the correlation strength between the candidate significant features and the scene category. Determine the candidate significant features with a counting ratio greater than the ratio threshold as the significant features; in the supplementary feature sub-layer, perform local separation on the image to be processed based on the contour saliency measurement of the central axis to obtain the supplementary features.
[0094] Second step: Import the significant features into the feature selector to output the target significant features.
[0095] In a specific application, use the selective search algorithm to determine the candidate significant features for the significant features, calculate the counting ratio of the correlation strength between the candidate significant features and the scene category according to the support vector machine, and determine the candidate significant features with a counting ratio greater than the ratio threshold as the significant features.
[0096] Third step: Extract the local representation information of the target significant features and the global representation information in the supplementary features respectively.
[0097] In a specific application, use the multi-resolution CNN convolutional neural network framework to extract the local representation information of the target significant features and the global representation information in the supplementary features respectively.
[0098] Fourth step: After splicing the local representation information and the global representation information, import them into the classification layer to output the scene type of the scene to be recognized.
[0099] In a specific application, use a fully connected layer to splice and classify the local representation information and the global representation information, and output the scene type of the scene to be recognized.
[0100] It can be understood that based on the idea that the human eye generally discriminates the category of a scene according to the most representative features in the image, the embodiments of the present application use the significant features and supplementary features in the image to recognize the scene type, improving the scene recognition accuracy.
[0101] In a specific application, as Figure 2 shown, it is the specific implementation flow diagram of step S104 in the complex scene modeling method provided by the embodiments of the present application. Based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized, generate a target scene model, including: Figure 1
[0102] Step S202: Construct a first dense point cloud based on the first image to be processed, the second camera pose corresponding to the first image to be processed, and a preset first scene reconstruction algorithm.
[0103] The preset first scene reconstruction algorithm includes an AKAZE feature point matching pair algorithm, a depth estimation algorithm, and a 3D point cloud registration algorithm.
[0104] In specific applications, use the first image to be processed and the second camera pose corresponding to the first image to be processed to calculate the relative position relationship of the camera at different shooting points in the scene to be recognized. Calculate the AKAZE feature point matching pairs between the image taken by the camera at a certain point and the image taken at the previous point, and transmit them as input to the depth estimation model. The model outputs the initial point cloud of the feature points matched between the images. Use the 3D point cloud registration algorithm to perform matching and comparison on the initial point cloud, place the initial point clouds belonging to different spaces in different positions, and obtain the first dense point cloud by means of distance and reprojection.
[0105] Step S204: If the scene type of the scene to be recognized is the first scene, form a target scene model based on the first dense point cloud.
[0106] The first scene is the simple scene described above.
[0107] In specific applications, take each dense point cloud as the starting point and the corresponding camera as the ending point to draw a virtual straight line. The spaces passed by multiple virtual straight lines are intertwined to form a visible space, and the space surrounded by the rays is cut out; make a closed space based on the shortest path of graph theory to generate a preliminary 3D scene model. Finally, paste the image to be processed taken by the camera at a certain position in the space to the corresponding position on the 3D model to generate the target scene model.
[0108] Step S206: If the scene type of the scene to be recognized is the second scene, construct a second dense point cloud, and generate a target scene model based on the first dense point cloud and the second dense point cloud.
[0109] The second scene is the special scene described above.
[0110] Exemplarily, taking the scene to be recognized as a multi-floor indoor scene as an example, which includes the simple scenes on each floor and the special scenes of the stairs between each floor. Then, the first dense point cloud is the point cloud information representing each floor, and the second dense point cloud is the point cloud information representing the stairs between each floor. Then, a scene model of the multi-floor indoor scene can be obtained according to the point cloud information of each floor and the point cloud information of the stairs between each floor.
[0111] Specifically, as Figure 3 shown, it is for the complex scene modeling method provided by the embodiments of the present applicationFigure 2 Specific process schematic diagram of step S206. If the scene type of the scene to be recognized is the second scene, a second dense point cloud is constructed, and a target scene model is generated based on the first dense point cloud and the second dense point cloud, including:
[0112] Step S302: If the scene type of the scene to be recognized is the second scene, a shooting reminder instruction is generated and sent to the camera to instruct the camera to display the predicted shooting position to the user.
[0113] Among them, the predicted shooting position is the position that prompts the user to shoot in the second scene, that is, the feature scene.
[0114] Exemplarily, taking the scene to be recognized as a multi-floor indoor scene as an example, it includes simple scenes on each floor and special scenes of stairs between each floor. When the user starts using the camera to shoot on a floor and the server recognizes that the scene where the camera is located transitions from a floor to a staircase based on the images collected by the camera, the next shooting position on the staircase is predicted based on the current shooting position, and the user is prompted to shoot at the predicted shooting position on the staircase until the server recognizes that the scene where the camera is located is a floor based on the images collected by the camera.
[0115] Specifically, as Figure 4 shown, it is the specific process schematic diagram of step S302 in the complex scene modeling method provided by the embodiment of the present application. If the scene type of the scene to be recognized is the second scene, a shooting reminder instruction is generated and the shooting reminder instruction is sent to the camera to instruct the camera to display the predicted shooting position to the user. The predicted shooting position is the position that instructs the user to shoot in the second scene, including: Figure 3
[0116] Step S402: Obtain the predicted shooting position according to a preset position prediction algorithm and the first camera pose.
[0117] Among them, the preset position prediction algorithm may include the Logistics algorithm, decision tree algorithm, etc. In a specific application, first calculate the predicted value of the next camera pose according to the LK optical flow method based on the first camera pose, and then substitute the predicted value into the preset position prediction algorithm to obtain the true value of the predicted shooting position.
[0118] Step S404: Generate a shooting reminder instruction according to the predicted shooting position.
[0119] Step S406: Send the shooting reminder instruction to the user.
[0120] In the embodiment of the present application, when it is recognized that the scene type of the scene to be recognized is a special scene, the next shooting position in the special scene can be predicted according to the current shooting position of the camera, which is convenient for the user to select a suitable shooting position for shooting.
[0121] Step S304: Obtain the second image to be processed and the second camera pose corresponding to the second image to be processed.
[0122] Among them, the second image to be processed is a depth image obtained by the camera at the shooting point in a special scenario, and the second camera pose is the IMU data collected by the IMU control unit of the camera when the camera captures the image to be processed.
[0123] Step S306: Generate a second dense point cloud according to the second image to be processed, the second camera pose corresponding to the second image to be processed, and a preset second scene reconstruction algorithm.
[0124] Among them, the preset second scene reconstruction algorithm includes the ORB feature descriptor algorithm, the TSDF algorithm, and the MVS algorithm.
[0125] Specifically, as Figure 5 shown, it is a schematic diagram of the specific process of step S306 in the complex scene modeling method provided by the embodiment of the present application. Generating a second dense point cloud according to the second image to be processed, the second camera pose corresponding to the second image to be processed, and a preset second scene reconstruction algorithm includes: Figure 3
[0126] Step S502: Generate a point cloud to be processed according to the image to be processed and the camera pose corresponding to the image to be processed.
[0127] In specific applications, key frames in the image to be processed are extracted according to the ORB feature descriptor algorithm, the time stamps and camera poses corresponding to the key frames are determined, and based on the SFM algorithm, a point cloud to be processed is generated according to the image to be processed and the camera pose corresponding to the image to be processed.
[0128] Step S504: Fuse the point cloud to be processed based on the TSDF algorithm to obtain a fused point cloud.
[0129] In specific applications, the TSDF reconstructs the point cloud data set according to the key frame information, realizes a controllable point cloud density and reduces unnecessary repeated calculations by pre - constructing a three - dimensional space to obtain a fused point cloud.
[0130] Step S506: Perform statistical filtering on the fused point cloud to obtain an optimized fused point cloud.
[0131] It can be understood that statistical filtering is performed on the reconstructed three - dimensional environmental point cloud to optimize the point cloud quality.
[0132] Step S508: Densify the optimized fused point cloud based on the MVS algorithm to obtain a second dense point cloud.
[0133] In specific applications, photometric consistency constraints and visibility constraints are imposed on the fused point cloud to obtain a second dense point cloud.
[0134] Step S308: Register the first dense point cloud and the second dense point cloud.
[0135] Among them, the registration algorithms include but are not limited to the iterative closest point algorithm, the second type of point cloud registration algorithm, the robust point matching algorithm, or the fifth type of point cloud registration method, etc.
[0136] It can be understood that registering the first dense point cloud and the second dense point cloud enables obtaining a target scene model based on the registered first dense point cloud and second dense point cloud.
[0137] Step S310: Form a target scene model based on the registered first dense point cloud and second dense point cloud.
[0138] In specific applications, based on the Marching Cube algorithm, the registered first dense point cloud and second dense point cloud are subjected to triangular meshing to obtain a target scene model.
[0139] It can be understood that point cloud data is discretely represented in three-dimensional space. The Marching Cube algorithm is used for the registered first dense point cloud and second dense point cloud to extract the isosurface, realize triangular mesh reconstruction, and obtain a target scene model.
[0140] In the embodiments of the present application, a first image to be processed and a first camera pose corresponding to the first image to be processed are obtained. The image to be processed is a depth image captured by a camera at different shooting positions in the scene to be recognized; based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized, a target scene model is generated. It can be seen that the present application can perform real-time automatic modeling in complex scenes. Since the connection between scenes in the directly generated target scene model is complete, relatively little computing resources are required for rendering processing in the later stage (i.e., after modeling), and the three-dimensional reconstruction rendering effect is also improved, making the viewing effect better when the user roams in the three-dimensional model of the complex scene from the first perspective.
[0141] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0142] Corresponding to the complex scene modeling method described in the above embodiments, Figure 6 The structural block diagram of a complex scene modeling device provided by the embodiments of the present application is shown. For ease of description, only the parts related to the embodiments of the present application are shown.
[0143] Reference Figure 6 , the device includes:
[0144] An acquisition module 61, configured to acquire a first image to be processed and a first camera pose corresponding to the first image to be processed, where the image to be processed is an image captured by a camera in a scene to be recognized in response to a user's shooting reminder instruction;
[0145] A generation module 62, configured to generate a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized.
[0146] In a possible implementation manner, the device further includes:
[0147] An identification module, configured to determine the scene type of the scene to be recognized according to a pre-trained scene recognition model.
[0148] In a possible implementation manner, the pre-trained scene recognition model includes a feature extraction layer, a feature selection layer, and a classification layer;
[0149] The identification module includes:
[0150] A first processing sub-module, configured to import the image to be processed into the feature extraction layer and output significant features and supplementary features;
[0151] A second processing sub-module, configured to import the significant features into a feature selector and output target significant features;
[0152] An extraction sub-module, configured to extract local representation information of the target significant features and global representation information in the supplementary features respectively;
[0153] A classification sub-module, configured to splice the local representation information and the global representation information and then import them into the classification layer to output the scene type of the scene to be recognized.
[0154] In a possible implementation manner, the generation module includes:
[0155] A construction sub-module, configured to construct a first dense point cloud according to the first image to be processed, a second camera pose corresponding to the first image to be processed, and a preset first scene reconstruction algorithm;
[0156] A first generation sub-module, configured to form a target scene model according to the first dense point cloud if the scene type of the scene to be recognized is a first scene;
[0157] A second generation sub-module, configured to, if the scene type of the scene to be recognized is a second scene, construct a second dense point cloud, and generate a target scene model according to the first dense point cloud and the second dense point cloud.
[0158] In a possible implementation manner, the first generation sub-module includes:
[0159] A generation unit, configured to, if the scene type of the scene to be recognized is a second scene, generate a shooting reminder instruction, and send the shooting reminder instruction to a camera to indicate the camera to display a predicted shooting position to the user, where the predicted shooting position is a position indicating the user to shoot in the second scene;
[0160] An acquisition unit, configured to acquire a second image to be processed and a second camera pose corresponding to the second image to be processed;
[0161] A second generation unit, configured to generate a second dense point cloud according to the second image to be processed, the second camera pose corresponding to the second image to be processed, and a preset second scene reconstruction algorithm;
[0162] A registration unit, configured to register the first dense point cloud and the second dense point cloud;
[0163] A third generation unit, configured to form a target scene model according to the registered first dense point cloud and the second dense point cloud.
[0164] In a possible implementation manner, the second generation sub-module includes:
[0165] A fourth generation unit, configured to generate a point cloud to be processed according to an image to be processed and a camera pose corresponding to the image to be processed;
[0166] A fusion unit, configured to fuse the point cloud to be processed based on the TSDF algorithm to obtain a fused point cloud;
[0167] An optimization unit, configured to perform statistical filtering on the fused point cloud to obtain an optimized fused point cloud;
[0168] A densification processing unit, configured to perform densification processing on the optimized fused point cloud based on the MVS algorithm to obtain a second dense point cloud.
[0169] In a possible implementation manner, the generation unit includes:
[0170] A prediction sub-unit, configured to obtain a predicted shooting position according to a preset position prediction algorithm and the first camera pose;
[0171] A generation sub-unit, configured to generate a shooting reminder instruction according to the predicted shooting position.
[0172] It should be noted that for the content such as information interaction and execution process between the above-mentioned devices / units, since it is based on the same concept as the method embodiments of this application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.
[0173] Figure 7 This is a schematic structural diagram of a server provided by an embodiment of this application. As Figure 7 shown, the server 7 in this embodiment includes: at least one processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the at least one processor 70. When the processor 70 executes the computer program 72, the steps in any of the above method embodiments are implemented.
[0174] The server 7 may be a computing device such as a cloud server. The server may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art can understand that Figure 7 merely examples of the server 7 are given, which do not constitute a limitation on the server 7. It may include more or fewer components than those shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0175] The so-called processor 70 may be a central processing unit (CPU). The processor 70 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0176] The memory 71 may be an internal storage unit of the server 7 in some embodiments, such as the hard disk or memory of the server 7. The memory 71 may also be an external storage device of the server 7 in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the server 7. Further, the memory 71 may also include both the internal storage unit of the server 7 and the external storage device. The memory 71 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 71 may also be used to temporarily store data that has been output or will be output.
[0177] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0178] The embodiment of the present application also provides a readable storage medium, specifically a computer-readable storage medium. The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.
[0179] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the server, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0180] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0181] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0182] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0183] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0184] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A complex scene modeling method, characterized in that Including: Obtain a first image to be processed and a first camera pose corresponding to the first image to be processed, where the image to be processed is a depth image obtained by the camera at different shooting positions in the scene to be recognized; Generate a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized, where the scene type of the scene to be recognized includes a first scene and a second scene; Generating a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized includes: constructing a first dense point cloud according to the first image to be processed, the second camera pose corresponding to the first image to be processed, and a preset first scene reconstruction algorithm; If the scene type of the scene to be recognized is the first scene, form a target scene model according to the first dense point cloud; if the scene type of the scene to be recognized is the second scene, construct a second dense point cloud, and generate a target scene model according to the first dense point cloud and the second dense point cloud.
2. The complex scene modeling method according to claim 1, wherein, Before generating a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized, it further includes: Determine the scene type of the scene to be recognized according to a pre-trained scene recognition model.
3. The complex scene modeling method according to claim 2, wherein The pre-trained scene recognition model includes a feature extraction layer, a feature selection layer, and a classification layer; Determining the scene type of the scene to be recognized according to a pre-trained scene recognition model includes: Import the image to be processed into the feature extraction layer, and output significant features and supplementary features; Import the significant features into a feature selector, and output target significant features; Extract local representation information of the target significant features and global representation information in the supplementary features respectively; Concatenate the local representation information and the global representation information and import them into the classification layer, and output the scene type of the scene to be recognized.
4. The complex scene modeling method according to any one of claims 1 to 3, characterized in that If the scene type of the scene to be recognized is the second scene, constructing a second dense point cloud, and forming a target scene model according to the first dense point cloud and the second dense point cloud includes: If the scene type of the scene to be recognized is the second scene, generate a shooting reminder instruction, and send the shooting reminder instruction to the camera to indicate the camera to display a predicted shooting position to the user, where the predicted shooting position is the position indicating the user to shoot in the second scene; Obtain a second image to be processed and a second camera pose corresponding to the second image to be processed; Generate a second dense point cloud according to the second image to be processed, the second camera pose corresponding to the second image to be processed, and a preset second scene reconstruction algorithm; Register the first dense point cloud and the second dense point cloud; Form a target scene model according to the registered first dense point cloud and the second dense point cloud.
5. The complex scene modeling method according to claim 4, wherein, Generating a second dense point cloud according to the second image to be processed, the second camera pose corresponding to the second image to be processed, and a preset second scene reconstruction algorithm includes: Generate a point cloud to be processed based on the image to be processed and the camera pose corresponding to the image to be processed; Fuse the point cloud to be processed based on the TSDF algorithm to obtain a fused point cloud; Perform statistical filtering on the fused point cloud to obtain an optimized fused point cloud; Densely process the optimized fused point cloud based on the MVS algorithm to obtain a second dense point cloud.
6. The complex scene modeling method according to claim 4, characterized in that If the scene type of the scene to be recognized is the second scene, generate a shooting reminder instruction, and send the shooting reminder instruction to the camera to instruct the camera to display the predicted shooting position to the user. The predicted shooting position is the position that prompts the user to shoot in the second scene, including: Obtain the predicted shooting position according to the preset position prediction algorithm and the first camera pose; Generate a shooting reminder instruction according to the predicted shooting position; Send the shooting reminder instruction to the user.
7. A complex scene modeling device, characterized in that, Including: An acquisition module for acquiring a first image to be processed and a first camera pose corresponding to the first image to be processed. The image to be processed is an image captured by the camera in response to the user's shooting reminder instruction in the scene to be recognized; A generation module for generating a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized. The scene type of the scene to be recognized includes a first scene and a second scene; Generating a target scene model based on the first image to be processed, the first camera pose corresponding to the first image to be processed, and the scene type of the scene to be recognized includes: constructing a first dense point cloud according to the first image to be processed, the second camera pose corresponding to the first image to be processed, and a preset first scene reconstruction algorithm; If the scene type of the scene to be recognized is the first scene, form a target scene model according to the first dense point cloud; if the scene type of the scene to be recognized is the second scene, construct a second dense point cloud, and generate a target scene model according to the first dense point cloud and the second dense point cloud.
8. A server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 6 is implemented.
9. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Monitoring and early warning method and system based on dynamic three-dimensional model and storage medium
CN112053391A
Method and device for training deep learning model
CN112639846A