Neural map building method and apparatus
By constructing neural maps and using data of aerial images and ground-based images, the problem that traditional map construction methods cannot achieve high-precision rendering and positioning is solved, high-precision rendering and positioning of the entire airspace is achieved, and the level of urban governance planning is improved.
Patent Information
- Application Number
- PCT/CN2024/099954
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-06-18
- Publication Date
- 2025-05-30
AI Technical Summary
The traditional map construction method is relatively single, and it is impossible to achieve high-precision rendering and positioning at the same time in large-scene modeling.
By acquiring data of aerial images and ground-based images, using position information to align pixel feature information, and construct neural maps based on this information to achieve high-precision rendering and positioning.
It realizes high-precision rendering and positioning in the entire airspace, meets the service needs of large-scene modeling, and improves the level of comprehensive urban governance planning.
Smart Images

Figure CN2024099954_30052025_PF_FP_ABST
Abstract
Description
Neural map construction method and device Technical Field
[0001] The present application relates to the field of cloud technology, and in particular to a method and device for constructing a neural map. Background Art
[0002] The concept of a city digital twin emphasizes the creation of a virtual city that interacts with the physical city in real time, accurately mapping the physical city's operations and forming a virtual-reality interaction pattern to enhance and optimize the city's comprehensive governance and planning. A city digital twin is not a single technology, but rather a building block assembly of numerous cutting-edge technologies. Mapping technology plays a crucial role, deeply involved in the construction of virtual models and the realization of virtual-reality interaction. Common map formats include 3D city maps, AR maps, and autonomous driving maps. Traditional map construction methods are relatively simple and cannot simultaneously achieve high-precision rendering and positioning capabilities in large-scale scene modeling.
[0003] Summary of the Invention
[0004] This application provides a method and device for constructing a neural map to meet the service needs of high-precision rendering and positioning of the entire airspace, so as to solve the problem that traditional map construction methods are relatively single and cannot achieve high-precision rendering and positioning simultaneously in large-scene modeling.
[0005] In a first aspect, the present application provides a method for constructing a neural map, which is applied to a cloud server, and includes: obtaining first acquisition device data collected by a first acquisition device located under the cloud, and second acquisition device data collected by a second acquisition device located under the cloud, wherein the first acquisition device data is an aerial image of the first shooting scene, and the second acquisition device data is a ground image of the first shooting scene; aligning the first pose information of the aerial image and the second pose information of the ground image to the same geographic coordinate system; extracting the first feature information of each pixel of the aerial image and the second feature information of each pixel of the ground image respectively; constructing a neural map of the first shooting scene based on the geographic coordinate system aligned with the first pose information and the second pose information, the first feature information and the second feature information, the neural map includes multiple grids, each grid includes coordinate position information and third feature information, the coordinate position information is the spatial position information of the pixel mapping of the aerial image and / or the ground image of the first shooting scene, and the third feature information is the fusion information of the first feature information of each pixel of the aerial image and the second feature information of each pixel of the ground image corresponding to the spatial position.
[0006] Based on the above method, different acquisition devices are used to collect aerial images and ground images of the same shooting scene respectively, and then the posture information and pixel feature information of the aerial images and ground images are used to construct a neural map. The neural map contains richer and more accurate information about aerial images and / or ground images, thereby meeting the service requirements of high-precision rendering and positioning of the entire airspace.
[0007] In a possible implementation of the first aspect, the method also includes: calculating the third pose information of the aerial image using the positioning model based on the neural map and the aerial image, obtaining first error data between the first pose information and the third pose information, and feeding back the first error data to the neural map and the positioning model; calculating the fourth pose information of the ground image using the positioning model based on the neural map and the ground image, obtaining second error data between the second pose information and the fourth pose information, and feeding back the second error data to the neural map and the positioning model; updating the feature information in the neural map and the parameter information of the positioning model based on the first error data and the second error data.
[0008] Based on the above method, aerial images and / or ground images are used to update the parameter information in the neural map through the positioning model, thereby realizing continuous training of the neural map in the positioning mode, making the data in the neural map and the positioning model more accurate.
[0009] In a possible implementation of the first aspect, the method further includes: calculating a first rendered image of the aerial image using a rendering model based on the neural map and the first pose information, obtaining third error data between the first rendered image and the aerial image, and feeding the third error data back to the neural map and the rendering model; calculating a second rendered image of the ground image using a rendering model based on the neural map and the second pose information, obtaining fourth error data between the second rendered image and the ground image, and feeding the fourth error data back to the neural map and the rendering model; and updating the feature information in the neural map and the parameter information of the rendering model based on the third error data and the fourth error data.
[0010] Based on the above method, aerial images and / or ground images are used to update the parameter information in the neural map through the rendering model, thereby realizing continuous training of the neural map in the rendering mode, making the data in the neural map and the rendering model more accurate.
[0011] In a possible implementation manner of the first aspect, the method further includes: encoding stylized labels on the aerial image and the terrestrial image, where the stylized labels are used to indicate image style types to which the aerial image and the terrestrial image belong.
[0012] Based on the above method, by encoding the image style types of aerial images and ground images with stylized labels, the information of aerial images and ground images stored in the neural map can be made more comprehensive and the classification more accurate, which is conducive to the subsequent rendering model to perform stylized rendering of images.
[0013] In a possible implementation of the first aspect, the method further includes: receiving an image to be positioned of a second shooting scene, the second shooting scene being located within the range of the first shooting scene, and calculating fifth pose information of the image to be positioned using a positioning model based on the neural map and the image to be positioned.
[0014] Based on the above method, the trained neural map and positioning model can be used to locate images within the same shooting scene. Since the neural map is pre-constructed and trained based on aerial and ground images of the same shooting scene, the data obtained in the positioning process is more accurate.
[0015] In a possible implementation manner of the first aspect, the method further includes: receiving sixth posture information of the second shooting scene, and calculating a third rendered image corresponding to the sixth posture information using a rendering model based on the neural map and the sixth posture information.
[0016] Based on the above method, the trained neural map and rendering model can be used to render images within the same shooting scene. Since the neural map is pre-constructed and rendered based on aerial and ground images of the same shooting scene, the data obtained in the rendering process is more accurate.
[0017] In a possible implementation of the first aspect, seventh pose information and a first stylized label of a second captured scene are received, and a fourth rendered image corresponding to the seventh pose information is calculated using a rendering model based on the neural map, the seventh pose information, and the first stylized label. The first stylized label is used to indicate an image style type to which the fourth rendered image belongs.
[0018] Based on the above method, by defining stylized tag information, the rendering style type of the image can be customized, making the image style type of the image rendering more diverse.
[0019] In a possible implementation of the first aspect, a partial area selection instruction for a first image of a third shooting scene is received, the third shooting scene is located within the range of the first shooting scene, first spatial information corresponding to the partial area of the first image in the first image is identified, and first grid information corresponding to the first spatial information in the neural map is located; a partial area migration instruction for a second image of the third shooting scene is received, second spatial information corresponding to the partial area of the second image in the second image is identified, and second grid information corresponding to the second spatial information in the neural map is located; the second grid information in the neural map is replaced with the first grid information, and the neural map is updated.
[0020] Based on the above method, by replacing different grid information in the neural map, the synchronous replacement of the image background can be achieved more quickly and accurately.
[0021] In a possible implementation manner of the first aspect, the first feature information, the second feature information, and the third feature information include semantic information, point, line, and surface information, density information, lighting information, color information, texture information, and material information corresponding to the first shooting scene.
[0022] Based on the above method, by extracting rich feature information of each pixel of aerial images and ground images, the types of fused information in the neural map are more diversified, the neural map is more accurate in image positioning, and the image rendering is more precise.
[0023] In a second aspect, the present application also provides a device for constructing a neural map, which is applied to a cloud server and includes: an acquisition module for acquiring first acquisition device data collected by a first acquisition device located under the cloud, and second acquisition device data collected by a second acquisition device located under the cloud, wherein the first acquisition device data is an aerial image of the first shooting scene, and the second acquisition device data is a ground image of the first shooting scene; an alignment module for aligning the first pose information of the aerial image and the second pose information of the ground image to the same geographic coordinate system; an extraction module for respectively extracting the first feature information of each pixel of the aerial image and the second feature information of each pixel of the ground image; a construction module for constructing a neural map of the first shooting scene based on the geographic coordinate system aligned with the first pose information and the second pose information, the first feature information and the second feature information, the neural map including multiple grids, each grid including coordinate position information and feature information, the coordinate position information being the spatial position information of the pixel point mapping of the aerial image and / or the ground image of the first shooting scene, and the feature information being the fusion information of the first feature information of each pixel of the aerial image and the second feature information of each pixel of the ground image corresponding to the spatial position.
[0024] The second aspect or any implementation method of the second aspect is the step implementation of the device corresponding to the first aspect or any implementation method of the first aspect. The description in the second aspect or any implementation method of the second aspect is applicable to the first aspect or any implementation method of the first aspect and will not be repeated here.
[0025] In a third aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method disclosed in the first aspect and any possible implementation of the first aspect.
[0026] In a fourth aspect, the present application provides a computer program product comprising instructions, which, when executed by a computer device cluster, enables the computer device cluster to implement the method disclosed in the first aspect and any possible implementation manner of the first aspect.
[0027] In a fifth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method disclosed in the first aspect and any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] FIG1 is a flow chart of a neural map construction method according to an embodiment of the present invention;
[0029] FIG2 is a schematic diagram of an image captured by a capture device according to an embodiment of the present invention;
[0030] FIG3 a is a schematic diagram of a point cloud of an aerial image according to an embodiment of the present invention;
[0031] FIG3 b is a schematic diagram of a point cloud of a ground-captured image according to an embodiment of the present invention;
[0032] FIG4 is a schematic diagram of a flow chart of positioning model training in a neural map construction method according to an embodiment of the present invention;
[0033] FIG5 is a schematic diagram of a process for training a rendering model in a neural map construction method according to an embodiment of the present invention;
[0034] FIG6 is a schematic diagram of implementing style transfer in a rendering application based on a neural map according to an embodiment of the present invention;
[0035] FIG7 is a schematic diagram of a map replacement application based on a neural map according to an embodiment of the present invention;
[0036] FIG8 is a schematic diagram of a neural map update according to an embodiment of the present invention;
[0037] FIG9 is a flowchart of a neural map construction method according to an embodiment of the present invention;
[0038] FIG10 is a schematic diagram of a neural map construction device according to an embodiment of the present invention;
[0039] FIG11 is a schematic diagram of the computing device structure of the method for constructing a neural map according to an embodiment of the present invention;
[0040] FIG12 is a schematic diagram of a computing device cluster structure of a method for constructing a neural map according to an embodiment of the present invention. DETAILED DESCRIPTION
[0041] The following describes the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0042] In order to facilitate understanding of the embodiments of the present application, first, some terms involved in the present invention are explained.
[0043] Neural map: A neural map is a map constructed by collecting various data information through sensors, extracting and fusing the features of this data information as needed, and then binding them with spatial coordinates.
[0044] Pose: Pose refers to the position and orientation of an object, robot, or person's posture or gesture in three-dimensional space. It consists of two components: position and orientation. Position represents the coordinates of an object's center or reference point in three-dimensional space and is typically represented using three real numbers. Orientation represents the orientation or direction of an object in three-dimensional space and is typically represented using rotation matrices, Euler angles, or quaternions.
[0045] Feature fusion: Feature fusion is the process of extracting features of different attributes from sensor data, using the complementarity between features and the advantages of fused features to construct new fused features, thereby improving the performance of the model.
[0046] FIG1 is a flow chart of a method for constructing a neural map according to an embodiment of the present invention. As shown in FIG1 , an embodiment of the present invention provides a method for constructing a neural map, which is applied to a cloud server and includes:
[0047] 1. Data collection stage:
[0048] Acquire the acquisition device data A1 of the device A located under the cloud, and the acquisition device data B1 of the acquisition device B located under the cloud, wherein the acquisition device data A1 is an aerial image of the shooting scene M, and the acquisition device data B1 is a ground image of the shooting scene M.
[0049] Specifically, the acquisition device A and the acquisition device B can be sensor devices, such as satellite collectors, drones, panoramic cameras, laser scanners, cameras, video cameras, infrared cameras, etc.
[0050] Figure 2 is a schematic diagram of images collected by the collection device in an embodiment of the present invention. As shown in Figure 2, the collection device data A1 and the collection device data B1 collected by the collection device A and the collection device B respectively can be multi-perspective satellite images collected by a satellite collector, drone bird's-eye views collected by a drone, image sequences collected by a panoramic camera, laser point clouds collected by a laser scanner, etc.
[0051] The acquisition device A and the acquisition device B can be the same type of acquisition device, or they can be different types of acquisition devices that have advantages in shooting in the aerial range and in the ground range respectively for the shooting scene M. For example, the acquisition device A is a drone that has advantages in shooting in the aerial range of the shooting scene M, and the acquisition device B is a panoramic camera that has advantages in shooting in the ground range of the shooting scene M.
[0052] It should be noted that the description of the device types of the above-mentioned acquisition devices A and acquisition devices B and the range of the pictures taken are all cases in the embodiments of the present invention. In other embodiments, there are other cases similar to the central idea of the present invention regarding the device types of the acquisition devices and the range of the pictures taken. The types of the acquisition devices and the range of the pictures taken are not limited here.
[0053] 2. Map construction phase:
[0054] 1) Align the pose information A2 of the aerial image and the pose information B2 of the ground image to the same geographic coordinate system.
[0055] Specifically, the aerial image is sparsely reconstructed and / or densely reconstructed to obtain a point cloud A4 and pose information A2 in the aerial image coordinate system, and the ground image is sparsely reconstructed and / or densely reconstructed to obtain a point cloud B4 and pose information B2 in the ground image coordinate system. The coordinate system of the ground image is used as a reference coordinate system, and the point cloud A4 of the aerial image and the point cloud B4 of the ground image are point cloud registered to obtain the transformation relationship from the aerial image coordinate system to the ground image coordinate system, and then the pose information A2 of the aerial image is aligned to the coordinate system of the ground image. Figure 3a is a schematic diagram of a point cloud of an aerial image according to an embodiment of the present invention. As shown in Figure 3a, the point cloud A4 of the aerial image includes the pose information A2 of the aerial image. Figure 3b is a schematic diagram of a point cloud of a ground image according to an embodiment of the present invention. As shown in Figure 3b, the point cloud B4 of the ground image includes the pose information B2 of the ground image. Among them, the posture information A2 can be understood as the shooting trajectory formed when the acquisition device A shoots the shooting scene M within the aerial range, and the posture information B2 can be understood as the shooting trajectory formed when the acquisition device B shoots the shooting scene M within the ground range.
[0056] Furthermore, in the coordinate system of the ground-captured image, image feature matching can be performed on the ground-captured image and the aerial-captured image, and the posture information A2 of the aerial-captured image can be further optimized, so that the result of aligning the posture information A2 of the aerial-captured image to the coordinate system of the ground-captured image is more accurate.
[0057] It should be noted that in the process of aligning the geographic coordinate system, the coordinate system of the aerial image can be aligned to the coordinate system of the ground image, the coordinate system of the ground image can be aligned to the coordinate system of the aerial image, or the aerial image and the ground image can be aligned to other geographic coordinate systems. There is no limitation on the alignment method of the geographic coordinate system here.
[0058] 2) Extract the feature information A3 of each pixel of the aerial image and the feature information B3 of each pixel of the ground image respectively.
[0059] Specifically, aerial images are fed into a deep convolutional neural network to obtain a feature vector for each pixel in the aerial image. Similarly, terrestrial images are fed into a deep convolutional neural network to obtain a feature vector for each pixel in the terrestrial image. For example, if an aerial or terrestrial image has a pixel height of 1000 and a pixel width of 2000, and a feature vector length of 256, the resulting feature information for that aerial or terrestrial image will have a dimension of 1000x2000x256.
[0060] 3) Based on the geographic coordinate system aligned with the posture information A2 and the posture information B2, the feature information A3 and the feature information B3, a neural map of the shooting scene M is constructed. The neural map includes multiple grids, each grid includes coordinate position information C1 and feature information C2. The coordinate position information C1 is the spatial position information of the pixel point mapping of the aerial image and / or the ground image of the shooting scene M, and the feature information C2 is the fusion information of the feature information A3 of each pixel in the aerial image and the feature information B3 of each pixel in the ground image corresponding to the spatial position.
[0061] Specifically, a neural map of the shooting scene M is constructed based on the coordinate system, feature information A3 and feature information B3 of the ground-captured image aligned with the pose information A2 and the pose information B2. The neural map is an initial neural map that has not been trained. The neural map represents the feature vectors stored in the three-dimensional space. For example, a three-dimensional grid is divided in a three-dimensional space with a length of L = 100m, a width of W = 50m, and a height of H = 10m. The resolution of the grid is d = 0.1m, and the grid dimension corresponding to the three-dimensional space is M = L / d = 1000, N = W / d = 500, and O = H / d = 100, that is, MxNxO = 1000x500x100. The length of the feature vector stored in each grid is D = 256, and the dimension of the feature vector stored in the neural map is MxNxOxD = 1000x500x100x256.
[0062] Furthermore, each grid coordinate in the three-dimensional space (i.e., coordinate position information C1, typically expressed as grid center coordinates) is projected onto the aerial image and / or terrestrial image to obtain a feature vector (i.e., feature information C2) for the corresponding pixel coordinate. The pixel coordinate is the intersection coordinate of the line connecting the grid coordinate and the camera optical center coordinate and the camera imaging plane. Because each grid within the imaging range may obtain one or more feature vectors from one or more aerial and / or terrestrial images, a network such as a multilayer perceptron is used to aggregate these multiple feature vectors into a single feature vector, thereby achieving feature fusion of feature information A3 for each pixel in the aerial image and feature information B3 for each pixel in the terrestrial image. Ultimately, each grid contains only one fused feature vector (i.e., feature information C2).
[0063] 3. Map training phase:
[0064] The constructed neural map still requires a corresponding training process before it can be used. The following examples illustrate the steps of positioning model training and rendering model training of the neural map. In other possible embodiments, the neural map can also be trained for other models as needed. The model training method of the neural map is not limited here.
[0065] 1) Positioning model training
[0066] Based on the neural map and the aerial image, the positioning model is used to calculate the posture information D1 of the aerial image, and the error data E1 between the posture information A2 and the posture information D1 is obtained, and the error data E1 is fed back to the neural map and the positioning model; based on the neural map and the ground image, the positioning model is used to calculate the posture information D2 of the ground image, and the error data E2 between the posture information B2 and the posture information D2 is obtained, and the error data E2 is fed back to the neural map and the positioning model; based on the error data E1 and the error data E2, the feature information C2 in the neural map and the parameter information of the positioning model are updated.
[0067] Specifically, after the neural map is constructed through a large number of image sets of aerial images and / or ground images, it is necessary to use any image in the image set as an input image for positioning model training to train the neural map. Figure 4 is a flow chart of positioning model training in a neural map construction method according to an embodiment of the present invention. As shown in Figure 4, for example: we input a ground image, and the neural map determines the preliminary retrieval range of the ground image in the neural map based on the GPS positioning information of the ground image itself, obtains the grid area L of the three-dimensional space corresponding to the retrieval range, extracts the coordinate position information and feature information contained in the grid area L, and decodes the feature information through the positioning decoder to obtain a feature vector for positioning. The feature vector includes semantic information, point, line, and surface information, and structured information of the ground image. Through the above information, finally The grid coordinates and positioning feature information of the three-dimensional space corresponding to the retrieval range, namely the positioning feature map W1, are obtained. The positioning model uses the feature extraction network to extract the three-dimensional cone feature map W2 of the ground-captured image. By matching the three-dimensional cone feature map W2 with the positioning feature map W1 corresponding to the retrieval range, the pose information D1 corresponding to the ground-captured image is calculated. At the same time, the error data E1 between the pose information D1 and the true pose information A2 of the ground-captured image is calculated. The error data E1 is the pose loss function. After the gradient is returned, the pose loss function continuously updates the feature information C2 in the neural map and the parameter information of the positioning model.
[0068] It should be noted that the process of training the neural map using the positioning model using aerial images is similar to the process of training the neural map using the ground-based images mentioned above, and will not be repeated here. At the same time, the calculation formula of the pose loss function shown in FIG4 is only a calculation formula of a loss function in the embodiment of the present invention. In other embodiments, the calculation formula of the pose loss function can also be other formulas, and the calculation formula of the pose loss function is not limited here.
[0069] 2) Rendering model training
[0070] Based on the neural map and posture information A2, the rendering model is used to calculate the rendered image F1 of the aerial image, and the error data G1 between the rendered image F1 and the aerial image is obtained, and the error data G1 is fed back to the neural map and the rendering model; based on the neural map and posture information B2, the rendering model is used to calculate the rendered image F2 of the ground image, and the error data G2 between the rendered image F2 and the ground image is obtained, and the error data G2 is fed back to the neural map and the rendering model; based on the error data G1 and the error data G2, the feature information C2 in the neural map and the parameter information of the rendering model are updated.
[0071] Figure 5 is a flow chart of rendering model training in a neural map construction method according to an embodiment of the present invention. As shown in Figure 5, the camera pose (position and angle) and camera parameters (focal length, image resolution, etc.) of the acquisition device B for acquiring aerial images are input, and the three-dimensional space area covered by the camera shooting angle range is determined by the camera pose and camera parameters. The grid area X corresponding to the three-dimensional space area in the neural map is extracted, and the coordinate position information and feature information contained in the grid area X are extracted. The feature information is decoded by a rendering decoder to obtain a feature vector for rendering. The feature vector includes density information, color information and signed distance field information of the ground shooting image. Through the above information, the grid coordinates and rendering feature information of the three-dimensional space corresponding to the camera shooting angle range are finally obtained, that is, the rendering feature map Z. Using the rendering model and neural rendering method, the rendered image F1 corresponding to the aerial image is calculated by rendering the feature map Z1. At the same time, we need to calculate the error data G2 between the rendered image F1 and the real image Z2 of the aerial image. The error data G2 is the pixel color loss function. After the pixel color loss function is returned through the gradient, the feature information C2 in the neural map and the parameter information of the rendering model are continuously updated.
[0072] It should be noted that the process of training the neural map by rendering the model using ground-captured images is similar to the process of training the neural map by the acquisition device B that inputs aerial images, and will not be repeated here.
[0073] Furthermore, in order to diversify the style of rendered images, diversified collection is performed for different weather, lighting and seasons when collecting aerial images and ground images. After the rendering training of aerial images and ground images is completed, the stylized label codes of aerial images and ground images can be obtained. The stylized label codes are used to indicate the image style types to which the aerial images and ground images belong, so that users can input the stylized label codes to customize the style type of image rendering when using the neural map later.
[0074] 4. Map application stage:
[0075] The trained neural map can be released to users for use. The following examples illustrate the steps of the neural map's image positioning application, image rendering application, image fusion application, and image replacement application. In other possible embodiments, the neural map can also be used for other applications as needed. The application method of the neural map is not limited here.
[0076] 1) Image positioning application
[0077] During image positioning, the user inputs an image to be positioned, which can be an aerial image or a ground image. The trained neural map determines the preliminary retrieval range of the image to be positioned in the neural map based on the GPS positioning information of the image to be positioned itself, obtains the grid area of the three-dimensional space corresponding to the retrieval range, extracts the coordinate position information and feature information contained in the grid area, and decodes the feature information through the positioning decoder to obtain a feature vector for positioning. The feature vector includes semantic information, point, line, and surface information, and structured information of the image to be positioned. Through the above information, the grid coordinates and positioning feature information of the three-dimensional space corresponding to the image to be positioned are finally obtained, that is, the positioning feature map. The positioning model uses the three-dimensional cone feature map calculation method to calculate the posture information corresponding to the image to be positioned through the positioning feature map, thereby completing the positioning process of the image to be positioned.
[0078] 2) Image rendering applications
[0079] When rendering an image, the user inputs the camera pose (pose and angle) and camera parameters (focal length, image resolution, etc.) of the acquisition device. The three-dimensional space area covered by the camera's shooting angle of view is determined by the camera pose and camera parameters. The grid area corresponding to the three-dimensional space area in the trained neural map is extracted, and the coordinate position information and feature information contained in the grid area are extracted. The feature information is decoded by the rendering decoder to obtain a feature vector for rendering. The feature vector includes density information, color information, and signed distance field information of the image to be rendered. Through the above information, the grid coordinates and rendering feature information of the three-dimensional space corresponding to the camera's shooting angle of view are finally obtained, that is, the rendering feature map. Using the rendering model and neural rendering method, the rendered image corresponding to the aerial image is calculated through the rendering feature map, and the rendering process of the image to be rendered is completed.
[0080] Furthermore, when rendering an image, users can input stylized tags as needed. The trained neural map will then transfer the style of the rendered image based on the stylized tags input by the user, thereby obtaining the desired rendered image. Figure 6 is a schematic diagram of an embodiment of the present invention implementing style transfer in a rendering application based on a neural map. As shown in Figure 6, different images are rendered when the stylized tags input are sunny, night, and winter.
[0081] It should be noted that the stylized labels here can be text labels such as sunny day, night, and winter, or they can be digital labels, letter labels, etc. There is no limitation on the labeling method here.
[0082] 3) Image replacement application
[0083] Figure 7 is a schematic diagram of an embodiment of the present invention that implements an image replacement application based on a neural map. As shown in Figure 7, a user can click on an object in image scene A and click on a space in image scene B to calculate the feature information A5 of the three-dimensional spatial area occupied in the neural map corresponding to the object in image scene A and the feature information B5 of the three-dimensional spatial area occupied in the neural map corresponding to the space in image scene B. The feature information B5 of the three-dimensional spatial area occupied in the neural map corresponding to the space in image scene B is replaced with the feature information A5, that is, the image scene replacement process is completed through the neural map.
[0084] It should be noted that the image scene replacement here may be a replacement for the same image scene, or a cross-scene replacement for different image scenes, and the manner of image scene replacement is not limited here.
[0085] 4) Map update application
[0086] As time changes, objects in the same scene will also change accordingly. For example, with the ever-accelerating construction of cities, the buildings in the cities will also continue to change. Therefore, the neural map also needs to be continuously updated as the scene changes. Figure 8 is a schematic diagram of the neural map update of an embodiment of the present invention. As shown in Figure 8, the feature information corresponding to the three-dimensional space area in the neural map before the update is fused with the feature information corresponding to the three-dimensional space in the neural map to be updated to obtain an updated neural map. The neural map before the update, the neural map to be updated, and the updated neural map respectively present different rendering effects when implementing image rendering.
[0087] It should be noted that the feature fusion process can refer to the feature fusion process of aerial photography maps and ground photography maps in the neural map during the aforementioned neural map construction process, which will not be repeated here.
[0088] FIG9 is a flow chart of a neural map construction method according to an embodiment of the present invention. As shown in FIG9 , the neural map construction method is applied to a cloud server. The neural map construction method includes:
[0089] S101: Acquire first acquisition device data acquired by a first acquisition device located under the cloud, and second acquisition device data acquired by a second acquisition device located under the cloud, wherein the first acquisition device data is an aerial image of a first shooting scene, and the second acquisition device data is a ground image of the first shooting scene.
[0090] S102: Aligning the first pose information of the aerial image and the second pose information of the ground image to the same geographic coordinate system.
[0091] S103: Extracting first feature information of each pixel of the aerial image and second feature information of each pixel of the ground image respectively.
[0092] S104: Constructing a neural map of the first shooting scene based on the geographic coordinate system aligned with the first pose information and the second pose information, the first feature information, and the second feature information.
[0093] The neural map includes multiple grids, each grid includes coordinate position information and feature information. The coordinate position information is the spatial position information of the pixel point mapping of the aerial image and / or the ground image of the first shooting scene, and the feature information is the fusion information of the first feature information of each pixel in the aerial image corresponding to the spatial position and the second feature information of each pixel in the ground image.
[0094] Steps S101 to S104 respectively collect aerial images and ground images of the same shooting scene through different acquisition devices, and then align the aerial images and the ground images to the same geographic coordinate system through pose calculation and alignment, extract the pose information and feature vectors of each pixel of the aerial images and the ground images, construct an initial neural map, cut the neural map into multiple three-dimensional spatial grids, correspond each pixel of the aerial images and / or the ground images to the three-dimensional spatial grid of the neural map, and fuse the feature vectors of each pixel of the aerial images and / or the ground images into the three-dimensional spatial grid of the neural map. At this time, each three-dimensional grid in the neural map contains the fused feature information of the aerial images and / or the ground images.
[0095] The above steps use different acquisition devices to capture images from different shooting angles of the same scene to construct a neural map through pose calculation, alignment, feature fusion and other processes, making the characteristics in the neural map more comprehensive and detailed, and can be used for subsequent positioning and rendering functions of images under multiple acquisition devices and multiple shooting angles.
[0096] S105: Based on the neural map and the aerial image, the third posture information of the aerial image is calculated using the positioning model to obtain first error data between the first posture information and the third posture information, and the first error data is fed back to the neural map and the positioning model; based on the neural map and the ground image, the fourth posture information of the ground image is calculated using the positioning model to obtain second error data between the second posture information and the fourth posture information, and the second error data is fed back to the neural map and the positioning model; based on the first error data and the second error data, the feature information in the neural map and the parameter information of the positioning model are updated.
[0097] S106: Based on the neural map and the first pose information, the rendering model is used to calculate the first rendered image of the aerial image, and the third error data between the first rendered image and the aerial image is obtained, and the third error data is fed back to the neural map and the rendering model; based on the neural map and the second pose information, the rendering model is used to calculate the second rendered image of the ground image, and the fourth error data between the second rendered image and the ground image is obtained, and the fourth error data is fed back to the neural map and the rendering model; based on the third error data and the fourth error data, the feature information in the neural map and the parameter information of the rendering model are updated.
[0098] Steps S105 and S106 train the constructed neural map through the positioning model and the rendering model respectively, and continuously update the neural map through error analysis based on the training results, so that the neural map is more accurate in the subsequent application stage.
[0099] Furthermore, the method further includes encoding stylized labels for the aerial and terrestrial images, wherein the stylized labels are used to indicate the image style types of the aerial and terrestrial images. The stylized label encoding facilitates users to input the stylized label encoding and customize the style type of image rendering when subsequently using the neural map.
[0100] The application phase of the trained map also includes the following steps:
[0101] An image of a second scene to be positioned is received, where the second scene is within the range of the first scene. Based on the neural map and the image to be positioned, fifth pose information of the image to be positioned is calculated using the positioning model. If a user inputs an image to be positioned, and the image to be positioned is within the range of the scene captured when constructing the neural map, the pose information of the image to be positioned can be obtained by combining the neural map with the positioning model.
[0102] A sixth pose information of the second captured scene is received, and a third rendered image corresponding to the sixth pose information is calculated using the rendering model based on the neural map and the sixth pose information. If a user inputs an image to be rendered, and the rendered image is within the range of the scene captured when constructing the neural map, a rendered image of the image to be rendered can be obtained by combining the neural map with the rendering model.
[0103] The system receives the seventh pose information and the first stylized label of the second captured scene. Based on the neural map, the seventh pose information, and the first stylized label, the rendering model calculates a fourth rendered image corresponding to the seventh pose information. The first stylized label indicates the image style type of the fourth rendered image. The user inputs the image to be rendered and a customized stylized label. The neural map is combined with the rendering model to generate a rendered image that matches the customized style.
[0104] Receive a partial area selection instruction for a first image of a third shooting scene, the third shooting scene being within the range of the first shooting scene, identify first spatial information corresponding to the partial area of the first image in the first image, and locate first grid information corresponding to the first spatial information in the neural map; receive a partial area migration instruction for a second image of the third shooting scene, identify second spatial information corresponding to the partial area of the second image in the second image, locate second grid information corresponding to the second spatial information in the neural map; replace the second grid information in the neural map with the first grid information, and update the neural map. The user selects partial areas in the two images and replaces the grid information corresponding to the partial areas in the neural map, thereby achieving the replacement operation of the partial areas in the two images.
[0105] Furthermore, the first, second, and third feature information include semantic information, point, line, and surface information, density information, lighting information, color information, texture information, and material information corresponding to the first captured scene. The different feature information allows for more comprehensive details when merging the aerial and terrestrial images.
[0106] The neural map construction method provided in this application meets the service requirements of high-precision rendering and positioning of the entire airspace, so as to solve the problem that the traditional map construction method is relatively single and cannot achieve high-precision rendering and positioning at the same time in large-scene modeling.
[0107] Based on the process of the neural map construction method provided above, an embodiment of the present invention further discloses a neural map construction device. Figure 10 is a schematic diagram of a neural map construction device according to an embodiment of the present invention. As shown in Figure 10, the neural map construction device is applied to a cloud server and includes:
[0108] The acquisition module 101 is used to acquire first acquisition device data collected by a first acquisition device located under the cloud, and second acquisition device data collected by a second acquisition device located under the cloud, wherein the first acquisition device data is an aerial image of the first shooting scene, and the second acquisition device data is a ground image of the first shooting scene.
[0109] The alignment module 102 is configured to align the first pose information of the aerial image and the second pose information of the ground image into the same geographic coordinate system.
[0110] The extraction module 103 is configured to extract first feature information of each pixel of the aerial image and second feature information of each pixel of the ground image.
[0111] A construction module 104 is used to construct a neural map of the first shooting scene based on a geographic coordinate system aligned with the first pose information and the second pose information, the first feature information, and the second feature information. The neural map includes multiple grids, each grid includes coordinate position information and feature information, the coordinate position information is the spatial position information of the pixel point mapping of the aerial image and / or the ground image of the first shooting scene, and the feature information is the fusion information of the first feature information of each pixel of the aerial image corresponding to the spatial position and the second feature information of each pixel of the ground image.
[0112] The error calculation module 105 is used to calculate the third pose information of the aerial image based on the neural map and the aerial image using the positioning model, obtain the first error data between the first pose information and the third pose information, and feed the first error data back to the neural map and the positioning model.
[0113] The error calculation module 105 is also used to calculate the fourth posture information of the ground-captured image based on the neural map and the ground-captured image using the positioning model, obtain the second error data between the second posture information and the fourth posture information, and feed the second error data back to the neural map and the positioning model.
[0114] The updating module 106 is configured to update the third feature information in the neural map and the parameter information of the positioning model based on the first error data and the second error data.
[0115] The error calculation module 105 is also used to calculate the first rendered image of the aerial photography posture using the rendering model based on the neural map and the first pose information, obtain third error data between the first rendered image and the aerial photography image, and feed the third error data back to the neural map and the rendering model.
[0116] The error calculation module 105 is also used to calculate the second rendered image of the ground shooting posture using the rendering model based on the neural map and the second posture information, obtain the fourth error data between the second rendered image and the ground shooting image, and feed the fourth error data back to the neural map and the rendering model.
[0117] The updating module 106 is further configured to update the third feature information in the neural map and the parameter information of the rendering model based on the third error data and the fourth error data.
[0118] The label encoding module 107 is used to perform stylized label encoding on the aerial image and the ground image, where the stylized label is used to indicate the image style type to which the aerial image and the ground image belong.
[0119] The positioning module 108 is used to receive the image to be positioned of the second shooting scene, where the second shooting scene is located within the range of the first shooting scene, and calculate the fifth pose information of the image to be positioned using the positioning model based on the neural map and the image to be positioned.
[0120] The rendering module 109 is configured to receive the sixth posture information of the second shooting scene, and calculate a third rendered image corresponding to the sixth posture information using a rendering model based on the neural map and the sixth posture information.
[0121] The rendering module 109 is further used to receive the seventh pose information and the first stylization label of the second shooting scene, and calculate the fourth rendered image corresponding to the seventh pose information using the rendering model based on the neural map, the seventh pose information and the first stylization label. The first stylization label is used to indicate the image style type to which the fourth rendered image belongs.
[0122] The updating module 106 is further configured to receive an instruction for selecting a partial region of the first image for a third shooting scene, the third shooting scene being within the range of the first shooting scene, identify first spatial information corresponding to the partial region of the first image in the first image, and locate first grid information corresponding to the first spatial information in the neural map;
[0123] The updating module 106 is further configured to receive a partial region migration instruction for the second image of the third shooting scene, identify second spatial information corresponding to the partial region of the second image in the second image, and locate second grid information corresponding to the second spatial information in the neural map;
[0124] The updating module 106 is further configured to replace the second grid information with the first grid information and update the neural map.
[0125] It should be noted that the acquisition module 101, alignment module 102, extraction module 103, construction module 104, error calculation module 105, update module 106, label encoding module 107, positioning module 108, and rendering module 109 can all be implemented through software or hardware. For example, the implementation of the cloud resource configuration interface acquisition module 101 will be described below using the cloud resource configuration interface acquisition module 101 as an example. Similarly, the implementation of the alignment module 102, extraction module 103, construction module 104, error calculation module 105, update module 106, label encoding module 107, positioning module 108, and rendering module 109 can refer to the implementation of the acquisition module 101.
[0126] When implemented by software, the acquisition module 101 can be an application or code block running on a computer device. The computer device can be at least one of a physical host, a virtual machine, a container, and other computing devices. Furthermore, the above-mentioned computer device can be one or more. For example, the acquisition module 101 can be an application running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application can be distributed in the same availability zone (AZ) or in different AZs. The multiple hosts / virtual machines / containers used to run the application can be distributed in the same region or in different regions. Generally, a region can include multiple AZs.
[0127] Similarly, multiple hosts / virtual machines / containers used to run the application can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Typically, a region can include multiple VPCs, and a VPC can include multiple AZs.
[0128] When implemented through hardware, the acquisition module 101 may include at least one computing device, such as a server. Alternatively, the cloud resource configuration interface providing module 101 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0129] The multiple computing devices included in acquisition module 101 can be distributed in the same AZ or in different AZs. The multiple computing devices included in cloud resource configuration interface provisioning module 101 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in acquisition module 101 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0130] It should be noted that the acquisition module 101, the alignment module 102, the extraction module 103, the construction module 104, the error calculation module 105, the update module 106, the label encoding module 107, the positioning module 108 and the rendering module 109 can all be used to execute some or all of the steps in the neural map construction method.
[0131] The various modules of the neural map construction device disclosed in the embodiment of the present invention have clear division of labor and close cooperation. The modules cooperate with each other to efficiently complete the construction of the neural network map.
[0132] The present invention also provides a computing device. See Figure 11 below, which is a schematic diagram of the computing device structure for the neural map construction method according to an embodiment of the present invention. Computing device 100 includes a bus 104, a processor 106, a memory 105, and a communication interface 107. Processor 106, memory 105, and communication interface 107 communicate with each other via bus 104. Computing device 100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 100.
[0133] Bus 104 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG11 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 104 may include a path for transmitting information between various components of computing device 100 (e.g., memory 105, processor 106, and communication interface 107).
[0134] The processor 106 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0135] The memory 105 may include a volatile memory, such as a random access memory (RAM). The processor 106 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0136] Memory 105 stores executable program code, which processor 106 executes to implement the functions of the aforementioned acquisition module 101, alignment module 102, extraction module 103, construction module 104, error calculation module 105, update module 106, label encoding module 107, positioning module 108, and rendering module 109, thereby implementing the method for configuring a cloud connection service based on a public cloud. Specifically, memory 105 stores instructions for the cloud management platform to execute the method for configuring a cloud connection service based on a public cloud.
[0137] The communication interface 107 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.
[0138] An embodiment of the present invention further provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0139] Please refer to Figure 12 below, which is a schematic diagram of the structure of a computing device cluster for the neural map construction method according to an embodiment of the present invention. The computing device cluster includes at least one computing device 100. The memory 105 of one or more computing devices 100 in the computing device cluster may store instructions for executing the method for configuring a cloud connection service based on a public cloud on the same cloud management platform.
[0140] In some possible implementations, one or more computing devices 100 in the computing device cluster may also be used to execute some of the instructions of the cloud management platform for executing the configuration method for a public cloud-based cloud connection service. In other words, the combination of one or more computing devices 100 may jointly execute the instructions of the cloud management platform for executing the configuration method for a public cloud-based cloud connection service.
[0141] It should be noted that the memory 105 in different computing devices 100 in the computing device cluster can store different instructions for executing some functions of the cloud management platform. That is, the instructions stored in the memory 105 in different computing devices 100 can implement the functions of one or more of the acquisition module 101, alignment module 102, extraction module 103, construction module 104, error calculation module 105, update module 106, label encoding module 107, positioning module 108, and rendering module 109.
[0142] An embodiment of the present invention further provides a computer program product containing instructions. This computer program product may be software or a program product containing instructions that can be executed on a computing device or stored in any available medium. When executed on at least one computing device, this computer program product causes the at least one computing device to execute the aforementioned configuration method for a cloud management platform for executing a public cloud-based cloud connection service.
[0143] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the configuration method described above for a cloud management platform for executing a public cloud-based cloud connection service.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for constructing a neural map, characterized in that: The method is applied to a cloud server, and the method comprises: Acquire first acquisition device data acquired by a first acquisition device located under the cloud, and second acquisition device data acquired by a second acquisition device located under the cloud, wherein the first acquisition device data is an aerial image of a first shooting scene, and the second acquisition device data is a ground image of the first shooting scene; Aligning the first pose information of the aerial image and the second pose information of the ground image to the same geographic coordinate system; Respectively extracting first feature information of each pixel of the aerial image and second feature information of each pixel of the ground image; A neural map of the first shooting scene is constructed based on the geographic coordinate system aligned with the first posture information and the second posture information, the first feature information and the second feature information, the neural map including a plurality of grids, each grid including coordinate position information and third feature information, the coordinate position information being spatial position information of pixel point mapping of the aerial image and / or the ground image of the first shooting scene, and the third feature information being fusion information of the first feature information of each pixel of the aerial image corresponding to the spatial position and the second feature information of each pixel of the ground image.
2. The method according to claim 1, characterized in that The method further comprises: Calculating third pose information of the aerial image using a positioning model based on the neural map and the aerial image, obtaining first error data between the first pose information and the third pose information, and feeding back the first error data to the neural map and the positioning model; Calculating fourth pose information of the ground-captured image using a positioning model based on the neural map and the ground-captured image, obtaining second error data between the second pose information and the fourth pose information, and feeding back the second error data to the neural map and the positioning model; Based on the first error data and the second error data, the third feature information in the neural map and the parameter information of the positioning model are updated.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: Calculating a first rendered image of the aerial image using a rendering model based on the neural map and the first pose information, obtaining third error data between the first rendered image and the aerial image, and feeding the third error data back to the neural map and the rendering model; Calculate a second rendered image of the ground-captured image using a rendering model based on the neural map and the second pose information, obtain fourth error data between the second rendered image and the ground-captured image, and feed the fourth error data back to the neural map and the rendering model; Based on the third error data and the fourth error data, the third feature information in the neural map and the parameter information of the rendering model are updated.
4. The method according to any one of claims 1 or 3, characterized in that: The method further comprises: Stylized label encoding is performed on the aerial image and the ground image, where the stylized label is used to indicate the image style type to which the aerial image and the ground image belong.
5. The method according to any one of claims 2 or 4, characterized in that: The method further comprises: An image to be positioned of a second shooting scene is received, where the second shooting scene is within the range of the first shooting scene, and fifth posture information of the image to be positioned is calculated using the positioning model based on the neural map and the image to be positioned.
6. The method according to claim 5, characterized in that The method further comprises: Receive sixth posture information of the second shooting scene, and calculate a third rendered image corresponding to the sixth posture information using the rendering model based on the neural map and the sixth posture information.
7. The method according to claim 5 or 6, characterized in that: The method further comprises: Receive the seventh pose information and the first stylized label of the second shooting scene, and calculate the fourth rendered image corresponding to the seventh pose information using the rendering model based on the neural map, the seventh pose information and the first stylized label, wherein the first stylized label is used to indicate the image style type to which the fourth rendered image belongs.
8. The method according to any one of claims 2 or 7, characterized in that: The method further comprises: receiving a partial area selection instruction for a first image of a third shooting scene, the third shooting scene being located within the range of the first shooting scene, identifying first spatial information corresponding to a partial area of the first image in the first image, and locating first grid information corresponding to the first spatial information in the neural map; receiving a partial area migration instruction for a second image of the third shooting scene, identifying second spatial information corresponding to the partial area of the second image in the second image, and locating second grid information corresponding to the second spatial information in the neural map; The second grid information in the neural map is replaced with the first grid information, and the neural map is updated.
9. The method according to any one of claims 1 or 8, characterized in that: The first feature information, the second feature information and the third feature information include semantic information, point, line and surface information, density information, lighting information, color information, texture information and material information corresponding to the first shooting scene.
10. A neural map construction device, characterized in that: The device is applied to a cloud server, and the device includes: an acquisition module, configured to acquire first acquisition device data acquired by a first acquisition device located under the cloud, and second acquisition device data acquired by a second acquisition device located under the cloud, wherein the first acquisition device data is an aerial image of a first shooting scene, and the second acquisition device data is a ground image of the first shooting scene; An alignment module, used for aligning the first pose information of the aerial image and the second pose information of the ground image into the same geographic coordinate system; An extraction module, used to extract first feature information of each pixel of the aerial image and second feature information of each pixel of the ground image respectively; A construction module is used to construct a neural map of the first shooting scene based on the geographic coordinate system aligned with the first posture information and the second posture information, the first feature information and the second feature information, the neural map including multiple grids, each grid including coordinate position information and third feature information, the coordinate position information is the spatial position information of the pixel point mapping of the aerial image and / or the ground image of the first shooting scene, and the third feature information is the fusion information of the first feature information of each pixel of the aerial image corresponding to the spatial position and the second feature information of each pixel of the ground image.
11. The device according to claim 10, characterized in that The device also includes: an error calculation module, configured to calculate third posture information of the aerial image using a positioning model based on the neural map and the aerial image, obtain first error data between the first posture information and the third posture information, and feed the first error data back to the neural map and the positioning model; The error calculation module is further used to calculate the fourth posture information of the ground-captured image using the positioning model based on the neural map and the ground-captured image, obtain the second error data between the second posture information and the fourth posture information, and feed the second error data back to the neural map and the positioning model; An updating module is used to update the third feature information in the neural map and the parameter information of the positioning model based on the first error data and the second error data.
12. The device according to claim 10 or 11, characterized in that The error calculation module is further used to calculate a first rendered image of the aerial image using a rendering model based on the neural map and the first pose information, obtain third error data between the first rendered image and the aerial image, and feed the third error data back to the neural map and the rendering model; The error calculation module is further used to calculate a second rendered image of the ground-captured image using a rendering model based on the neural map and the second pose information, obtain fourth error data between the second rendered image and the ground-captured image, and feed the fourth error data back to the neural map and the rendering model; The updating module is further used to update the third feature information in the neural map and the parameter information of the rendering model based on the third error data and the fourth error data.
13. The device according to any one of claims 10 or 12, characterized in that The device also includes: The label encoding module is used to perform stylized label encoding on the aerial image and the ground image, wherein the stylized label is used to indicate the image style type to which the aerial image and the ground image belong.
14. The device according to any one of claims 11 or 13, characterized in that The device also includes: A positioning module is used to receive an image to be positioned of a second shooting scene, where the second shooting scene is located within the range of the first shooting scene, and calculate fifth posture information of the image to be positioned using the positioning model based on the neural map and the image to be positioned.
15. The device according to claim 14, characterized in that The device also includes: A rendering module is used to receive the sixth posture information of the second shooting scene, and calculate a third rendered image corresponding to the sixth posture information using the rendering model based on the neural map and the sixth posture information.
16. The device according to claim 14 or 15, characterized in that The rendering module is further configured to receive seventh pose information and a first stylized label of the second shooting scene, and based on the neural map Figure, the seventh pose information and the first stylized label, use the rendering model to calculate the fourth rendered image corresponding to the seventh pose information, and the first stylized label is used to indicate the image style type to which the fourth rendered image belongs.
17. The device according to claim 14 or 16, characterized in that The updating module is further configured to receive a partial area selection instruction for a first image of a third shooting scene, the third shooting scene being located within the range of the first shooting scene, identifying first spatial information corresponding to a partial area of the first image in the first image, and locating first grid information corresponding to the first spatial information in the neural map; The updating module is further used to receive a partial area migration instruction for the second image of the third shooting scene, identify second spatial information corresponding to the partial area of the second image in the second image, and locate second grid information corresponding to the second spatial information in the neural map; The updating module is further used to replace the second grid information with the first grid information and update the neural map.
18. The method according to any one of claims 10 or 17, characterized in that: The first feature information, the second feature information and the third feature information include semantic information, point, line and surface information, density information, lighting information, color information, texture information and material information corresponding to the first shooting scene.
19. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 9.
20. A computer program product comprising instructions, characterized in that When the instructions are executed by a computer device cluster, the computer device cluster executes the method according to any one of claims 1 to 9.
21. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Positioning and mapping method and electronic equipment
CN111340922A
Map generation method and device, electronic equipment and storage medium
CN116450761A
Dicing tape with adhesive film
KR1020200108785A
Using a neural network scene representation for mapping
WO2023094271A1
Three-dimensional scene rendering method, device, and storage medium
WO2023138471A1