Neural map construction method and device

By constructing neural maps, using the position information and feature information of aerial images and ground-based images, the problem that traditional map construction methods cannot achieve high-precision rendering and positioning is solved, and high-precision rendering and positioning services are realized in the entire airspace.

CN120047553APending Publication Date: 2025-05-27HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311744565.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2023-12-18
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The traditional map construction method is relatively single, and it is impossible to achieve high-precision rendering and positioning functions simultaneously in large-scene modeling.

Method used

By obtaining data of aerial images and ground-based images, using position information to align, extract feature information of each pixel, and construct a neural map, including multiple grids, each grid includes coordinate position information and feature information, achieving high-precision rendering and positioning.

Benefits of technology

It realizes high-precision rendering and positioning in the entire airspace, meeting the service needs of high-precision rendering and positioning in large-scene modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047553A_ABST
    Figure CN120047553A_ABST
Patent Text Reader

Abstract

The invention provides a construction method of a neural map, and the method comprises the steps: collecting an aerial shot image and a ground shot image through a collection device, and enabling the first pose information of the aerial shot image and the second pose information of the ground shot image to be aligned to a same geographic coordinate system; and respectively extracting first feature information of each pixel of the aerial shot image and second feature information of each pixel of the ground shot image, and constructing a neural map based on the geographic coordinate system, the first feature information and the second feature information. The neural map comprises a plurality of grids, each grid comprises coordinate position information and third feature information, and the coordinate position information is spatial position information mapped by pixel points of an aerial shot image and / or a ground shot image; the third feature information is fusion information of the first feature information of each pixel of the aerial shot image corresponding to the spatial position and the second feature information of each pixel of the ground shot image. The neural map provided by the invention meets the service requirements of high-precision rendering and positioning of the whole airspace.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud technology, and in particular, to a method and apparatus for constructing a neural map. Background Art

[0002] City digital twin emphasizes establishing a virtual city that interacts with the physical city in real time, accurately mapping the operation of the physical city, forming a pattern of virtual-real interaction, so as to improve and optimize the comprehensive governance and planning level of the city. City digital twin is not a single technology, but a modular assembly of many cutting-edge technologies. Among them, mapping technology plays an important role and is deeply involved in the construction of virtual models and the realization of virtual-real interaction. Common map forms include urban 3D maps, AR maps, and autonomous driving maps, etc. The traditional map construction method is relatively single and cannot simultaneously achieve the functions of high-precision rendering and positioning in large-scale scene modeling. Summary of the Invention

[0003] This application provides a method and apparatus for constructing a neural map, which meet the service requirements of high-precision rendering and positioning in the whole airspace, so as to solve the problem that the traditional map construction method is relatively single and cannot simultaneously achieve high-precision rendering and positioning in large-scale scene modeling.

[0004] In a first aspect, this application provides a method for constructing a neural map. The method is applied to a cloud server and includes: obtaining first acquisition device data acquired by a first acquisition device located under the cloud and second acquisition device data acquired by a second acquisition device located under the cloud, where the first acquisition device data is an aerial image of a first shooting scene, and the second acquisition device data is a ground image of the first shooting scene; aligning the first pose information of the aerial image and the second pose information of the ground image to the same geographical coordinate system; respectively extracting first feature information of each pixel of the aerial image and second feature information of each pixel of the ground image; constructing a neural map of the first shooting scene based on the geographical coordinate system aligned with the first pose information and the second pose information, the first feature information, and the second feature information. The neural map includes a plurality of grids, and each grid includes coordinate position information and third feature information. The coordinate position information is the spatial position information mapped by pixel points of the aerial image and / or the ground image of the first shooting scene, and the third feature information is the fusion information of the first feature information of each pixel of the aerial image corresponding to the spatial position and the second feature information of each pixel of the ground image.

[0005] Based on the above method, aerial captured images and ground captured images are respectively captured for the same shooting scene using different capture devices, and then a neural map is constructed using the pose information and pixel feature information of the aerial captured images and the ground captured images. The neural map contains richer and more accurate information of the aerial captured images and / or the ground captured images, thereby meeting the service requirements for high-precision rendering and positioning in the entire airspace.

[0006] In a possible implementation manner of the first aspect, the method further includes: calculating the third pose information of the aerial captured image using a positioning model based on the neural map and the aerial captured image, obtaining first error data between the first pose information and the third pose information, and feeding back the first error data to the neural map and the positioning model; calculating the fourth pose information of the ground captured image using the positioning model based on the neural map and the ground captured image, obtaining second error data between the second pose information and the fourth pose information, and feeding back the second error data to the neural map and the positioning model; updating the feature information in the neural map and the parameter information of the positioning model based on the first error data and the second error data.

[0007] Based on the above method, the parameter information in the neural map is updated using the aerial captured image and / or the ground captured image through the positioning model, thereby realizing continuous training of the neural map in the positioning mode and making the data in the neural map and the positioning model more accurate.

[0008] In a possible implementation manner of the first aspect, the method further includes: calculating the first rendered image of the aerial captured image using a rendering model based on the neural map and the first pose information, obtaining third error data between the first rendered image and the aerial captured image, and feeding back the third error data to the neural map and the rendering model; calculating the second rendered image of the ground captured image using the rendering model based on the neural map and the second pose information, obtaining fourth error data between the second rendered image and the ground captured image, and feeding back the fourth error data to the neural map and the rendering model; updating the feature information in the neural map and the parameter information of the rendering model based on the third error data and the fourth error data.

[0009] Based on the above method, the parameter information in the neural map is updated using the aerial captured image and / or the ground captured image through the rendering model, thereby realizing continuous training of the neural map in the rendering mode and making the data in the neural map and the rendering model more accurate.

[0010] In a possible implementation manner of the first aspect, the method further includes: performing stylized label encoding on the aerial captured image and the ground captured image, and the stylized label is used to indicate the image style type to which the aerial captured image and the ground captured image belong.

[0011] Based on the above method, by performing stylized label encoding on the image style types to which the aerial photography images and the ground photography images belong, the information of the aerial photography images and the ground photography images stored in the neural map can be made more comprehensive and the classification can be more accurate, which is beneficial for the subsequent rendering model to perform stylized rendering on the images.

[0012] In a possible implementation manner of the first aspect, the method further includes: receiving a to-be-localized image of a second shooting scene, where the second shooting scene is within the range of the first shooting scene, and based on the neural map and the to-be-localized image, using a localization model to calculate the fifth pose information of the to-be-localized image.

[0013] Based on the above method, using the trained neural map and the localization model can complete the localization of the images within the same shooting scene range. Since the neural map is pre-constructed based on the aerial photography images and the ground photography images of the same shooting scene and undergoes localization training, the data obtained in this localization process is more accurate.

[0014] In a possible implementation manner of the first aspect, the method further includes: receiving the sixth pose information of the second shooting scene, and based on the neural map and the sixth pose information, using a rendering model to calculate the third rendered image corresponding to the sixth pose information.

[0015] Based on the above method, using the trained neural map and the rendering model can complete the rendering of the images within the same shooting scene range. Since the neural map is pre-constructed based on the aerial photography images and the ground photography images of the same shooting scene and undergoes rendering training, the data obtained in this rendering process is more accurate.

[0016] In a possible implementation manner of the first aspect, receive the seventh pose information of the second shooting scene and the first stylized label, and based on the neural map, the seventh pose information, and the first stylized label, use a rendering model to calculate the fourth rendered image corresponding to the seventh pose information, where the first stylized label is used to indicate the image style type to which the fourth rendered image belongs.

[0017] Based on the above method, by defining the stylized label information, the custom of the image rendering style type can be performed, making the image style types of image rendering more diverse.

[0018] In a possible implementation of the first aspect, a partial area selection instruction for a first image of a third shooting scene is received. The third shooting scene is within the range of the first shooting scene. The first spatial information corresponding to the partial area of the first image in the first image is identified, and the first grid information corresponding to the first spatial information in the neural map is located. A partial area migration instruction for a second image of the third shooting scene is received. The second spatial information corresponding to the partial area of the second image in the second image is identified, and the second grid information corresponding to the second spatial information in the neural map is located. The second grid information in the neural map is replaced with the first grid information, and the neural map is updated.

[0019] Based on the above method, by replacing different grid information in the neural map, the synchronous replacement of the image background can be achieved more quickly and accurately.

[0020] In a possible implementation of the first aspect, the first feature information, the second feature information, and the third feature information include semantic information, point-line-plane information, density information, illumination information, color information, texture information, and material information corresponding to the first shooting scene.

[0021] Based on the above method, by extracting rich feature information of each pixel of the aerial shooting image and the ground shooting image, the types of fusion information in the neural map are more diversified, the neural map is more accurate in image positioning, and the image rendering is more precise.

[0022] In a second aspect, the present application further provides a neural map construction device, which is applied to a cloud server. The device includes: an acquisition module, configured to acquire first acquisition device data acquired by a first acquisition device located under the cloud and second acquisition device data acquired by a second acquisition device located under the cloud, where the first acquisition device data is an aerial shooting image of a first shooting scene, and the second acquisition device data is a ground shooting image of the first shooting scene; an alignment module, configured to align the first pose information of the aerial shooting image and the second pose information of the ground shooting image to the same geographical coordinate system; an extraction module, configured to extract first feature information of each pixel of the aerial shooting image and second feature information of each pixel of the ground shooting image respectively; a construction module, configured to construct a neural map of the first shooting scene based on the geographical coordinate system aligned with the first pose information and the second pose information, the first feature information, and the second feature information. The neural map includes a plurality of grids, each grid includes coordinate position information and feature information, the coordinate position information is the spatial position information mapped by pixel points of the aerial shooting image and / or the ground shooting image of the first shooting scene, and the feature information is the fusion information of the first feature information of each pixel of the aerial shooting image and the second feature information of each pixel of the ground shooting image corresponding to the spatial position.

[0023] The second aspect or any implementation manner of the second aspect is implemented by the steps of the device corresponding to the first aspect or any implementation manner of the first aspect. The descriptions in the second aspect or any implementation manner of the second aspect are applicable to the first aspect or any implementation manner of the first aspect, and will not be repeated here.

[0024] In a third aspect, the present application provides a computing device cluster, including at least one computing device. Each computing device includes a processor and a memory. The processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method disclosed in the first aspect and any possible implementation manner of the first aspect.

[0025] In a fourth aspect, the present application provides a computer program product containing instructions. When the instructions are run by a computing device cluster, the computing device cluster is enabled to implement the method disclosed in the first aspect and any possible implementation manner of the first aspect.

[0026] In a fifth aspect, the present application provides a computer-readable storage medium, including computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster is enabled to execute the method disclosed in the first aspect and any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic flowchart of a method for constructing a neural map according to an embodiment of the present invention;

[0028] Figure 2 is a schematic diagram of an image collected by a collection device in an embodiment of the present invention;

[0029] Figure 3a is a schematic diagram of a point cloud of an aerial photography image in an embodiment of the present invention;

[0030] Figure 3b is a schematic diagram of a point cloud of a ground photography image in an embodiment of the present invention;

[0031] Figure 4 is a schematic flowchart of training a positioning model in a method for constructing a neural map according to an embodiment of the present invention;

[0032] Figure 5 is a schematic flowchart of training a rendering model in a method for constructing a neural map according to an embodiment of the present invention;

[0033] Figure 6 is a schematic diagram of style transfer when implementing a rendering application based on a neural map according to an embodiment of the present invention;

[0034] Figure 7 is a schematic diagram of implementing a map replacement application based on a neural map according to an embodiment of the present invention;

[0035] Figure 8 It is a schematic diagram of the neural map update in the embodiment of the present invention;

[0036] Figure 9 It is a flowchart of a method for constructing a neural map in the embodiment of the present invention;

[0037] Figure 10 It is a schematic diagram of a device for constructing a neural map in the embodiment of the present invention;

[0038] Figure 11 It is a schematic diagram of the structure of a computing device for the method of constructing a neural map in the embodiment of the present invention;

[0039] Figure 12 It is a schematic diagram of the structure of a computing device cluster for the method of constructing a neural map in the embodiment of the present invention. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] To facilitate the understanding of the embodiments of the present application, first, some terms related to the present invention will be explained.

[0042] Neural map: A neural map is a map constructed by collecting various data information through sensors, extracting and fusing the characteristics of these data information according to needs, and then binding them to spatial coordinates.

[0043] Pose: Pose refers to the position and orientation of the posture or pose of an object, robot, or person in three-dimensional space. It consists of two elements: position and orientation. The position represents the coordinates of the center or reference point of the object in three-dimensional space, usually represented by three real numbers. The orientation represents the orientation or direction of the object in three-dimensional space, usually represented by a rotation matrix, Euler angles, or quaternions, etc.

[0044] Feature fusion: Feature fusion is a process of extracting features of different attributes from sensor data, using the complementarity between features, and fusing the advantages between features to construct new fused features, thereby improving the performance of the model.

[0045] Figure 1 It is a schematic flowchart of a method for constructing a neural map in the embodiment of the present invention. As Figure 1 shown, the embodiment of the present invention provides a method for constructing a neural map. This method is applied to a cloud server and includes:

[0046] 1. Data collection stage:

[0047] Acquire acquisition device data A1 of device A located under the cloud, and acquisition device data B1 of acquisition device B located under the cloud, wherein acquisition device data A1 is an aerial image of shooting scene M, and acquisition device data B1 is a ground image of shooting scene M.

[0048] Specifically, the acquisition device A and the acquisition device B can be sensor devices, such as: satellite collectors, drones, panoramic cameras, laser scanners, cameras, video cameras, infrared cameras, etc.

[0049] Figure 2 is a schematic diagram of an image collected by a collection device in an embodiment of the present invention, such as Figure 2 As shown, the acquisition device data A1 and the acquisition device data B1 collected by the acquisition device A and the acquisition device B respectively can be multi-perspective satellite images collected by a satellite collector, drone bird's-eye views collected by a drone, image sequences collected by a panoramic camera, laser point clouds collected by a laser scanner, etc.

[0050] The acquisition device A and the acquisition device B can be the same type of acquisition devices, or they can be different types of acquisition devices that have advantages in shooting in the aerial range and in the ground range respectively for the shooting scene M. For example, the acquisition device A is a drone that has advantages in shooting within the aerial range of the shooting scene M, and the acquisition device B is a panoramic camera that has advantages in shooting within the ground range of the shooting scene M.

[0051] It should be noted that the description of the device types and the range of pictures taken of the above-mentioned acquisition devices A and B are all cases in the embodiments of the present invention. In other embodiments, the device types of the acquisition devices and the range of pictures taken may also exist in other cases similar to the central idea of ​​the present invention. The types of the acquisition devices and the range of pictures taken are not limited here.

[0052] 2. Map construction phase:

[0053] 1) Align the position information A2 of the aerial image and the position information B2 of the ground image to the same geographic coordinate system.

[0054] Specifically, the aerial captured image is used to obtain the point cloud A4 and pose information A2 in the coordinate system of the aerial captured image through sparse reconstruction and / or dense reconstruction. The ground captured image is used to obtain the point cloud B4 and pose information B2 in the coordinate system of the ground captured image through sparse reconstruction and / or dense reconstruction. Taking the coordinate system of the ground captured image as the reference coordinate system, the point cloud A4 of the aerial captured image is registered with the point cloud B4 of the ground captured image, so as to obtain the transformation relationship from the coordinate system of the aerial captured image to the coordinate system of the ground captured image, and then the pose information A2 of the aerial captured image is aligned into the coordinate system of the ground captured image. Figure 3a is a schematic diagram of the point cloud of the aerial captured image in the embodiment of the present invention, as Figure 3a shown, the point cloud A4 of the aerial captured image contains the pose information A2 of the aerial captured image. Figure 3b is a schematic diagram of the point cloud of the ground captured image in the embodiment of the present invention, as Figure 3b shown, the point cloud B4 of the ground captured image contains the pose information B2 of the ground captured image. Among them, the pose information A2 can be understood as the shooting trajectory formed when the acquisition device A shoots the shooting scene M within the aerial range, and the pose information B2 can be understood as the shooting trajectory formed when the acquisition device B shoots the shooting scene M within the ground range.

[0055] Further, in the coordinate system of the ground captured image, feature matching of the ground captured image and the aerial captured image can be performed to further optimize the pose information A2 of the aerial captured image, so that the result of aligning the pose information A2 of the aerial captured image into the coordinate system of the ground captured image is more accurate.

[0056] It should be noted that during the alignment of the geographical coordinate systems, it can be to align the coordinate system of the aerial captured image to the coordinate system of the ground captured image, or to align the coordinate system of the ground captured image to the coordinate system of the aerial captured image, or to align the aerial captured image and the ground captured image to other geographical coordinate systems. The alignment method of the geographical coordinate systems is not limited here.

[0057] 2) Extract the feature information A3 of each pixel of the aerial captured image and the feature information B3 of each pixel of the ground captured image respectively.

[0058] Specifically, the aerial captured image is input into a deep convolutional neural network to obtain the feature vectors of each pixel of the aerial captured image, and the ground captured image is input into a deep convolutional neural network to obtain the feature vectors of each pixel of the ground captured image. For example: the pixel height of the aerial captured image or the ground captured image is 1000, the pixel width is 2000, and the length of the feature vector is 256, then the dimension of the corresponding feature information of the aerial captured image or the ground captured image finally obtained is: 1000x2000x256.

[0059] 3) Based on the geographical coordinate system with pose information A2 and pose information B2 aligned, and feature information A3 and feature information B3, construct a neural map of the shooting scene M. The neural map includes multiple grids, and each grid includes coordinate position information C1 and feature information C2. The coordinate position information C1 is the spatial position information mapped by pixel points of the aerial photography image and / or ground photography image of the shooting scene M, and the feature information C2 is the fusion information of the feature information A3 of each pixel of the aerial photography image corresponding to the spatial position and the feature information B3 of each pixel of the ground photography image.

[0060] Specifically, based on the coordinate system of the ground photography image with pose information A2 and pose information B2 aligned, and feature information A3 and feature information B3, construct a neural map of the shooting scene M. This neural map is an untrained initial neural map, and this neural map represents the feature vectors stored in the three-dimensional space. For example: in a three-dimensional space with a length L = 100m, a width W = 50m, and a height H = 10m, divide a three-dimensional grid. The resolution size of the grid is d = 0.1m. Then the grid dimension corresponding to this three-dimensional space is M = L / d = 1000, N = W / d = 500, O = H / d = 100, that is, MxNxO = 1000x500x100. The length of the feature vector stored in each grid is D = 256. Then the dimension of the feature vector stored in this neural map is MxNxOxD = 1000x500x100x256.

[0061] Furthermore, project the coordinates of each grid in this three-dimensional space (that is, the coordinate position information C1, usually represented by the grid center coordinates) onto the aerial photography image and / or ground photography image to obtain the feature vectors (that is, the feature information C2) of the corresponding pixel coordinates. The pixel coordinates are the intersection coordinates of the line connecting the grid coordinates and the camera optical center coordinates and the camera imaging plane. Since each grid covered by the image photography range may obtain one or more feature vectors from one or more aerial photography images and / or ground photography images, use networks such as multi-layer perceptrons to aggregate multiple feature vectors into one feature vector, so as to realize the feature fusion of the feature information A3 of each pixel of the aerial photography image and the feature information B3 of each pixel of the ground photography image. Finally, each grid only contains one fused feature vector (that is, the feature information C2).

[0062] 3. Map training stage:

[0063] The constructed neural map also needs a corresponding training process to be used. The following respectively gives examples of the steps for training the positioning model and rendering model of the neural map. In other possible embodiments, the neural map can also be trained corresponding to other models as needed. The model training method of the neural map is not limited here.

[0064] 1) Localization model training

[0065] Based on the neural map and aerial images, the pose information D1 of the aerial images is calculated using the localization model, the error data E1 between the pose information A2 and the pose information D1 is obtained, and the error data E1 is fed back to the neural map and the localization model; based on the neural map and ground images, the pose information D2 of the ground images is calculated using the localization model, the error data E2 between the pose information B2 and the pose information D2 is obtained, and the error data E2 is fed back to the neural map and the localization model; based on the error data E1 and the error data E2, the feature information C2 in the neural map and the parameter information of the localization model are updated.

[0066] Specifically, after the neural map is constructed through a set of a large number of aerial images and / or ground images, any image in the set of images needs to be used as the input image for the localization model training to perform training operations on the neural map. Figure 4 It is a schematic flowchart of the localization model training in a method for constructing a neural map according to an embodiment of the present invention. As Figure 4 shown, for example: we input a ground image, and the neural map determines the initial retrieval range of the ground image in the neural map based on the GPS positioning information carried by the ground image itself, obtains the grid area L of the corresponding three-dimensional space in the retrieval range, extracts the coordinate position information and feature information included in the grid area L, and the feature information is decoded by a localization decoder to obtain a feature vector for localization. The feature vector includes semantic information, point-line-plane information, and structured information of the ground image, etc. Through the above information, the grid coordinates and localization feature information of the corresponding three-dimensional space in the retrieval range, that is, the localization feature map W1, are finally obtained. The localization model uses a feature extraction network to extract the three-dimensional frustum feature map W2 of the ground image. By matching the three-dimensional frustum feature map W2 and the localization feature map W1 corresponding to the retrieval range, the pose information D1 corresponding to the ground image is calculated. At the same time, we calculate the error data E1 between the pose information D1 and the true pose information A2 of the ground image. The error data E1 is a pose loss function. After the pose loss function is backpropagated by the gradient, the feature information C2 in the neural map and the parameter information of the localization model are continuously updated.

[0067] It should be noted that the process of training the neural map using aerial images through the localization model is similar to the process of training the neural map by inputting the ground images as described above, and will not be elaborated here. At the same time, Figure 4The calculation formula of the pose loss function shown is only the calculation formula of one loss function in the embodiments of the present invention. In other embodiments, the calculation formula of the pose loss function can also be other formulas, and the calculation formula of the pose loss function is not limited this time.

[0068] 2) Rendering model training

[0069] Based on the neural map and pose information A2, use the rendering model to calculate the rendered image F1 of the aerial captured image, obtain the error data G1 between the rendered image F1 and the aerial captured image, and feedback the error data G1 to the neural map and the rendering model; based on the neural map and pose information B2, use the rendering model to calculate the rendered image F2 of the ground captured image, obtain the error data G2 between the rendered image F2 and the ground captured image, and feedback the error data G2 to the neural map and the rendering model; based on the error data G1 and the error data G2, update the feature information C2 in the neural map and the parameter information of the rendering model.

[0070] Figure 5 is a schematic flowchart of the rendering model training in a neural map construction method according to an embodiment of the present invention. As Figure 5 shown, input the camera pose (position and angle) and camera parameters (focal length, image resolution, etc.) of the acquisition device B for capturing the aerial captured image, determine the three-dimensional space area covered by the shooting perspective range of this camera through the camera pose and camera parameters, extract the grid area X corresponding to this three-dimensional space area in the neural map, extract the coordinate position information and feature information included in this grid area X, the feature information is decoded by the rendering decoder to obtain the feature vector for rendering, and the feature vector includes the density information, color information, and signed distance field information of the ground captured image. Through the above information, finally obtain the grid coordinates and rendering feature information of the three-dimensional space corresponding to the shooting perspective range of this camera, that is, the rendering feature map Z. Using the rendering model and the neural rendering method, calculate the rendered image F1 corresponding to this aerial captured image through the rendering feature map Z1. At the same time, we need to calculate the error data G2 between this rendered image F1 and the real image Z2 of the aerial captured image. This error data G2 is the pixel color loss function. After the pixel color loss function is backpropagated by the gradient, continuously update the feature information C2 in the neural map and the parameter information of the rendering model.

[0071] It should be noted that the process of training the neural map using the ground captured image through the rendering model is similar to the process of training the neural map using the acquisition device B for the aerial captured image input above, and will not be elaborated here.

[0072] Furthermore, in order to diversify the styles of the rendered images, the aerial and ground captured images are collected in a diverse manner for different weather conditions, lighting, and seasons. After the rendering training of the aerial and ground captured images, stylized label encodings for the aerial and ground captured images can be obtained. The stylized label encodings are used to indicate the image style types to which the aerial and ground captured images belong, facilitating subsequent use by the user of the neural map. The user can input the stylized label encodings to customize the style type of the image rendering.

[0073] 4. Map application stage:

[0074] The trained neural map can be released for users to use. The following are examples of the steps for the image localization application, image rendering application, image fusion application, and image replacement application of the neural map respectively. In other possible embodiments, the neural map can also be used for other applications as needed. The application methods of the neural map are not limited here.

[0075] 1) Image localization application

[0076] During image localization, the user inputs an image to be localized. The image to be localized can be an aerial captured image or a ground captured image. The trained neural map determines the initial retrieval range of the image to be localized in the neural map based on the GPS positioning information carried by the image itself, obtains the grid area of the three-dimensional space corresponding to the retrieval range, extracts the coordinate position information and feature information contained in the grid area. The feature information is decoded by a localization decoder to obtain a feature vector for localization. The feature vector includes semantic information, point-line-plane information, and structured information of the image to be localized, etc. Through the above information, the grid coordinates and localization feature information of the three-dimensional space corresponding to the image to be localized, that is, the localization feature map, are finally obtained. The localization model uses the three-dimensional visual cone feature map calculation method to calculate the pose information corresponding to the image to be localized through the localization feature map, thus completing the localization process of the image to be localized.

[0077] 2) Image rendering application

[0078] During image rendering, the user inputs the camera pose (pose and angle) and camera parameters (focal length, image resolution, etc.) of the acquisition device. Based on the camera pose and camera parameters, the three-dimensional space area covered by the shooting perspective range of the camera is determined. The grid area corresponding to this three-dimensional space area in the trained neural map is extracted, and the coordinate position information and feature information contained in this grid area are extracted. The feature information is decoded by a rendering decoder to obtain a feature vector for rendering. The feature vector includes the density information, color information, and signed distance field information of the image to be rendered. Through the above information, the grid coordinates and rendering feature information of the three-dimensional space corresponding to the shooting perspective range of the camera are finally obtained, that is, the rendering feature map. Using the rendering model and neural rendering method, the rendering image corresponding to the aerial photography image is calculated through the rendering feature map, and thus the rendering process of the image to be rendered is completed.

[0079] Furthermore, during image rendering, the user can also input a stylization label as needed. When the trained neural map is rendering, it will perform style transfer on the image to be rendered according to the stylization label input by the user, so as to obtain the rendering image required by the user. Figure 6 It is a schematic diagram of style transfer when the embodiment of the present invention realizes a rendering application based on a neural map, as Figure 6 shown. Respectively, it shows the images with different rendering effects when the input stylization labels are sunny, night, and winter.

[0080] It should be noted that the stylization label here can be a text label such as sunny, night, winter, or a digital label, a letter label, etc. The labeling method of the label is not limited here.

[0081] 3) Image replacement application

[0082] Figure 7 It is a schematic diagram of the image replacement application realized by the embodiment of the present invention based on a neural map, as Figure 7 shown. The user can click on an object in image scene A and click on a space in image scene B. The feature information A5 of the three-dimensional space area occupied by the object in image scene A in the corresponding neural map and the feature information B5 of the three-dimensional space area occupied by the space in image scene B in the corresponding neural map are calculated. The feature information B5 of the three-dimensional space area occupied by the space in image scene B in the corresponding neural map is replaced with the feature information A5, that is, the image scene replacement process is completed through the neural map.

[0083] It should be noted that the image scene replacement here can be a replacement for the same image scene or a cross-scene replacement for different image scenes. The way of image scene replacement is not limited here.

[0084] 4) Map update application

[0085] As time changes, the objects in the same scene will also change accordingly. For example, during the accelerating construction of a city, the buildings in the city will constantly change. Therefore, the neural map also needs to be continuously updated as the scene changes. Figure 8 It is a schematic diagram of the neural map update in an embodiment of the present invention. As Figure 8 shown, the feature information corresponding to the three-dimensional space region in the neural map before update is fused with the feature information corresponding to the three-dimensional space in the neural map to be updated, so as to obtain the updated neural map. The neural map before update, the neural map to be updated, and the updated neural map present different rendering effects respectively when implementing image rendering.

[0086] It should be noted that the process of feature fusion can refer to the feature fusion process of the aerial photography map and the ground photography map in the neural map construction process described above, and will not be elaborated here.

[0087] Figure 9 It is a flowchart of a neural map construction method in an embodiment of the present invention. As Figure 9 shown, this neural map construction method is applied to a cloud server, and this neural map construction method includes:

[0088] S101: Obtain the first acquisition device data collected by the first acquisition device located under the cloud and the second acquisition device data collected by the second acquisition device located under the cloud. Among them, the first acquisition device data is an aerial photography image for the first shooting scene, and the second acquisition device data is a ground photography image for the first shooting scene.

[0089] S102: Align the first pose information of the aerial photography image and the second pose information of the ground photography image to the same geographic coordinate system.

[0090] S103: Extract the first feature information of each pixel of the aerial photography image and the second feature information of each pixel of the ground photography image respectively.

[0091] S104: Construct a neural map of the first shooting scene based on the geographic coordinate system aligned with the first pose information and the second pose information, the first feature information, and the second feature information.

[0092] This neural map includes multiple grids, and each grid includes coordinate position information and feature information. The coordinate position information is the spatial position information mapped by the pixel points of the aerial photography image and / or the ground photography image of the first shooting scene, and the feature information is the fusion information of the first feature information of each pixel of the aerial photography image and the second feature information of each pixel of the ground photography image corresponding to the spatial position.

[0093] Steps S101 - S104 use different acquisition devices to collect aerial images and ground images for the same shooting scene respectively. Then, the aerial images and ground images are aligned to the same geographic coordinate system through pose calculation and registration, and the pose information and feature vectors of each pixel of the aerial images and ground images are extracted to construct an initial neural map. The neural map is cut into multiple three-dimensional space grids, and each pixel of the aerial images and / or ground images is mapped to the three-dimensional space grids of the neural map. The feature vectors of each pixel of the aerial images and / or ground images are fused into the three-dimensional space grids of the neural map. At this time, each three-dimensional grid in the neural map contains the fused feature information of the aerial images and / or ground images.

[0094] The above steps construct a neural map through processes such as pose calculation, registration, and feature fusion for the captured images at different shooting angles of the same scene by different acquisition devices, making the features in the neural map more comprehensive and detailed, and applicable to the subsequent positioning and rendering functions of images under multiple acquisition devices and multiple shooting angles.

[0095] S105: Based on the neural map and the aerial image, use the positioning model to calculate the third pose information of the aerial image, obtain the first error data between the first pose information and the third pose information, and feedback the first error data to the neural map and the positioning model; based on the neural map and the ground image, use the positioning model to calculate the fourth pose information of the ground image, obtain the second error data between the second pose information and the fourth pose information, and feedback the second error data to the neural map and the positioning model; based on the first error data and the second error data, update the feature information in the neural map and the parameter information of the positioning model.

[0096] S106: Based on the neural map and the first pose information, use the rendering model to calculate the first rendered image of the aerial image, obtain the third error data between the first rendered image and the aerial image, and feedback the third error data to the neural map and the rendering model; based on the neural map and the second pose information, use the rendering model to calculate the second rendered image of the ground image, obtain the fourth error data between the second rendered image and the ground image, and feedback the fourth error data to the neural map and the rendering model; based on the third error data and the fourth error data, update the feature information in the neural map and the parameter information of the rendering model.

[0097] Steps S105 - S106 train the constructed neural map through the positioning model and the rendering model respectively, and continuously update the neural map through error analysis based on the training results, so that the neural map is more accurate in the subsequent application stage.

[0098] Further, the method further includes performing stylized label encoding on the aerial captured image and the ground captured image, where the stylized label is used to indicate the image style type to which the aerial captured image and the ground captured image belong. The stylized label encoding facilitates the user to customize the style type of image rendering by inputting the stylized label encoding when using the neural map subsequently.

[0099] The trained map in the application stage further includes the following steps:

[0100] Receiving the image to be located in the second capture scene, where the second capture scene is within the range of the first capture scene, and calculating the fifth pose information of the image to be located based on the neural map and the image to be located by using the positioning model. When the user inputs the image to be located, and the image to be located is within the range of the scene captured when constructing the neural map, the pose information of the image to be located can be obtained by combining the neural map with the positioning model.

[0101] Receiving the sixth pose information of the second capture scene, and calculating the third rendered image corresponding to the sixth pose information based on the neural map and the sixth pose information by using the rendering model. When the user inputs the image to be rendered, and the rendered image is within the range of the scene captured when constructing the neural map, the rendered image of the image to be rendered can be obtained by combining the neural map with the rendering model.

[0102] Receiving the seventh pose information of the second capture scene and the first stylized label, and calculating the fourth rendered image corresponding to the seventh pose information based on the neural map, the seventh pose information and the first stylized label by using the rendering model, where the first stylized label is used to indicate the image style type to which the fourth rendered image belongs. When the user inputs the image to be rendered and the customized stylized label, the rendered image conforming to the customized style can be obtained by combining the neural map with the rendering model.

[0103] Receiving a partial area selection instruction for the first image of the third capture scene, where the third capture scene is within the range of the first capture scene, identifying the first spatial information corresponding to the partial area of the first image in the first image, and locating the first grid information corresponding to the first spatial information in the neural map; receiving a partial area migration instruction for the second image of the third capture scene, identifying the second spatial information corresponding to the partial area of the second image in the second image, and locating the second grid information corresponding to the second spatial information in the neural map; replacing the second grid information in the neural map with the first grid information, and updating the neural map. The user can realize the replacement operation of partial areas in two images by performing partial area selection operations on the two images and replacing the grid information corresponding to the partial areas in the neural map.

[0104] Further, the above first feature information, second feature information, and third feature information include semantic information, point-line-plane information, density information, illumination information, color information, texture information, and material information corresponding to the first shooting scene. Different feature information enables more comprehensive details in the feature fusion of aerial shooting images and ground shooting images.

[0105] The method for constructing a neural map provided by this application meets the service requirements of high-precision rendering and positioning in the entire airspace, so as to solve the problem that the traditional map construction method is relatively single and cannot simultaneously achieve high-precision rendering and positioning in large-scale scene modeling.

[0106] Based on the process of the method for constructing a neural map provided above, an embodiment of the present invention further discloses a neural map construction device. Figure 10 It is a schematic diagram of a neural map construction device according to an embodiment of the present invention, as Figure 10 shown. This neural map construction device is applied to a cloud server, and this neural map construction device includes:

[0107] An acquisition module 101, configured to acquire first acquisition device data acquired by a first acquisition device located under the cloud and second acquisition device data acquired by a second acquisition device located under the cloud. Among them, the first acquisition device data is an aerial shooting image for the first shooting scene, and the second acquisition device data is a ground shooting image for the first shooting scene.

[0108] An alignment module 102, configured to align the first pose information of the aerial shooting image and the second pose information of the ground shooting image to the same geographic coordinate system.

[0109] An extraction module 103, configured to respectively extract first feature information of each pixel of the aerial shooting image and second feature information of each pixel of the ground shooting image.

[0110] A construction module 104, configured to construct a neural map of the first shooting scene based on the geographic coordinate system aligned with the first pose information and the second pose information, the first feature information, and the second feature information. The neural map includes multiple grids, and each grid includes coordinate position information and feature information. The coordinate position information is the spatial position information mapped by pixel points of the aerial shooting image and / or the ground shooting image of the first shooting scene, and the feature information is the fusion information of the first feature information of each pixel of the aerial shooting image corresponding to the spatial position and the second feature information of each pixel of the ground shooting image.

[0111] An error calculation module 105, configured to calculate the third pose information of the aerial shooting image based on the neural map and the aerial shooting image by using a positioning model, obtain first error data between the first pose information and the third pose information, and feedback the first error data to the neural map and the positioning model.

[0112] The error calculation module 105 is further configured to calculate the fourth pose information of the ground captured image by using the positioning model based on the neural map and the ground captured image, obtain the second error data between the second pose information and the fourth pose information, and feed back the second error data to the neural map and the positioning model.

[0113] The update module 106 is configured to update the third feature information in the neural map and the parameter information of the positioning model based on the first error data and the second error data.

[0114] The error calculation module 105 is further configured to calculate the first rendered image of the aerial capture pose by using the rendering model based on the neural map and the first pose information, obtain the third error data between the first rendered image and the aerial capture image, and feed back the third error data to the neural map and the rendering model.

[0115] The error calculation module 105 is further configured to calculate the second rendered image of the ground capture pose by using the rendering model based on the neural map and the second pose information, obtain the fourth error data between the second rendered image and the ground capture image, and feed back the fourth error data to the neural map and the rendering model.

[0116] The update module 106 is further configured to update the third feature information in the neural map and the parameter information of the rendering model based on the third error data and the fourth error data.

[0117] The label encoding module 107 is configured to perform stylized label encoding on the aerial capture image and the ground capture image, and the stylized label is used to indicate the image style type to which the aerial capture image and the ground capture image belong.

[0118] The positioning module 108 is configured to receive the image to be positioned in the second capture scene, where the second capture scene is within the range of the first capture scene, and calculate the fifth pose information of the image to be positioned by using the positioning model based on the neural map and the image to be positioned.

[0119] The rendering module 109 is configured to receive the sixth pose information of the second capture scene, and calculate the third rendered image corresponding to the sixth pose information by using the rendering model based on the neural map and the sixth pose information.

[0120] The rendering module 109 is further configured to receive the seventh pose information of the second capture scene and the first stylized label, and calculate the fourth rendered image corresponding to the seventh pose information by using the rendering model based on the neural map, the seventh pose information and the first stylized label, where the first stylized label is used to indicate the image style type to which the fourth rendered image belongs.

[0121] The update module 106 is further configured to receive a partial area selection instruction for a first image of a third shooting scene, where the third shooting scene is within the range of the first shooting scene, identify first spatial information corresponding to the partial area of the first image in the first image, and locate first grid information corresponding to the first spatial information in the neural map;

[0122] The update module 106 is further configured to receive a partial area migration instruction for a second image of the third shooting scene, identify second spatial information corresponding to the partial area of the second image in the second image, and locate second grid information corresponding to the second spatial information in the neural map;

[0123] The update module 106 is further configured to replace the second grid information with the first grid information and update the neural map.

[0124] It should be noted that the acquisition module 101, the alignment module 102, the extraction module 103, the construction module 104, the error calculation module 105, the update module 106, the label encoding module 107, the positioning module 108, and the rendering module 109 can all be implemented by software or by hardware. Exemplarily, next, taking the cloud resource configuration interface acquisition module 101 as an example, the implementation manner of the interface acquisition module 101 will be introduced. Similarly, the implementation manners of the alignment module 102, the extraction module 103, the construction module 104, the error calculation module 105, the update module 106, the label encoding module 107, the positioning module 108, and the rendering module 109 can refer to the implementation manner of the acquisition module 101.

[0125] When implemented by software, the acquisition module 101 can be an application program or a code block running on a computer device. Among them, the computer device can be at least one of computing devices such as a physical host, a virtual machine, and a container. Further, the above computer device can be one or more. For example, the acquisition module 101 can be an application program running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the application program can be distributed in the same availability zone (AZ), or can be distributed in different AZs. The multiple hosts / virtual machines / containers for running the application program can be distributed in the same region, or can be distributed in different regions. Among them, generally, one region can include multiple AZs.

[0126] Similarly, the multiple hosts / virtual machines / containers for running the application program can be distributed in the same virtual private cloud (VPC), or can be distributed in multiple VPCs. Among them, generally, one region can include multiple VPCs, and one VPC can include multiple AZs.

[0127] When implemented by hardware, the obtaining module 101 may include at least one computing device, such as a server, etc. Alternatively, the cloud resource configuration interface providing module 101 may also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0128] The multiple computing devices included in the obtaining module 101 may be distributed in the same AZ or in different AZs. The multiple computing devices included in the cloud resource configuration interface providing module 101 may be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the obtaining module 101 may be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0129] It should be noted that the obtaining module 101, the alignment module 102, the extraction module 103, the construction module 104, the error calculation module 105, the update module 106, the label encoding module 107, the positioning module 108, and the rendering module 109 can all be used to execute some or all of the steps in the method for constructing a neural map.

[0130] Each module of the neural map construction device disclosed in the embodiments of the present invention has a clear division of labor and close cooperation. The cooperation of each module with each other can efficiently complete the construction service for the neural network map.

[0131] The present invention also provides a computing device. Please refer to the following Figure 11 , Figure 11 is a schematic structural diagram of a computing device for the method of constructing a neural map according to an embodiment of the present invention. The computing device 100 includes: a bus 104, a processor 106, a memory 105, and a communication interface 107. The processor 106, the memory 105, and the communication interface 107 communicate with each other through the bus 104. The computing device 100 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 100.

[0132] The bus 104 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 11 only one line is used in the figure, but it does not mean that there is only one bus or one type of bus. The bus 104 can include a path for transmitting information between various components of the computing device 100 (such as the memory 105, the processor 106, and the communication interface 107).

[0133] The processor 106 can include any one or more of processors such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Micro Processor (MP), or a Digital Signal Processor (DSP).

[0134] The memory 105 can include volatile memory, such as Random Access Memory (RAM). The processor 106 can also include non-volatile memory, such as Read-Only Memory (ROM), flash memory, a Hard Disk Drive (HDD), or a Solid State Drive (SSD).

[0135] The executable program code is stored in the memory 105, and the processor 106 executes the executable program code to respectively implement the functions of the foregoing acquisition module 101, alignment module 102, extraction module 103, construction module 104, error calculation module 105, update module 106, label encoding module 107, positioning module 108, and rendering module 109, so as to implement the configuration method of the cloud connection service based on the public cloud. That is, the instructions for the cloud management platform to execute the configuration method of the cloud connection service based on the public cloud are stored on the memory 105.

[0136] The communication interface 107 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 100 and other devices or a communication network.

[0137] An embodiment of the present invention further provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0138] Please refer to the following Figure 12 , Figure 12 which is a schematic structural diagram of a computing device cluster for the method of constructing a neural map according to an embodiment of the present invention. The computing device cluster includes at least one computing device 100. Instructions for executing a configuration method for cloud connection services based on a public cloud by the same cloud management platform can be stored in the memory 105 of one or more of the computing devices 100 in the computing device cluster.

[0139] In some possible implementation manners, one or more of the computing devices 100 in the computing device cluster can also be used to execute some instructions of the cloud management platform for executing the configuration method for cloud connection services based on a public cloud. In other words, a combination of one or more computing devices 100 can jointly execute the instructions of the cloud management platform for executing the configuration method for cloud connection services based on a public cloud.

[0140] It should be noted that the memories 105 in different computing devices 100 in the computing device cluster can store different instructions for executing some functions of the cloud management platform. That is, the instructions stored in the memories 105 of different computing devices 100 can implement the functions of one or more of the acquisition module 101, the alignment module 102, the extraction module 103, the construction module 104, the error calculation module 105, the update module 106, the label encoding module 107, the positioning module 108, and the rendering module 109.

[0141] An embodiment of the present invention further provides a computer program product including instructions. The computer program product can be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computer device, at least one computer device is caused to execute the above-mentioned configuration method for cloud connection services based on a public cloud applied to the cloud management platform.

[0142] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the above-described configuration method applied to a cloud management platform for performing cloud connection services based on a public cloud.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a neural map, characterized in that, the method is applied to a cloud server, and the method includes: obtaining first acquisition device data acquired by a first acquisition device located under the cloud and second acquisition device data acquired by a second acquisition device located under the cloud, wherein the first acquisition device data is an aerial image of a first shooting scene, and the second acquisition device data is a ground image of the first shooting scene; aligning the first pose information of the aerial image and the second pose information of the ground image to the same geographic coordinate system; extracting first feature information of each pixel of the aerial image and second feature information of each pixel of the ground image respectively; constructing a neural map of the first shooting scene based on the geographic coordinate system aligned with the first pose information and the second pose information, the first feature information and the second feature information, the neural map includes a plurality of grids, each grid includes coordinate position information and third feature information, the coordinate position information is the spatial position information mapped by pixel points of the aerial image and / or the ground image of the first shooting scene, and the third feature information is the fusion information of the first feature information of each pixel of the aerial image and the second feature information of each pixel of the ground image corresponding to the spatial position.

2. The method according to claim 1, characterized in that, the method further includes: calculating third pose information of the aerial image based on the neural map and the aerial image by using a positioning model, obtaining first error data between the first pose information and the third pose information, and feeding back the first error data to the neural map and the positioning model; calculating fourth pose information of the ground image based on the neural map and the ground image by using a positioning model, obtaining second error data between the second pose information and the fourth pose information, and feeding back the second error data to the neural map and the positioning model; updating the third feature information in the neural map and the parameter information of the positioning model based on the first error data and the second error data.

3. The method according to claim 1 or 2, characterized in that, the method further includes: calculating a first rendered image of the aerial image based on the neural map and the first pose information by using a rendering model, obtaining third error data between the first rendered image and the aerial image, and feeding back the third error data to the neural map and the rendering model; calculating a second rendered image of the ground image based on the neural map and the second pose information by using a rendering model, obtaining fourth error data between the second rendered image and the ground image, and feeding back the fourth error data to the neural map and the rendering model; updating the third feature information in the neural map and the parameter information of the rendering model based on the third error data and the fourth error data.

4. The method according to any one of claims 1 or 3, It is characterized in that The method further includes: Performing stylized label encoding on the aerial captured image and the ground captured image, where the stylized label is used to indicate the image style type to which the aerial captured image and the ground captured image belong.

5. The method according to any one of claims 2 or 4, It is characterized in that The method further includes: Receiving a to-be-localized image of a second shooting scene, where the second shooting scene is within the range of the first shooting scene, and based on the neural map and the to-be-localized image, using the localization model to calculate the fifth pose information of the to-be-localized image.

6. The method according to claim 5, It is characterized in that The method further includes: Receiving the sixth pose information of the second shooting scene, and based on the neural map and the sixth pose information, using the rendering model to calculate a third rendered image corresponding to the sixth pose information.

7. The method according to any one of claims 5 or 6, It is characterized in that The method further includes: Receiving the seventh pose information and a first stylized label of the second shooting scene, and based on the neural map, the seventh pose information and the first stylized label, using the rendering model to calculate a fourth rendered image corresponding to the seventh pose information, where the first stylized label is used to indicate the image style type to which the fourth rendered image belongs.

8. The method according to any one of claims 2 or 7, It is characterized in that The method further includes: Receiving a partial area selection instruction for a first image of a third shooting scene, where the third shooting scene is within the range of the first shooting scene, identifying first spatial information corresponding to a partial area of the first image in the first image, and locating first grid information corresponding to the first spatial information in the neural map; Receiving a partial area migration instruction for a second image of the third shooting scene, identifying second spatial information corresponding to a partial area of the second image in the second image, and locating second grid information corresponding to the second spatial information in the neural map; Replacing the second grid information in the neural map with the first grid information, and updating the neural map.

9. The method according to any one of claims 1 or 8, It is characterized in that The first feature information, the second feature information and the third feature information include semantic information, point-line-plane information, density information, illumination information, color information, texture information and material information corresponding to the first shooting scene.

10. A neural map construction device, It is characterized in that The device is applied to a cloud server, and the device includes: An acquisition module, configured to acquire first acquisition device data acquired by a first acquisition device located under the cloud and second acquisition device data acquired by a second acquisition device located under the cloud, where the first acquisition device data is an aerial captured image of a first shooting scene, and the second acquisition device data is a ground captured image of the first shooting scene; An alignment module, configured to align the first pose information of the aerial captured image and the second pose information of the ground captured image to the same geographic coordinate system; An extraction module for separately extracting first feature information of each pixel of the aerial captured image and second feature information of each pixel of the ground captured image; A construction module for constructing a neural map of the first capture scene based on the geographical coordinate system aligned with the first pose information and the second pose information, the first feature information, and the second feature information. The neural map includes a plurality of grids, each grid including coordinate position information and third feature information. The coordinate position information is the spatial position information mapped by pixel points of the aerial captured image and / or the ground captured image of the first capture scene, and the third feature information is the fusion information of the first feature information of each pixel of the aerial captured image and the second feature information of each pixel of the ground captured image corresponding to the spatial position.

11. The apparatus according to claim 10, wherein, the apparatus further includes: An error calculation module for calculating third pose information of the aerial captured image based on the neural map and the aerial captured image using a positioning model, obtaining first error data between the first pose information and the third pose information, and feeding back the first error data to the neural map and the positioning model; The error calculation module is further configured to calculate fourth pose information of the ground captured image based on the neural map and the ground captured image using a positioning model, obtain second error data between the second pose information and the fourth pose information, and feed back the second error data to the neural map and the positioning model; An update module for updating the third feature information in the neural map and the parameter information of the positioning model based on the first error data and the second error data.

12. The apparatus according to claim 10 or 11, wherein, The error calculation module is further configured to calculate a first rendered image of the aerial captured image based on the neural map and the first pose information using a rendering model, obtain third error data between the first rendered image and the aerial captured image, and feed back the third error data to the neural map and the rendering model; The error calculation module is further configured to calculate a second rendered image of the ground captured image based on the neural map and the second pose information using a rendering model, obtain fourth error data between the second rendered image and the ground captured image, and feed back the fourth error data to the neural map and the rendering model; The update module is further configured to update the third feature information in the neural map and the parameter information of the rendering model based on the third error data and the fourth error data.

13. The apparatus according to any one of claims 10 or 12, wherein, the apparatus further includes: A label encoding module for performing stylized label encoding on the aerial captured image and the ground captured image. The stylized label is used to indicate the image style type to which the aerial captured image and the ground captured image belong.

14. The apparatus according to any one of claims 11 or 13, It is characterized in that the device further comprises: a positioning module, configured to receive a to-be-positioned image of a second shooting scene, where the second shooting scene is within the range of the first shooting scene, and calculate fifth pose information of the to-be-positioned image based on the neural map and the to-be-positioned image by using the positioning model.

15. The device according to claim 14, It is characterized in that the device further comprises: a rendering module, configured to receive sixth pose information of the second shooting scene, and calculate a third rendered image corresponding to the sixth pose information based on the neural map and the sixth pose information by using the rendering model.

16. The device according to claim 14 or 15, It is characterized in that the rendering module is further configured to receive seventh pose information of the second shooting scene and a first stylization label, and calculate a fourth rendered image corresponding to the seventh pose information based on the neural map, the seventh pose information and the first stylization label by using the rendering model, where the first stylization label is used to indicate the image style type to which the fourth rendered image belongs.

17. The device according to claim 14 or 16, It is characterized in that the updating module is further configured to receive a partial area selection instruction for a first image of a third shooting scene, where the third shooting scene is within the range of the first shooting scene, identify first spatial information corresponding to a partial area of the first image in the first image, and locate first grid information corresponding to the first spatial information in the neural map; the updating module is further configured to receive a partial area migration instruction for a second image of the third shooting scene, identify second spatial information corresponding to a partial area of the second image in the second image, and locate second grid information corresponding to the second spatial information in the neural map; the updating module is further configured to replace the second grid information with the first grid information and update the neural map.

18. The method according to any one of claims 10 or 17, It is characterized in that the first feature information, the second feature information and the third feature information include semantic information, point-line-plane information, density information, illumination information, color information, texture information and material information corresponding to the first shooting scene.

19. A computing device cluster, It is characterized in that it includes at least one computing device, and each computing device includes a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 9.

20. A computer program product containing instructions, It is characterized in that when the instructions are run by a computer device cluster, the computer device cluster is caused to execute the method according to any one of claims 1 to 9.

21. A computer-readable storage medium, It is characterized in that it includes computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 9.