Visual environment generation method, system and device based on neural radiation field and storage medium

Through multi-camera synchronous calibration and semantic segmentation technology, semantic point clouds are constructed and redundant voxels are cropped, which solves the problems of surge in model parameters and semantic mapping errors in existing technologies, realizes efficient visual environment generation, and meets real-time rendering requirements.

CN120672896AActive Publication Date: 2025-09-19ZHUHAI XIANG YI AVIATION TECH CO LTD

Patent Information

Application Number
CN202511190548.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-19
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing technologies have a surge in model parameters in large urban scenes and rely on dense perspective input, resulting in semantic mapping errors and artifacts in sparse scenes. The lack of scene structure priors leads to the accumulation of invalid voxels, limiting real-time rendering efficiency.

Method used

Through multi-camera synchronous calibration, multi-view images of wide-angle lenses are obtained, semantic segmentation is performed to generate image data with semantic labels, a semantic point cloud is constructed and non-critical areas are identified for redundant voxel cropping, and the cropped semantic point cloud is used to drive the neural radiation field model to generate arbitrary perspective scenes.

Benefits of technology

It achieves the goal of improving the efficiency and quality of visual environment generation under limited computing power conditions, meeting real-time rendering requirements, optimizing computing resource allocation, reducing model complexity, and eliminating semantic mapping errors and artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672896A_ABST
    Figure CN120672896A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle control, provides a visual environment generation method, system and device based on a neural radiation field and a storage medium, and solves the problem of poor visual environment generation effect. The method comprises the steps that semantic segmentation is carried out on a corrected image, image data with semantic tags are generated, and the semantic tags are used for marking categories of objects in an outdoor scene; updating the neural radiation field model based on the image data with the semantic tag, mapping the semantic tag to a voxel feature space corresponding to the neural radiation field model, and constructing a semantic point cloud; identifying a non-key area in the multi-view image, performing redundant voxel cutting on the non-key area, and generating a cut semantic point cloud; and utilizing the clipped semantic point cloud to drive a neural radiation field model, and generating any visual angle scene of the outdoor scene. According to the technical scheme, lightweight semantic modeling and efficient free viewpoint rendering of the outdoor visual environment are achieved, and the new visual angle generation efficiency and quality of the complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drone control technology, and in particular to a method and system for generating a visual environment based on a neural radiation field. Background Art

[0002] In application areas such as drone aerial surveys, strict requirements are placed on the efficient three-dimensional reconstruction and real-time vision generation of large-scale outdoor scenes. Large-scale reconstruction capabilities are required to fully restore the geometric structure and semantic information of kilometer-level scenes, while effectively overcoming lighting changes, object occlusion, and dynamic interference in the actual environment. Three-dimensional semantic consistency must be achieved to ensure that different types of objects such as buildings, roads, and vegetation can maintain unified recognition of semantic labels from any observation perspective. Moreover, under limited computing power conditions, it is necessary to accurately balance computing efficiency and output quality to ensure both reconstruction accuracy and real-time rendering performance requirements while avoiding memory overload problems.

[0003] The existing solution first decouples the dynamic scene, splitting the entire scene into two parts: "object" and "background". Each object instance is modeled using an independent small multi-layer perceptron, and its three-dimensional bounding box and radiation field parameters are combined; secondly, a multimodal data fusion strategy is adopted, using existing mature algorithms to pre-acquire the camera pose, object motion trajectory and two-dimensional semantic segmentation map, and the weight parameters of the multi-layer perceptron are jointly optimized by fusing self-supervised and pseudo-supervised learning; finally, an implicit scene representation is constructed, and comprehensive information including voxel density, radiosity and semantic labels is output, thereby supporting key tasks such as new perspective synthesis and three-dimensional scene editing, which is particularly suitable for real road environment scenes.

[0004] However, existing solutions will lead to a linear growth of model parameters when the number of instances in large urban scenes is huge, increasing the computational burden and memory usage; this solution relies heavily on dense multi-angle image input for training to ensure reconstruction integrity. When the input perspective is sparse, it is easy to produce semantic label mapping errors and false artifacts in non-critical areas due to insufficient geometric constraints; in addition, this solution lacks the use of prior knowledge of scene structure and does not introduce a spatial redundancy judgment mechanism, resulting in equal allocation of computing resources to areas far away from the center of view or with low visual contribution, causing invalid voxel accumulation, which in turn limits the efficiency of real-time rendering. Summary of the Invention

[0005] The present application provides a method and system for generating a visual environment based on a neural radiation field, which is used to solve the problems in the prior art such as insufficient modeling scalability leading to a surge in the number of model parameters, reliance on dense perspective input leading to semantic mapping errors and artifacts in sparse scenes, and lack of scene structure priors resulting in the accumulation of invalid voxels that limits real-time rendering efficiency, ultimately leading to poor visual environment generation.

[0006] In a first aspect, the present application provides a method for generating a visual environment based on a neural radiation field, comprising: Acquire multi-view images of an outdoor scene captured by a wide-angle lens, perform multi-camera synchronous calibration on the multi-view images, and generate corrected images; Performing semantic segmentation on the corrected image to generate image data with semantic labels, wherein the semantic labels are used to mark the categories of objects in the outdoor scene; In a cascaded rendering computer cluster, a neural radiation field model is updated based on the image data with the semantic labels, and the semantic labels are mapped to a voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud; Based on the semantic point cloud, identifying non-critical areas in the multi-view image, performing redundant voxel cropping on the non-critical areas, and generating a cropped semantic point cloud; The cropped semantic point cloud is used to drive the neural radiance field model to generate an arbitrary perspective view of the outdoor scene.

[0007] Optionally, the using the cropped semantic point cloud to drive the neural radiation field model to generate an arbitrary perspective view of an outdoor scene includes: Defining virtual camera parameters of the target viewing angle, wherein the virtual camera parameters include the virtual camera center position, posture and intrinsic parameter matrix; Based on the virtual camera parameters, starting from the center position of the virtual camera, emitting rays along each pixel direction to form a ray set; For each ray in the ray set, using the cropped semantic point cloud to drive adaptive sampling along the ray direction to generate a sampling point sequence; For each sampling point in the sampling point sequence, obtaining a color attribute and a density attribute of each sampling point from a voxel feature space of the neural radiation field model; Integrating the color attribute and density attribute of each sampling point along the ray direction to generate a final color value of each pixel; Aggregating the final color values ​​of all the pixels to form an arbitrary viewing angle view of the outdoor scene.

[0008] Optionally, for each ray in the ray set, using the cropped semantic point cloud to drive adaptive sampling along the ray direction to generate a sampling point sequence includes: Based on the spatial distribution of the cropped semantic point cloud, the ray is divided into a plurality of initial sampling intervals along the ray direction, and an initial sampling point set is generated in each initial sampling interval according to a uniform step size; In a neighborhood of the initial sampling point set, a dynamic sampling step is calculated based on the voxel density distribution of the cropped semantic point cloud and the spatial continuity of the semantic labels, and candidate sampling points are generated around the initial sampling point according to the dynamic sampling step; Performing semantic label consistency check on the candidate sampling point. When the similarity between the predicted semantic label of the candidate sampling point in the voxel feature space of the neural radiation field and the label of the semantic point cloud closest to the initial sampling point is less than a preset similarity, expanding the preset distance in the ray direction with the candidate sampling point as the center to generate a label difference area. Alternatively, when the similarity is greater than the preset similarity, retaining the candidate sampling point. Inserting encrypted sampling points in the ray direction of the label difference area; The initial sampling point set, the retained candidate sampling points and the encrypted sampling points are merged to generate a sampling point sequence along the ray direction.

[0009] Optionally, identifying a non-critical area in the multi-view image based on the semantic point cloud, performing redundant voxel cropping on the non-critical area, and generating a cropped semantic point cloud includes: Based on the object categories in the semantic labels, counting the occurrence frequencies of each object category in the spatial region, and marking regions where the occurrence frequencies of all physical categories are lower than a preset frequency threshold; Merging adjacent marked areas that meet a preset continuity condition to generate a non-critical area, wherein the preset continuity condition includes: the comprehensive similarity of the adjacent marked areas exceeds a preset similarity threshold; Locating a voxel set corresponding to the non-critical area in the semantic point cloud, and removing voxels with density values ​​lower than a density threshold; The retained voxels are recombined with the voxels in the semantic point cloud that do not belong to the non-critical area to generate a cropped semantic point cloud.

[0010] Optionally, merging adjacent marked areas that meet a preset continuity condition to generate a non-critical area includes: Each marked area is abstracted as a graph node. If the Euclidean distance between two adjacent marked areas is less than the preset spatial distance threshold, an undirected edge is established between the graph nodes to form an adjacency graph. Calculating semantic label similarity, boundary geometric similarity, and density distribution similarity for each undirected edge in the adjacency graph, and weighting the semantic label similarity, boundary geometric similarity, and density distribution similarity to obtain a corresponding comprehensive similarity; Iteratively merging the adjacent marked regions whose comprehensive similarities are greater than a preset similarity threshold to update the adjacency graph until the comprehensive similarities between all adjacent marked regions in the adjacency graph are lower than the preset similarity threshold, thereby obtaining candidate regions; Calculating the semantic entropy of the candidate regions, and taking the candidate regions whose semantic entropy is less than a preset semantic entropy threshold as valid merged regions; A screening process is performed on the valid merged area to generate a non-critical area, wherein the screening process includes at least one of the following: area threshold screening, aspect ratio constraint, and length-width ratio constraint.

[0011] Optionally, updating the neural radiation field model based on the image data with the semantic labels, and mapping the semantic labels to a voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud includes: Extracting geometric features and semantic features of each voxel from the image data with semantic labels; Updating parameters of the neural radiation field model by jointly optimizing the geometric features and the semantic features; Mapping the semantic labels to corresponding voxel locations in a voxel feature space of the neural radiation field model; The spatial coordinates of the voxel positions and the semantic labels are fused to construct a semantic point cloud.

[0012] Optionally, updating the parameters of the neural radiation field model by jointly optimizing the geometric features and the semantic features includes: calculating a matching error between the geometric feature and a corresponding three-dimensional geometric representation in the neural radiation field model to generate a geometric loss; Calculating the consistency error between the semantic feature and the semantic label to generate a semantic loss; Jointly optimizing the geometric loss and the semantic loss to generate a joint optimization loss; Parameters of the neural radiation field model are updated according to the joint optimization loss.

[0013] In a second aspect, the present application provides a visual environment generation system based on neural radiation fields, comprising: An acquisition module is used to acquire multi-view images of an outdoor scene captured by a wide-angle lens, perform multi-camera synchronous calibration on the multi-view images, and generate a corrected image; a generation module, configured to perform semantic segmentation on the corrected image to generate image data with semantic labels, wherein the semantic labels are used to mark the categories of objects in the outdoor scene; A construction module is configured to update a neural radiation field model based on the image data with semantic labels in a cascade rendering computer cluster, and map the semantic labels to a voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud; an identification module, configured to identify non-critical areas in the multi-view image based on the semantic point cloud, perform redundant voxel cropping on the non-critical areas, and generate a cropped semantic point cloud; A driving module is used to drive the neural radiation field model using the cropped semantic point cloud to generate an arbitrary perspective view of the outdoor scene.

[0014] In a third aspect, the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a method for generating a visual environment based on a neural radiation field as described in any one of the first aspects.

[0015] In a fourth aspect, the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements a method for generating a visual environment based on a neural radiation field as described in any one of the first aspects.

[0016] In the present application, a method for generating a visual environment based on a neural radiation field is provided, which includes: obtaining multi-perspective images of an outdoor scene captured by a wide-angle lens, performing multi-camera synchronous calibration on the multi-perspective images, and generating a corrected image; performing semantic segmentation on the corrected image to generate image data with semantic labels, and the semantic labels are used to mark the categories of objects in the outdoor scene; in a cascade rendering computer cluster, updating a neural radiation field model based on the image data with semantic labels, and mapping the semantic labels to the voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud; based on the semantic point cloud, identifying non-critical areas in the multi-perspective images, performing redundant voxel cropping on the non-critical areas, and generating a cropped semantic point cloud; using the cropped semantic point cloud to drive the neural radiation field model to generate an arbitrary perspective visual scene of the outdoor scene, thereby improving the generation effect of the visual environment.

[0017] Beneficial effects of this application: This application uses multi-camera synchronous calibration and wide-angle lens multi-view images for precise spatial and temporal alignment, effectively eliminating lens distortion and shooting timing deviation, and laying the foundation for geometric consistency for subsequent processing; on this basis, semantic segmentation technology is used to enhance data expression, accurately label object categories at the pixel level, and elevate visual information to the semantic level, thus laying the foundation for structured understanding of the scene; then a semantic neural radiation field is constructed, and cascade computing capabilities are used to fuse two-dimensional semantic label information into three-dimensional voxel space to form a point cloud model that simultaneously contains geometric shape, radiosity, and semantic attributes, realizing implicit semantic expression of the scene; then, non-critical areas with low visual contribution are identified based on semantic labels, and redundant invalid voxels are dynamically cropped to effectively reduce the overall complexity of the model; finally, the streamlined semantic point cloud is used to drive the rendering pipeline, which ensures the reconstruction accuracy of key areas while improving the synthesis speed of new perspective images to meet the application needs of real-time interaction.

[0018] Furthermore, a set of rays covering all pixel directions is first emitted according to the position, posture and internal parameters of the virtual camera from the target perspective; then semantic-driven adaptive sampling is implemented, the initial interval is divided along each ray direction and initial sampling points are generated with a uniform step size, and then the dynamic step size is calculated in the neighborhood of the initial point according to the spatial distribution and semantic continuity of the semantic point cloud to generate candidate sampling points; then semantic verification and expansion are performed: when the similarity between the semantic label predicted by the candidate point and the label of the nearest initial point is lower than the set threshold, the preset distance is expanded with the point as the center to generate a label difference area; if the similarity reaches or exceeds the threshold, the candidate point is retained; next, encrypted sampling points are inserted along the ray direction in the identified label difference area; the initial point, retained candidate points and encrypted points are merged to form the final sampling point sequence; finally, the color and density attributes of the sampling points in the sequence are queried, and the pixel color values ​​of any perspective are generated by integrating and accumulating along the ray. This process implements semantically guided sampling optimization, that is, dynamically adjusting the sampling density based on semantic information, automatically encrypting sampling in areas with rich geometric features, and sparsely sampling in flat areas, effectively reducing invalid sampling points; at the same time, it enhances boundary accuracy, accurately identifies and encrypts semantic mutation areas through semantic consistency verification, and eliminates blurring artifacts at object contours caused by traditional methods; ultimately, it optimizes the redistribution of computing resources, focusing limited computing power on key areas with significant semantic differences, avoiding the waste of resources from uniform sampling, and accelerating the ray integration process while maintaining rendering quality to meet the needs of high-frame-rate real-time scene generation.

[0019] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 A flowchart of a method for generating a visual environment based on a neural radiation field provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a visual environment generation system based on neural radiation fields provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0023] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 11, 12, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.

[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0025] In order to solve the problems in the existing technology of insufficient modeling scalability leading to a surge in the number of model parameters, reliance on dense perspective input causing semantic mapping errors and artifacts in sparse scenes, and lack of scene structure prior causing invalid voxel accumulation to limit real-time rendering efficiency, which ultimately lead to poor visual environment generation, the embodiment of the present application provides a visual environment generation method based on neural radiation field, which adopts the following ideas: through wide-angle lens multi-perspective image acquisition and multi-camera spatiotemporal synchronous calibration, distortion and timing deviation are eliminated to establish a geometric consistency foundation; semantic segmentation technology is used to annotate pixel-level object categories on the corrected image, and the dimension is upgraded to the semantic information layer to support scene structured analysis; semantic labels and neural radiation field voxel feature space are integrated in the cascade computing power architecture to construct a three-dimensional point cloud model with geometric, radiometric and semantic attributes; non-critical areas are identified based on semantic importance, and redundant voxels are dynamically cropped to compress model complexity; finally, a streamlined semantic point cloud is used to drive the neural radiation field rendering pipeline to achieve high-precision real-time visual generation of large outdoor scenes from any perspective, while ensuring the reconstruction quality of key areas and optimizing the efficiency of computing resource allocation.

[0026] Figure 1 A flowchart of a method for generating a visual environment based on a neural radiation field is provided in an embodiment of the present application, such as Figure 1 As shown, the method includes: S11 , acquiring multi-view images of an outdoor scene captured by a wide-angle lens, performing multi-camera synchronous calibration on the multi-view images, and generating a corrected image.

[0027] Among them, a wide-angle lens refers to an optical device with a focal length shorter than a standard lens, which is used to capture wide-field images of a large range of outdoor scenes, including fisheye lenses or ultra-wide-angle lens types. Outdoor scenes refer to open space areas in natural or urban environments, containing static objects such as buildings, roads, and vegetation, as well as dynamic factors such as lighting changes. Multi-perspective images refer to a collection of images of the same scene collected synchronously from different angles by multiple spatially distributed cameras. Multi-camera synchronous calibration refers to the process of calculating the relative position and timing deviation between cameras based on common view calibration objects, which is used to eliminate lens distortion and shooting time difference. Corrected images refer to multi-perspective images that have undergone geometric distortion correction and spatiotemporal alignment processing, with pixel-level coordinate consistency.

[0028] In an embodiment of the present application, multi-perspective images are first captured synchronously by a wide-angle lens array arranged in an outdoor scene. Secondly, multi-camera synchronous calibration technology is used to perform spatiotemporal alignment processing on the images, specifically including lens distortion correction and timing deviation compensation based on a checkerboard calibration plate, to finally generate a corrected image.

[0029] S12. Perform semantic segmentation on the corrected image to generate image data with semantic labels. The semantic labels are used to mark the categories of objects in the outdoor scene.

[0030] Semantic segmentation is a technique that uses deep learning models to classify images at the pixel level, identifying object category boundaries. Semantic labeling is a discrete classification identifier assigned to each pixel, used to label the specific category attributes of objects in an image. Image data refers to the digital image carrier containing a matrix of pixel values ​​and associated metadata.

[0031] In an embodiment of the present application, the corrected image is first input into a pre-trained semantic segmentation network, and then the image content is analyzed pixel by pixel through a convolutional neural network to identify and label the category attributes of objects such as buildings, roads, and vegetation, and finally generate image data with semantic labels.

[0032] S13. In the cascade rendering computer cluster, the neural radiation field model is updated based on the image data with semantic labels, and the semantic labels are mapped to the voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud.

[0033] Among them, the cascade rendering computer cluster refers to a distributed computing architecture composed of multi-node servers, which is used to parallelize the processing of neural radiation field training tasks. The neural radiation field model refers to a machine learning model that implicitly expresses three-dimensional scenes through a multi-layer perceptron, and can output the density and radiance of any spatial point. The voxel feature space refers to a data structure that discretizes three-dimensional space into grid cells, each of which stores geometric, semantic, and other feature vectors. The semantic point cloud refers to a set of three-dimensional points that integrates spatial coordinates, radiation attributes, and semantic labels to form an explicit semantic expression of the scene.

[0034] In an embodiment of the present application, a neural radiation field model is first loaded into a cascaded rendering computer cluster, and then the model parameters are iteratively optimized using image data with semantic labels. The two-dimensional semantic labels are then mapped to corresponding coordinates in the three-dimensional voxel feature space, and finally a semantic point cloud is constructed.

[0035] S14. Based on the semantic point cloud, identify non-critical areas in the multi-view image, perform redundant voxel cropping on the non-critical areas, and generate a cropped semantic point cloud.

[0036] Non-critical regions refer to scene subspaces with low visual contribution, such as distant sky or a single-textured wall, identified through semantic frequency statistics. Redundant voxel clipping is an operation that removes invalid voxels based on density thresholds and semantic importance to reduce model complexity. The cropped semantic point cloud is a streamlined 3D point set after redundant voxel clipping, preserving the complete semantic attributes of key regions.

[0037] In an embodiment of the present application, the frequency of occurrence of spatial regions is first counted based on the object category distribution of the semantic point cloud. Secondly, the regions with a frequency lower than a threshold are marked as candidate regions. Then, adjacent candidate regions whose geometric continuity meets preset conditions are merged to form non-critical regions. Finally, a redundant voxel cropping operation is performed on the region to remove low-density voxels and reorganize the retained voxels to generate a cropped semantic point cloud.

[0038] S15. Use the cropped semantic point cloud to drive the neural radiation field model to generate arbitrary perspective views of outdoor scenes.

[0039] Among them, arbitrary perspective view refers to the free perspective rendering image generated by virtual camera parameters to meet the needs of real-time interactive applications.

[0040] In an embodiment of the present application, the virtual camera parameters are first set according to the target perspective, and then a set of rays covering the pixel direction is emitted. Then, adaptive sampling is driven based on the cropped semantic point cloud to generate a sampling sequence. Finally, the color and density attributes of the sampling points are integrated along the rays to aggregate and generate an arbitrary perspective view of the outdoor scene.

[0041] The following is a specific example: First, multi-perspective images of urban roads are collected through wide-angle lenses installed on a swarm of drones, and the spatiotemporal parameters of each lens are aligned using a calibration plate to generate a corrected image. Secondly, the Deep Lab v3+ model is used to perform semantic segmentation on the corrected image and annotate category labels such as vehicles, pedestrians, and traffic signs. Then, in a cascade rendering computer cluster equipped with four graphics processor servers, the semantic labels are mapped to the neural radiation field voxel space to construct a semantic point cloud. Subsequently, the distant billboard areas that appear infrequently in the point cloud are identified as non-critical areas, and their low-density voxels are cropped to form a streamlined point cloud. Finally, the virtual camera's bird's-eye view parameters are set, and the color values ​​of the sampling points along the pixel rays are integrated to generate a panoramic bird's-eye view rendering of the road.

[0042] By executing S11 to S15, the embodiment of the present application ensures geometric consistency through multi-camera calibration, injects structured understanding capabilities through semantic segmentation, and constructs a three-dimensional neural field expression that integrates semantics under the support of distributed computing power; dynamically crops voxels in low-value areas based on semantic importance, reducing the computational load; and finally, through a semantically driven adaptive sampling mechanism, encrypted sampling in key areas improves boundary accuracy, and sparse sampling in non-key areas optimizes resource allocation, thereby achieving high-quality real-time rendering of large-scale outdoor scenes.

[0043] In one possible embodiment, S15, using the cropped semantic point cloud to drive the neural radiance field model to generate an arbitrary perspective view of the outdoor scene, includes: Step 151: Define the virtual camera parameters of the target viewing angle, where the virtual camera parameters include the virtual camera center position, posture, and intrinsic parameter matrix.

[0044] The target viewing angle refers to the virtual observation orientation specified by the user or automatically generated by the system, including pitch angle, azimuth angle, and observation distance parameters. The virtual camera center position refers to the coordinate values ​​of the virtual camera's optical center on the X, Y, and Z axes in the three-dimensional world coordinate system. The pose refers to the rotation state of the virtual camera coordinate system relative to the world coordinate system, expressed using a rotation matrix or quaternion. The intrinsic parameter matrix refers to the matrix that describes the imaging characteristics of the virtual camera, including focal length, principal point offset, and pixel aspect ratio parameters.

[0045] In an embodiment of the present application, the target observation angle is first determined according to user interaction instructions or preset path planning, and then the spatial coordinates of the center position of the virtual camera, the posture rotation matrix of the lens orientation, and the intrinsic parameter matrix of the focal length and imaging plane are defined, and finally a complete set of virtual camera parameters is formed.

[0046] Step 152: Based on the virtual camera parameters, rays are emitted from the center position of the virtual camera along the direction of each pixel to form a ray set.

[0047] Here, a pixel direction is the normalized 3D direction vector pointing from the camera center to a specific pixel in the imaging plane. A ray set is a cluster of spatial rays covering all pixels in the virtual camera imaging plane, with each ray corresponding to a pixel.

[0048] In an embodiment of the present application, the direction vector corresponding to each pixel in the imaging plane is first calculated based on the virtual camera intrinsic parameter matrix, and then spatial rays are emitted along the direction vector from the center position of the virtual camera, and finally a set of rays covering all pixels is generated.

[0049] Step 153: For each ray in the ray set, use the cropped semantic point cloud to drive adaptive sampling along the ray direction to generate a sampling point sequence.

[0050] Adaptive sampling is a strategy that dynamically adjusts the density of sampling points based on scene characteristics, driven by the geometric and semantic features of the semantic point cloud. A sampling point sequence is a set of 3D points arranged in spatial order along a single ray, used for subsequent attribute queries and integral calculations.

[0051] In an embodiment of the present application, an initial sampling interval with a uniform step size is first divided along the ray direction and an initial sampling point set is generated. Then, a dynamic step size is calculated based on the voxel density distribution and semantic continuity of the cropped semantic point cloud, and candidate sampling points are generated in the neighborhood of the initial point. Then, the semantic label similarity between the candidate point and the adjacent initial point is checked. If it is lower than a threshold, the label difference area is expanded and encrypted sampling points are inserted. Finally, the initial point, retained candidate points and encrypted points are merged to form a sampling point sequence.

[0052] Step 154: For each sampling point in the sampling point sequence, obtain the color attribute and density attribute of each sampling point from the voxel feature space of the neural radiation field model.

[0053] The color attribute refers to the RGB three-channel brightness value output by the neural radiation field, reflecting the reflective properties of the object's surface. The density attribute refers to the scalar absorptivity of light at a spatial point, which determines the object's opacity and light penetration ability.

[0054] In an embodiment of the present application, the three-dimensional coordinates of the sampling point in the neural radiation field voxel feature space are first located, and then the color vector and light absorption rate density attributes stored in the coordinates are queried, and finally the complete optical property data of each sampling point is obtained.

[0055] Step 155: Integrate and calculate the color attribute and density attribute of each sampling point along the ray direction, and accumulate to generate the final color value of each pixel.

[0056] Integration calculation is the process of accumulating the colors of sample points along a ray path according to transmittance weights, simulating the physical laws of light propagation. The final color value is the result of the integration calculation of a single ray, representing the color of the corresponding pixel on the virtual imaging plane.

[0057] In an embodiment of the present application, the sampling point sequence is first sorted by spatial distance along the ray direction, and then the density attributes of adjacent sampling points are converted into transmittance weights. Then, the transmittance-weighted cumulative integral calculation is performed on the color attribute to finally generate the final color value of the pixel corresponding to the current ray.

[0058] Step 156: Aggregate the final color values ​​of all pixels to form an arbitrary perspective view of the outdoor scene.

[0059] Aggregation refers to the operation of reorganizing the final color values ​​of all pixels into a two-dimensional image according to the imaging plane grid topology.

[0060] In the embodiment of the present application, the pixel color values ​​of all rays are first calculated, and then the color data is aggregated according to the pixel arrangement order of the imaging plane, and finally a complete visual image of the outdoor scene that meets the target viewing angle is generated.

[0061] The following is a specific example: First, a swarm of drones equipped with wide-angle lenses is used to collect multi-angle images of urban roads. Then, a calibration plate is used to synchronously calibrate multiple cameras to correct lens distortion and shooting time differences to generate geometrically aligned corrected images. Secondly, an image recognition model is used to annotate the object categories of the corrected images pixel by pixel, and semantic labels are added to elements such as vehicles, pedestrians, and traffic signs to form structured image data. Then, in a computing cluster consisting of four graphics processor servers, the two-dimensional semantic labels are fused with the three-dimensional model of the neural radiation field to construct a semantic point cloud containing spatial coordinates and object attributes. Non-critical areas are then identified based on the frequency of object appearance in the point cloud. For example, for billboard areas that appear less frequently in the distance, the low-density three-dimensional pixels are removed to form a streamlined point cloud. Finally, the virtual camera parameters are set according to traffic monitoring requirements: the camera center is located 100 meters above the ground directly above the road, with the lens pointing vertically downward, and imaging parameters using 1920×1080 resolution. Probe rays are emitted from the camera center to each pixel on the road surface, and the density of sampling points is automatically increased in the vehicle outline area using a simplified point cloud. After querying the color and transparency attributes of each sampling point, high-precision color accumulation is performed on the top of the vehicle, and low-precision calculations are performed on the sky area. Finally, a panoramic bird's-eye view of the traffic situation is synthesized, clearly showing the vehicle position and road markings.

[0062] By executing steps 151 to 156, the embodiment of the present application uses an adaptive sampling mechanism driven by semantic point clouds to dynamically encrypt sampling in geometrically complex areas to improve reconstruction accuracy, and sparsely sample in flat areas to reduce computational load; combines semantic boundary verification to eliminate object contour artifacts and achieve high-fidelity rendering of key areas; and finally generates high-quality views from any perspective through ray integration and pixel aggregation, meeting real-time requirements while ensuring visual realism.

[0063] In a possible embodiment, step 153, for each ray in the ray set, using the cropped semantic point cloud to drive adaptive sampling along the ray direction to generate a sampling point sequence, includes: Step a1: Based on the spatial distribution of the cropped semantic point cloud, the ray is divided into multiple initial sampling intervals along the ray direction, and an initial sampling point set is generated in each initial sampling interval according to a uniform step size.

[0064] Spatial distribution refers to the coordinate distribution characteristics of the cropped semantic point cloud within the 3D scene, reflecting the clustering and geometric complexity of objects. The initial sampling interval refers to the equal-length subsegments along the ray direction, used to initialize the uniform sampling framework. A uniform step size refers to the fixed sampling interval within the initial sampling interval to ensure basic sampling coverage. The initial sampling point set refers to the initial 3D point set generated using a uniform step size, which constitutes the basic sampling framework.

[0065] In an embodiment of the present application, the complexity of the scene passed by the ray is first analyzed based on the spatial distribution characteristics of the cropped semantic point cloud, and then the initial sampling intervals are divided into equidistant intervals along the ray direction, and then the initial sampling point set is generated in each interval with a fixed uniform step size.

[0066] Step a2: In the neighborhood of the initial sampling point set, a dynamic sampling step is calculated based on the voxel density distribution of the cropped semantic point cloud and the spatial continuity of the semantic labels, and candidate sampling points are generated around the initial sampling point according to the dynamic sampling step.

[0067] Voxel density distribution refers to the rate of change in the number of voxels per unit volume within a neighborhood, reflecting the richness of local geometric features. Spatial continuity of semantic labels refers to the degree of spatial clustering of voxels with the same semantic label within a neighborhood. Dynamic sampling step size refers to the sampling interval dynamically calculated based on density distribution and semantic continuity, with the step size automatically reduced in feature-rich areas. Candidate sampling points are 3D sampling points generated in the neighborhood of the initial point using the dynamic step size.

[0068] In an embodiment of the present application, a spherical neighborhood space is first delineated with the initial sampling point as the center. Then, the voxel density distribution gradient of the semantic point cloud and the spatial continuity index of the semantic label in the neighborhood are calculated. Then, the dynamic sampling step size is derived based on the inverse relationship between the gradient and the continuity index. Finally, candidate sampling points are generated around the initial sampling point according to the dynamic step size.

[0069] Step a3: Perform semantic label consistency check on the candidate sampling points. If the similarity between the predicted semantic label of the candidate sampling point in the voxel feature space of the neural radiation field and the label of the semantic point cloud closest to the initial sampling point is less than a preset similarity, a label difference region is generated by extending the distance in the ray direction with the candidate sampling point as the center. Otherwise, if the similarity is greater than the preset similarity, the candidate sampling point is retained.

[0070] Semantic label consistency verification is a verification process that compares the semantic label similarity of a candidate point with that of its neighboring initial points. The predicted semantic label is the semantic category identifier of the sampling point inferred by the neural radiance field model. The closest initial sampling point is the initial sampling frame reference point with the closest Euclidean distance. Similarity is the cosine similarity measure between the semantic label vectors of two sampling points. The preset similarity is a discrimination threshold set based on scene complexity.

[0071] In an embodiment of the present application, the semantic label of the candidate sampling point is first predicted by the neural radiation field model, and then the similarity between the label and the label of the nearest initial sampling point is calculated. Then, when the similarity is lower than a preset similarity, a label difference area is generated by extending a preset distance along a ray with the candidate point as the center. When the similarity reaches or exceeds a threshold, the candidate sampling point is retained.

[0072] Step a4: insert encrypted sampling points in the ray direction of the label difference area.

[0073] Among them, the label difference area refers to the ray sub-segment with obvious semantic mutation, which is generated by expanding the candidate points whose similarity is lower than the threshold.

[0074] In the embodiment of the present application, the start and end boundaries of the label difference area are first located, and then encrypted sampling points with a density higher than the initial sampling are inserted within the boundary along the ray direction, and finally a refined sampling cluster for the semantic mutation area is formed.

[0075] Step a5: Merge the initial sampling point set, the retained candidate sampling points, and the encrypted sampling points to generate a sampling point sequence along the ray direction.

[0076] The preset distance refers to a fixed length parameter for extending the label difference region along the ray direction. The encrypted sampling points refer to high-density three-dimensional sampling points inserted in the label difference region.

[0077] In the embodiment of the present application, the initial sampling point set is first merged with the retained candidate sampling points, and then the encrypted sampling points are inserted into the merged sequence in spatial order, and finally a final sampling point sequence arranged in order along the ray direction is generated.

[0078] The following is a specific example: First, a swarm of drones equipped with wide-angle lenses collects multi-angle images of urban roads, and uses a calibration plate to align the camera positions and shooting time differences to generate geometrically aligned corrected images. Secondly, an image recognition model is used to annotate the object categories of the corrected images pixel by pixel, adding semantic labels to elements such as vehicles, pedestrians, and traffic signs. Then, in a computing cluster consisting of four graphics processor servers, the two-dimensional semantic labels are fused with the three-dimensional scene model to construct point cloud data containing spatial coordinates and object attributes. Subsequently, areas with low frequency of occurrence in the point cloud are identified, such as billboard areas that appear less frequently in the distance, and their low-density three-dimensional pixels are removed to form a streamlined point cloud. Finally, the virtual camera's bird's-eye view parameters are set. The camera center is located 100 meters above the ground directly above the road. This height is calculated based on the relationship between the road width and the lens's vertical field of view. The formula is that the monitoring height is equal to the road width divided by the tangent of half the lens's vertical field of view twice. The lens is vertically downward, and high-definition resolution imaging parameters are used. A detection ray is emitted from the camera center to the road surface. The initial sampling interval is divided along the ray based on the spatial distribution of the streamlined point cloud, and initial sampling points are generated at a fixed interval. When high-density point cloud distribution and object category mutation features are detected in the vehicle edge area, the sampling interval is automatically shortened to generate candidate points. When the category label difference between the vehicle candidate point and the adjacent road initial point exceeds the set threshold, encrypted sampling points are inserted in the difference area. The initial point, candidate point, and encrypted point are merged to form a sampling sequence. After querying the color and transparency attributes of each point, the vehicle contour area is refined and color accumulation is performed, and the road area is conventionally calculated. Finally, a panoramic bird's-eye view of the traffic condition is synthesized, clearly showing the sharp transition between the vehicle edge and the road marking.

[0079] By executing steps a1 to a5, the embodiment of the present application uses a dynamic sampling mechanism driven by semantic point clouds to automatically encrypt sampling in areas with complex geometric features to improve reconstruction accuracy, and sparsely sample in semantically smooth areas to reduce computational load; combines semantic boundary verification to accurately identify areas where object contours suddenly change, and inserts encrypted sampling points in a targeted manner to eliminate boundary artifacts; ultimately, it achieves intelligent allocation of computing resources, optimizing rendering efficiency while ensuring visual quality in key areas.

[0080] In a possible embodiment, S14, identifying non-critical areas in the multi-view image based on the semantic point cloud, performing redundant voxel cropping on the non-critical areas, and generating a cropped semantic point cloud, includes: Step 141 : Based on the object categories in the semantic labels, the occurrence frequency of each object category in the spatial region is counted, and the region where the occurrence frequency of all physical categories is lower than a preset frequency threshold is used as a marked region.

[0081] Frequency of occurrence refers to the proportion of semantic labels for a specific object category within a unit spatial area, reflecting the importance of that category in the local space. A spatial region is a cubic grid unit that divides a three-dimensional scene into fixed-size units and is used for local statistical analysis. The preset frequency threshold is the critical frequency value used to determine whether a region is of low importance and needs to be dynamically set based on the scene scale. A marked region is a spatial grid unit where the frequency of occurrence of all object categories is below the preset frequency threshold.

[0082] In an embodiment of the present application, the three-dimensional space is first divided into regular grid units as spatial regions, and then the number of occurrences of semantic labels of different object categories in each grid is counted and converted into frequency values, and then the grids whose frequencies of all categories are lower than a preset frequency threshold are marked as marked areas.

[0083] Step 142: Merge adjacent marked areas that meet a preset continuity condition to generate a non-critical area. The preset continuity condition includes: the comprehensive similarity of adjacent marked areas exceeds a preset similarity threshold.

[0084] The preset continuity condition refers to the geometric and semantic association rules that must be met to merge adjacent labeled regions. Adjacent labeled regions are defined as a set of labeled grid cells that are spatially adjacent and share a common boundary. Comprehensive similarity is a weighted evaluation metric that integrates semantic composition similarity, boundary geometry similarity, and density distribution similarity. The preset similarity threshold is the critical value of comprehensive similarity that triggers the region merging operation.

[0085] In an embodiment of the present application, spatially adjacent marked areas are first detected, and then the semantic composition similarity, boundary shape similarity and density distribution similarity between adjacent areas are calculated. Subsequently, the weighted summation is performed to obtain the comprehensive similarity, and finally, the adjacent marked areas whose comprehensive similarity exceeds the preset similarity threshold are merged into continuous non-critical areas.

[0086] Step 143: Locate the voxel set corresponding to the non-critical area in the semantic point cloud, and remove the voxels with density values ​​lower than the density threshold.

[0087] The voxel set refers to all three-dimensional rasterized cells belonging to the same non-critical region. The density value is the light absorptivity scalar stored in the voxel cell, reflecting the probability of the presence of a geometric entity in a spatial location. The density threshold is the density critical value that distinguishes valid voxels from invalid voxels.

[0088] In an embodiment of the present application, first, all voxel sets covered by non-critical areas are located in the semantic point cloud, then the density value of each voxel is read, and then voxels with density values ​​lower than the density threshold are removed, retaining high-density voxels.

[0089] Step 144 : Recombining the retained voxels with the voxels in the semantic point cloud that do not belong to the non-critical area to generate a cropped semantic point cloud.

[0090] In the embodiment of the present application, first, all voxels in the semantic point cloud that are not marked as non-critical areas are extracted, and then the retained high-density voxels are recombined with the voxels outside the non-critical areas according to the spatial coordinates, and finally a cropped semantic point cloud with a complete structure is generated.

[0091] The following is a specific example: First, a swarm of drones equipped with wide-angle lenses collects multi-angle images of urban roads. Calibration plates are used to align the camera positions and capture times to generate rectified images. Next, an image recognition model is used to annotate the rectified images pixel by pixel with object categories, adding semantic labels for elements such as vehicles, pedestrians, and traffic signs. Then, in a computing cluster consisting of four GPU servers, the 2D semantic labels are fused with a 3D scene model to construct a labeled 3D point set. The three-dimensional space is then divided into cubic grid units with a side length of five meters. This size is determined based on the minimum building size on the road, and the calculation formula is that the grid side length is equal to one-tenth of the typical building width; the occurrence ratio of categories such as vehicles and pedestrians in each grid is counted, and the grid is marked when the ratio of all categories is less than three percent; adjacent marked grids are detected, and the similarity of their building surface features and the boundary shape matching are calculated. The weighted summation formula is that the comprehensive similarity is equal to the surface similarity multiplied by 0.6 plus the boundary matching multiplied by 0.4. When the result exceeds 0.7, it is merged into a non-critical area; the three-dimensional pixel point set corresponding to the non-critical area is located, and the billboard bracket pixels with a density value less than 0.1 point per cubic meter are removed. The density threshold is set according to the air density baseline value; finally, the retained high-density wall pixels are recombined with the road area pixels to form a streamlined three-dimensional point set. Finally, the virtual camera parameters are set, with the camera center located 35 meters above the ground directly above the road. This height is calculated based on the road width of 30 meters and the vertical viewing angle of the lens of 45 degrees. The calculation formula is that the monitoring height is equal to the road width divided by twice the tangent function value multiplied by half the vertical viewing angle of the lens; the streamlined point set is sampled and rendered along the pixel ray to generate a panoramic bird's-eye view of the traffic conditions, clearly showing vehicles and road markings.

[0092] By executing steps 141 to 144, the embodiment of the present application identifies low-value areas through semantic frequency statistics, combines geometric and semantic continuity to expand the cropping range; accurately removes invalid voxels through density screening, and compresses the model size while retaining key structures; and finally achieves targeted optimization of computing resources, providing a lightweight data foundation for real-time rendering.

[0093] In a possible embodiment, step 142, merging adjacent marked areas that meet a preset continuity condition to generate a non-critical area, includes: Step b1: abstract each marked area into a graph node. If the Euclidean distance between two adjacent marked areas is less than a preset spatial distance threshold, an undirected edge is established between the graph nodes to form an adjacency graph.

[0094] A graph node is an abstract topological unit representing a single labeled region and is used to construct a regional adjacency network. Euclidean distance, the straight-line distance between two points in three-dimensional space, is used to quantify the spatial proximity of labeled regions. The preset spatial distance threshold is the maximum spatial separation threshold used to determine whether regions are adjacent. An undirected edge is a bidirectional connection line connecting two graph nodes, representing spatial bordering relationships. An adjacency graph is a topological network composed of nodes and undirected edges that expresses the spatial adjacency of labeled regions.

[0095] In an embodiment of the present application, each marked area is first abstracted as an independent node in the graph structure, and then the straight-line distance between the center points of the spatial areas corresponding to any two nodes is detected. Then, when the distance is less than a preset spatial distance threshold, a bidirectional connection edge is established between the corresponding nodes, and finally a connection graph expressing the adjacency relationship of the areas is formed.

[0096] Step b2: Calculate the semantic label similarity, boundary geometry similarity, and density distribution similarity for each undirected edge in the adjacency graph, and perform weighted processing on the semantic label similarity, boundary geometry similarity, and density distribution similarity to obtain the corresponding comprehensive similarity.

[0097] Semantic label similarity measures the cosine similarity of the semantic category distribution between two regions. Boundary geometry similarity measures the Hausdorff distance of the polygonal shape matching between the two regions. Density distribution similarity measures the Pearson correlation coefficient of the point cloud density trends between the two regions. Weighted processing involves the linear combination of multiple metrics using pre-set weights.

[0098] In an embodiment of the present application, first, the semantic label composition similarity, boundary contour geometric similarity and point cloud density distribution similarity of the two end regions of each connecting edge in the connection graph are calculated, and then the three types of similarity indicators are weighted and summed according to the preset weight coefficients, and finally a comprehensive similarity value representing the overall correlation between regions is obtained.

[0099] Step b3: iteratively merge adjacent marked regions whose comprehensive similarity is greater than a preset similarity threshold to update the adjacency graph until the comprehensive similarity between all adjacent marked regions in the adjacency graph is lower than the preset similarity threshold, thereby obtaining candidate regions.

[0100] Iterative merging refers to the optimization process of cyclically executing region fusion and graph structure update. Candidate regions refer to the set of regions to be screened after similarity-driven merging.

[0101] In an embodiment of the present application, the connection graph is first traversed to screen connection edges whose comprehensive similarity exceeds a preset threshold. Secondly, the marked areas represented by the corresponding nodes are merged and the node attributes are updated. Then, the connection graph structure is rebuilt and the similarity is recalculated. This process is repeated until no connection edges exceeding the threshold exist, and finally a merged candidate area set is obtained.

[0102] Step b4: Calculate the semantic entropy of the candidate regions, and select the candidate regions whose semantic entropy is less than a preset semantic entropy threshold as valid merged regions.

[0103] Semantic entropy is a measure of the uncertainty in the distribution of semantic labels within a region. Lower entropy values ​​indicate a more homogeneous classification. The preset semantic entropy threshold is the critical entropy value for determining semantic consistency. Valid merge regions are defined as a subset of candidate regions whose semantic entropy is below the threshold and that meet the merge requirements.

[0104] In an embodiment of the present application, the semantic entropy value of the candidate area is first calculated, which reflects the degree of confusion of the semantic categories in the area. Secondly, the candidate area with an entropy value lower than a preset semantic entropy threshold is determined as a valid merged area that meets the semantic consistency standard.

[0105] Step b5: Perform screening processing on the valid merged area to generate a non-critical area, the screening processing including at least one of the following: area threshold screening, aspect ratio constraint, and length-width ratio constraint.

[0106] Screening refers to the secondary selection of regions based on geometric rules. Area threshold screening removes invalid small regions with an area smaller than a preset value. Aspect ratio constraint excludes irregular regions with an aspect ratio exceeding a preset limit.

[0107] In an embodiment of the present application, first, the effective merged areas are screened by area to remove areas that are too small, and then narrow and irregular areas are excluded according to the aspect ratio constraint, and finally a set of non-critical areas that meet the spatial regularity requirements is generated.

[0108] The following is a specific example: First, a swarm of drones equipped with wide-angle lenses captures multi-angle images of a city's main roads. Calibration plates are used to align the camera positions and shooting timing to generate calibrated images. Next, an image recognition model is used to annotate the calibrated images pixel by pixel with object categories, adding semantic labels to elements such as vehicles, streetlights, and billboards. Then, in a computing cluster consisting of four GPU servers, the 2D semantic labels are integrated with the 3D scene model to construct a 3D point set with category labels. Distant billboard areas with low frequency are then marked as processing units. Each billboard area is treated as an independent node. A network diagram is formed by connecting two billboard centers within a straight-line distance of less than 10 meters. This distance threshold is derived from statistics on the minimum distance between billboards on urban streets. The similarity of advertising category composition, outer contour shape matching, and point cloud density correlation in the connected areas are then calculated. The combined weighted sum is calculated in a 5:3:2 ratio to produce a comprehensive similarity value, with the weights assigned based on the importance of the billboard's structural features. Next, adjacent regions with a comprehensive similarity exceeding 0.85 are merged in a loop and the network graph is updated. This threshold is optimized by testing the merging effect in different scenarios until no excessive connections exist. A mathematical formula is then used to calculate the semantic confusion index of the merged regions, retaining regions with an index value below 1.2. This threshold is set based on the requirement that the index meets the requirement when the proportion of advertising categories within the region exceeds 85%. Finally, regions with an area smaller than five square meters and an aspect ratio exceeding three to one are removed. This size limit is determined by the minimum building facade size. Non-critical regions are generated and their low-density support points are removed. Finally, the virtual camera parameters are set, with the camera center located at a height of 50 meters above the ground directly above the road. This height is calculated by taking the road width of 40 meters and the vertical angle of view of the lens at 60 degrees. The calculation formula is: the monitoring height is equal to the road width divided by twice the tangent function value multiplied by half the vertical angle of view of the lens. The reduced point set is sampled and rendered along pixel rays, generating a panoramic bird's-eye view of the road, clearly showing the positional relationship between vehicles and billboards.

[0109] By executing steps b1 to b5, the embodiment of the present application accurately expresses regional spatial relationships through topological modeling, and drives progressive region merging in combination with multi-dimensional similarity evaluation; uses semantic entropy quantization to screen semantically consistent regions to avoid loss of effective information caused by excessive merging; and finally generates structurally regular non-critical regions through geometric rule constraints, laying the foundation for refined voxel cropping.

[0110] In a possible embodiment, S13, updating the neural radiation field model based on the image data with semantic labels, and mapping the semantic labels to the voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud, includes: Step 131: Extract the geometric features and semantic features of each voxel from the image data with semantic labels.

[0111] Geometric features are mathematical vectors that describe the spatial structural properties of voxels, including parameters that reflect three-dimensional shape, such as center point coordinates, boundary dimensions, and surface curvature. Semantic features are discrete or continuous vectors that represent voxel category attributes, such as the probability distribution of object categories and material attribute encoding, as well as other high-level semantic information.

[0112] In an embodiment of the present application, each three-dimensional voxel unit is first located from the image data with semantic labels, and then the geometric feature vector describing the spatial position and shape is extracted, and the semantic feature vector representing the object category attribute is extracted at the same time, and finally the bimodal feature data of all voxels are obtained.

[0113] Step 132: Update the parameters of the neural radiation field model by jointly optimizing the geometric features and the semantic features.

[0114] Among them, joint optimization refers to a training strategy that simultaneously minimizes the geometric reconstruction error and semantic classification error, and realizes cross-modal knowledge transfer by sharing the underlying network features.

[0115] In an embodiment of the present application, a geometric feature reconstruction error function and a semantic feature classification error function are first constructed, and then the two errors are jointly optimized through a multi-task learning framework. Subsequently, the gradient backpropagation algorithm is used to update the weight parameters of the fully connected layer of the neural radiation field model, and finally the collaborative expression optimization of geometry and semantics is achieved.

[0116] Step 133: Map the semantic labels to corresponding voxel positions in the voxel feature space of the neural radiation field model.

[0117] The voxel position refers to the integer index value of the voxel unit in the three-dimensional grid coordinate system, which is used to locate the storage position in the feature space.

[0118] In an embodiment of the present application, the three-dimensional index coordinate system of the neural radiation field voxel feature space is first determined, and then the semantic labels in the two-dimensional image data are associated with the corresponding voxel index positions through projection transformation, and finally the precise mapping of the semantic labels in the three-dimensional space is completed.

[0119] Step 134: Fuse the spatial coordinates of the voxel position with the semantic label to construct a semantic point cloud.

[0120] The spatial coordinates refer to the floating-point physical coordinate values ​​of the voxel center point in the three-dimensional world coordinate system, which are used to describe the absolute spatial position.

[0121] In an embodiment of the present application, the spatial position coordinates of the voxel in the three-dimensional world coordinate system are first read, and then the coordinates are feature-joined with the mapped semantic labels to finally generate a semantic point cloud dataset that contains both spatial position and category attributes.

[0122] The following is a specific example: First, a swarm of drones equipped with wide-angle lenses captures multi-angle images of an urban intersection. A calibration plate is used to align the camera lens positions and capture time differences to generate a geometrically corrected image. Next, an image recognition model is used to annotate the corrected images pixel by pixel with object categories, assigning category labels to traffic elements such as vehicles, pedestrians, and traffic lights. Next, a computing cluster consisting of four GPU servers extracts 3D dimensional features and vehicle type attributes from the labeled images. A joint training approach simultaneously optimizes the accuracy of vehicle shape reconstruction and type recognition, updating the internal parameters of the 3D scene model. The vehicle labels in the 2D images are then converted to 3D grid cell indexes using a spatial mapping process. This mapping is based on camera imaging principles, specifically calculating the correspondence between 2D pixels and 3D space using the camera's internal and external parameter matrices. The latitude, longitude, and altitude coordinates of the 3D grid cell center are then combined with the vehicle type label to construct a vehicle point set containing both spatial location and semantic attributes. Then, distant billboard areas with a point concentration frequency of less than 3% are marked as non-critical areas. This frequency threshold is determined based on the distribution statistics of the main objects in the road scene. Low-density areas of billboard supports are identified, and grid cells with a density value less than 0.05 points per cubic meter are removed. This density threshold is set based on the air density baseline value. Finally, the virtual camera parameters are set, with the camera center located 45 meters above the ground directly above the intersection. This height is calculated based on the maximum road width of 50 meters and the vertical angle of view of the lens of 50 degrees. The calculation formula is that the monitoring height is equal to the maximum road width divided by twice the tangent function value multiplied by half the vertical angle of view of the lens. Ray tracing calculations are performed on the simplified point set along all pixel directions. First, basic sampling points are generated. The sampling density is automatically increased in the vehicle outline area based on the point set distribution, and the number of sampling points is reduced in the sky area. After querying the color and transparency properties of each sampling point, the vehicle surface is refined and color accumulation is performed. Conventional calculations are performed on the background area. Finally, a panoramic bird's-eye view of the intersection is synthesized, clearly showing the vehicle position and traffic light status.

[0123] By executing steps 131 to 134, the embodiment of the present application enhances the consistency of three-dimensional expression through the joint optimization of geometric and semantic features, and uses multi-task learning to improve the model's modeling ability for complex scenes; establishes cross-modal associations through precise three-dimensional mapping of semantic labels, and constructs a point cloud model with both spatial accuracy and semantic understanding capabilities, providing a structured data foundation for subsequent scene editing and analysis.

[0124] In a possible embodiment, step 132, updating the parameters of the neural radiation field model by jointly optimizing geometric features and semantic features, includes: Step c1: Calculate the matching error between the geometric features and the corresponding three-dimensional geometric representation in the neural radiation field model to generate a geometric loss.

[0125] The 3D geometric representation refers to the spatial density field function output by the neural radiance field model, which is used to implicitly mathematically describe the 3D shape of the scene. The matching error measures the positional deviation between the geometric shape predicted by the model and the spatial structure of the actual scene. The geometric loss is a numerical function constructed based on the matching error, reflecting the goal of optimizing 3D reconstruction accuracy.

[0126] In an embodiment of the present application, the three-dimensional geometric expression data predicted by the neural radiation field model is first obtained, then the spatial position difference value between the data and the geometric features extracted from the image is calculated, and finally the difference value is quantified as the output of the geometric loss function.

[0127] Step c2: Calculate the consistency error between the semantic features and the semantic labels to generate the semantic loss.

[0128] The consistency error refers to the degree of information discrepancy between the semantic category predicted by the model and the true label. The semantic loss is a loss function term that quantifies the deviation in semantic prediction and is used to improve category recognition accuracy.

[0129] In an embodiment of the present application, the probability distribution of the semantic feature vector and the true semantic label is first compared, and then the category recognition deviation value is calculated through the cross entropy algorithm, and finally a semantic loss function output reflecting the accuracy of semantic prediction is generated.

[0130] Step c3: Jointly optimize the geometric loss and semantic loss to generate a joint optimization loss.

[0131] Among them, the joint optimization loss refers to a composite objective function that integrates geometric loss and semantic loss to achieve multi-task collaborative training.

[0132] In an embodiment of the present application, the weight coefficients of the geometric loss and the semantic loss are first set, and then the two weighted losses are added together to finally form a joint loss function that drives model optimization.

[0133] Step c4: Update the parameters of the neural radiation field model based on the joint optimization loss.

[0134] Among them, in the embodiment of the present application, the gradient of the joint loss function to the model parameters is first calculated, and then the adaptive learning rate optimization algorithm is used to update the fully connected layer weights along the negative gradient direction, and finally the iterative optimization of the neural radiation field model parameters is achieved.

[0135] The following is a specific example: First, a swarm of drones equipped with wide-angle lenses captures multi-angle images of a city intersection. Calibration plates are used to align the camera lens positions and capture time differences to generate geometrically corrected images. Next, an image recognition model is used to annotate the corrected images pixel by pixel with object categories, adding type labels to traffic elements such as vehicles, pedestrians, and traffic lights. Next, a computing cluster consisting of four GPU servers performs the following operations: Laser scanning equipment is used to obtain precise vehicle shape data as a benchmark. The spatial positional difference between the vehicle body curve predicted by the 3D scene model and the actual vehicle shape is calculated to generate a shape error. The vehicle type prediction output by the model is compared with the type difference from the manually annotated labels to generate a type error. The shape error is multiplied by 0.7 and the type error is multiplied by 0.3 to obtain a comprehensive optimization target value. This weighting ratio has been tested repeatedly and found to balance the requirements of shape accuracy and type recognition. Based on this target value, the model's internal parameter adjustment algorithm is used to optimize the 3D scene model. Subsequently, the vehicle labels in the 2D image are converted to 3D grid cell locations through spatial mapping. A semantic point set is constructed by combining the grid center's latitude and longitude, altitude, and vehicle type. Next, the statistical point concentration frequency in the distant billboard area is less than 5%. This threshold is set based on the minimum frequency of major traffic elements. Support structure points in the billboard area with a density value less than 0.1 point per cubic meter are removed. This threshold is determined based on air density measurement data. Finally, the virtual camera parameters are set. The camera center is located 50 meters above the ground directly above the intersection. This height is calculated based on the maximum road width of 60 meters and the vertical angle of view of the lens of 55 degrees. The calculation formula is that the monitoring height is equal to the maximum road width divided by twice the tangent function value multiplied by half the vertical angle of view of the lens. Ray calculations are performed on the simplified point set along all pixel directions, and the sampling point density is automatically increased in the vehicle outline area. After querying the color and transparency attributes of each point, the vehicle surface is refined and accumulated. Finally, a panoramic bird's-eye view of the intersection is generated, clearly showing the color and outline details of the vehicle and the status of the traffic light.

[0136] By executing steps c1 to c4, the embodiment of the present application balances the training objectives of geometric reconstruction and semantic recognition through a joint optimization strategy, and uses gradient updates to simultaneously improve spatial accuracy and semantic consistency; the constructed composite loss function effectively coordinates the multi-task learning process and enhances the model's ability to express complex outdoor scenes.

[0137] Figure 2 A schematic diagram of the structure of a visual environment generation system based on neural radiation field provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the system includes: The acquisition module 21 is used to acquire multi-view images of an outdoor scene captured by a wide-angle lens, perform multi-camera synchronous calibration on the multi-view images, and generate a corrected image.

[0138] The generation module 22 is used to perform semantic segmentation on the corrected image to generate image data with semantic labels, where the semantic labels are used to mark the categories of objects in the outdoor scene.

[0139] The construction module 23 is used to update the neural radiation field model based on the image data with semantic labels in the cascade rendering computer cluster, and map the semantic labels to the voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud.

[0140] The recognition module 24 is configured to recognize non-critical areas in the multi-view image based on the semantic point cloud, perform redundant voxel cropping on the non-critical areas, and generate a cropped semantic point cloud.

[0141] The driving module 25 is used to drive the neural radiation field model using the cropped semantic point cloud to generate an arbitrary perspective view of the outdoor scene.

[0142] Figure 2 The visual environment generation system based on neural radiation field can be executed Figure 1 The implementation principles and technical effects of the neural radiation field-based visual environment generation method described in the illustrated embodiment will not be elaborated on here. The specific manner in which the various modules and units in the neural radiation field-based visual environment generation system in the above embodiment perform operations has been described in detail in the relevant embodiments of the method and will not be elaborated on here.

[0143] In one possible design, Figure 2 The visual environment generation system based on neural radiation field of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32 .

[0144] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0145] The processing component 32 is used to: obtain multi-perspective images of outdoor scenes captured by a wide-angle lens, perform multi-camera synchronous calibration on the multi-perspective images, and generate corrected images. Perform semantic segmentation on the corrected images to generate image data with semantic labels, and the semantic labels are used to mark the categories of objects in the outdoor scene. In the cascade rendering computer cluster, the neural radiation field model is updated based on the image data with semantic labels, and the semantic labels are mapped to the voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud. Based on the semantic point cloud, non-critical areas in the multi-perspective images are identified, and redundant voxel cropping is performed on the non-critical areas to generate a cropped semantic point cloud. The cropped semantic point cloud is used to drive the neural radiation field model to generate an arbitrary perspective view of the outdoor scene.

[0146] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.

[0147] The storage component 31 is configured to store various types of data to support operations on the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as random access memory (RAM), static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0148] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0149] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0150] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0151] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0152] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment provides a method for generating a visual environment based on a neural radiation field.

[0153] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0154] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0155] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating a visual environment based on a neural radiation field, characterized in that: include: Acquire multi-view images of an outdoor scene captured by a wide-angle lens, perform multi-camera synchronous calibration on the multi-view images, and generate corrected images; Performing semantic segmentation on the corrected image to generate image data with semantic labels, wherein the semantic labels are used to mark the categories of objects in the outdoor scene; In a cascaded rendering computer cluster, a neural radiation field model is updated based on the image data with the semantic labels, and the semantic labels are mapped to a voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud; Based on the semantic point cloud, identifying non-critical areas in the multi-view image, performing redundant voxel cropping on the non-critical areas, and generating a cropped semantic point cloud; The cropped semantic point cloud is used to drive the neural radiance field model to generate an arbitrary perspective view of the outdoor scene.

2. The method for generating a visual environment based on a neural radiation field according to claim 1, characterized in that: The method of using the cropped semantic point cloud to drive the neural radiation field model to generate an arbitrary perspective view of an outdoor scene includes: Defining virtual camera parameters of the target viewing angle, wherein the virtual camera parameters include the virtual camera center position, posture and intrinsic parameter matrix; Based on the virtual camera parameters, starting from the center position of the virtual camera, emitting rays along each pixel direction to form a ray set; For each ray in the ray set, using the cropped semantic point cloud to drive adaptive sampling along the ray direction to generate a sampling point sequence; For each sampling point in the sampling point sequence, obtaining a color attribute and a density attribute of each sampling point from a voxel feature space of the neural radiation field model; Integrating the color attribute and density attribute of each sampling point along the ray direction to generate a final color value of each pixel; Aggregating the final color values ​​of all the pixels to form an arbitrary viewing angle view of the outdoor scene.

3. The method for generating a visual environment based on a neural radiation field according to claim 2, characterized in that: The method of driving adaptive sampling along the ray direction using the cropped semantic point cloud for each ray in the ray set to generate a sampling point sequence includes: Based on the spatial distribution of the cropped semantic point cloud, the ray is divided into a plurality of initial sampling intervals along the ray direction, and an initial sampling point set is generated in each initial sampling interval according to a uniform step size; In a neighborhood of the initial sampling point set, a dynamic sampling step is calculated based on the voxel density distribution of the cropped semantic point cloud and the spatial continuity of the semantic labels, and candidate sampling points are generated around the initial sampling point according to the dynamic sampling step; Performing semantic label consistency check on the candidate sampling point. When the similarity between the predicted semantic label of the candidate sampling point in the voxel feature space of the neural radiation field and the label of the semantic point cloud closest to the initial sampling point is less than a preset similarity, expanding the preset distance in the ray direction with the candidate sampling point as the center to generate a label difference area. Alternatively, when the similarity is greater than the preset similarity, retaining the candidate sampling point. Inserting encrypted sampling points in the ray direction of the label difference area; The initial sampling point set, the retained candidate sampling points and the encrypted sampling points are merged to generate a sampling point sequence along the ray direction.

4. The method for generating a visual environment based on a neural radiation field according to claim 1, characterized in that: The method of identifying a non-critical area in the multi-view image based on the semantic point cloud, performing redundant voxel cropping on the non-critical area, and generating a cropped semantic point cloud includes: Based on the object categories in the semantic labels, counting the occurrence frequencies of each object category in the spatial region, and marking regions where the occurrence frequencies of all physical categories are lower than a preset frequency threshold; Merging adjacent marked areas that meet a preset continuity condition to generate a non-critical area, wherein the preset continuity condition includes: the comprehensive similarity of the adjacent marked areas exceeds a preset similarity threshold; Locating a voxel set corresponding to the non-critical area in the semantic point cloud, and removing voxels with density values ​​lower than a density threshold; The retained voxels are recombined with the voxels in the semantic point cloud that do not belong to the non-critical area to generate a cropped semantic point cloud.

5. The method for generating a visual environment based on a neural radiation field according to claim 4, characterized in that: The step of merging adjacent marked areas that meet a preset continuity condition to generate a non-critical area includes: Each marked area is abstracted as a graph node. If the Euclidean distance between two adjacent marked areas is less than the preset spatial distance threshold, an undirected edge is established between the graph nodes to form an adjacency graph. Calculating semantic label similarity, boundary geometric similarity, and density distribution similarity for each undirected edge in the adjacency graph, and weighting the semantic label similarity, boundary geometric similarity, and density distribution similarity to obtain a corresponding comprehensive similarity; Iteratively merging the adjacent marked regions whose comprehensive similarities are greater than a preset similarity threshold to update the adjacency graph until the comprehensive similarities between all adjacent marked regions in the adjacency graph are lower than the preset similarity threshold, thereby obtaining candidate regions; Calculating the semantic entropy of the candidate regions, and taking the candidate regions whose semantic entropy is less than a preset semantic entropy threshold as valid merged regions; A screening process is performed on the valid merged area to generate a non-critical area, wherein the screening process includes at least one of the following: area threshold screening, aspect ratio constraint, and length-width ratio constraint.

6. The method for generating a visual environment based on a neural radiation field according to claim 1, characterized in that: The updating of the neural radiation field model based on the image data with the semantic labels, and mapping the semantic labels to the voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud, includes: Extracting geometric features and semantic features of each voxel from the image data with semantic labels; Updating parameters of the neural radiation field model by jointly optimizing the geometric features and the semantic features; Mapping the semantic labels to corresponding voxel locations in a voxel feature space of the neural radiation field model; The spatial coordinates of the voxel positions and the semantic labels are fused to construct a semantic point cloud.

7. The method for generating a visual environment based on a neural radiation field according to claim 6, characterized in that: The updating of the parameters of the neural radiation field model by jointly optimizing the geometric features and the semantic features includes: calculating a matching error between the geometric feature and a corresponding three-dimensional geometric representation in the neural radiation field model to generate a geometric loss; Calculating the consistency error between the semantic feature and the semantic label to generate a semantic loss; Jointly optimizing the geometric loss and the semantic loss to generate a joint optimization loss; Parameters of the neural radiation field model are updated according to the joint optimization loss.

8. A visual environment generation system based on neural radiation field, characterized in that: include: An acquisition module is used to acquire multi-view images of an outdoor scene captured by a wide-angle lens, perform multi-camera synchronous calibration on the multi-view images, and generate a corrected image; a generation module, configured to perform semantic segmentation on the corrected image to generate image data with semantic labels, wherein the semantic labels are used to mark the categories of objects in the outdoor scene; A construction module is configured to update a neural radiation field model based on the image data with semantic labels in a cascade rendering computer cluster, and map the semantic labels to a voxel feature space corresponding to the neural radiation field model to construct a semantic point cloud; an identification module, configured to identify non-critical areas in the multi-view image based on the semantic point cloud, perform redundant voxel cropping on the non-critical areas, and generate a cropped semantic point cloud; A driving module is used to drive the neural radiation field model using the cropped semantic point cloud to generate an arbitrary perspective view of the outdoor scene.

9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a visual environment generation method based on a neural radiation field as described in any one of claims 1-7.

10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, the method for generating a visual environment based on a neural radiation field according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Urban large-scale scene reconstruction method based on neural radiation field

    CN115841559A

  • Intelligent computing power recommendation method under heterogeneous computing power integration system

    CN118312329A

Cited By

  • Panoramic picture generation method and device based on structured prompt, equipment and medium

    CN120852576A

  • Method, apparatus and medium for generating panoramic picture based on structured prompt

    CN120852576B

  • Unmanned aerial vehicle live-action three-dimensional model construction optimization method and system for urban planning

    CN121392195A

  • Urban planning unmanned aerial vehicle real scene three-dimensional model construction optimization method and system

    CN121392195B

  • Efficient decoupling and object removing method based on neural radiation field scene

    CN121937615A