Neural feature field based target three-dimensional construction and real scale estimation system and method

By using a neural feature field-based approach combined with multi-view image acquisition and deep learning, we have achieved fast and accurate 3D construction and real-scale estimation, solving the problems of slow speed, low accuracy and high resource consumption in existing technologies. This approach is applicable to fields such as VR, AR and autonomous driving.

CN115761179BActive Publication Date: 2026-07-31SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI UNIV
Filing Date
2022-10-17
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing target 3D modeling technologies are slow, have low accuracy, and are costly, making it difficult to quickly and accurately build 3D models in complex scenes. They also cannot effectively handle occlusions and the problem of costly repetitive construction.

Method used

By employing a neural feature field-based approach, through multi-view image acquisition, feature information extraction, standard neural feature field registration, and virtual neural feature field generation, combined with deep learning concepts, we can customize 3D feature information extraction and rendering to achieve fast and accurate 3D construction and realistic scale estimation.

Benefits of technology

It achieves fast, accurate, and low-resource-consumption 3D construction, reduces equipment requirements, and solves the problems of slow construction speed, low accuracy, and high resource consumption in traditional methods. It also has the ability to construct personalized target 3D structures and estimate true scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761179B_ABST
    Figure CN115761179B_ABST
Patent Text Reader

Abstract

This invention discloses a system and method for target 3D construction and realistic scale estimation based on neural feature fields. The system consists of units for multi-view image acquisition, feature information extraction, target recognition, standard neural feature field registration, virtual neural feature field generation, visualization rendering, and 3D construction and realistic scale estimation of the target. The method includes the following steps: multi-view image acquisition, feature information extraction, standard neural feature field registration, virtual neural feature field generation, and realistic scale estimation of the target. This invention not only significantly reduces the demand for target 3D construction equipment but also enables new view synthesis, greatly reducing construction complexity and saving resources consumed by repetitive construction. It possesses faster, more accurate, and more personalized target 3D construction capabilities and realistic scale information perception capabilities. The system of this invention has a simple structure, is easy to operate, has superior performance, can adapt to various target 3D construction tasks, has strong practical value, and a wide range of applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a system and method for target 3D construction and true scale estimation, and in particular to a system and method for target 3D construction and true scale estimation based on neural feature fields. Background Technology

[0002] The 3D target reconstruction process integrates numerous advanced technologies such as image processing, feature extraction, 3D reconstruction, and deep learning, making it a research hotspot in the field of computer vision. Target 3D reconstruction technology involves extracting the 3D feature information of a target from a single or multiple views and then visualizing and rendering the target in a virtual 3D space based on this feature information. This technology has been widely applied in many fields: in virtual reality, it can be used for 3D scene construction and metaverse building; in autonomous driving, it can be used for driving environment construction and obstacle detection; in game development, it can be used for virtual character creation, game environment construction, and rendering; and in clinical medicine, it can be used for 3D reconstruction of human organs, prosthetic modeling, and 3D printing. Furthermore, the guiding opinions on accelerating the development of the virtual reality industry point out that target 3D reconstruction technology will serve as an important supporting technology and inevitable trend for many industries, greatly enhancing my country's core technological competitiveness.

[0003] To date, target 3D construction technology has made great progress, but it remains a difficult task to complete target 3D construction quickly, accurately, and efficiently and to visualize and render the target in a virtual 3D scene. This difficulty mainly comes from three aspects: (1) The problem of complex 3D construction equipment. Extracting and saving the 3D feature information such as the shape, appearance, and texture of the construction target requires professional computing equipment. In the process of 3D feature information extraction, multiple features are often intertwined, making it impossible to extract the 3D information required for a specific construction task. In the process of target 3D construction, a large amount of storage space is consumed to save all 3D feature information; (2) The problems of slow construction speed, poor real-time performance, low accuracy, and inability to construct occlusions (partial occlusion). Target 3D construction tasks require computing equipment to spend a lot of time and computing power to process the relative positions of the target in the entire environment space, including the spatial relationship with the background, other objects, possible occlusions, shadows, etc. The disadvantages of long time and high computing power make it difficult for traditional 3D construction tasks to be widely used; (3) The problem of high construction cost. 3D construction tasks can often only be constructed once for unchanging targets or scenes. Once the target or the scene changes, the entire construction process needs to be repeated, thus consuming a lot of resources. These issues pose various challenges to target 3D construction technology. Therefore, it is necessary to conduct in-depth research on rapid target 3D construction technology in complex scenarios to improve the speed, accuracy, and real-time performance of target 3D construction, while reducing resource consumption and construction costs. Summary of the Invention

[0004] The purpose of this invention is to address the problems of slow speed, low accuracy, and high resource consumption in 3D construction tasks. It proposes a target 3D construction and real scale estimation system and method based on neural feature fields. Using this technology, it is possible to achieve active identification of the construction target, extraction of customized 3D information, matching of standard neural feature fields, and generation of virtual neural feature fields based on limited real data input. It has fast construction speed, high accuracy, low resource consumption, and wide application fields.

[0005] To solve the above problems, the concept of this invention is as follows:

[0006] This paper employs a key technology for 3D target construction based on neural feature fields to improve construction efficiency. The key technology involves first capturing a set of images of the real target from a limited viewpoint, then extracting the target's 3D feature information from these images and storing this information in a fully connected deep neural network. This generates a neural feature field that meets the target scale and construction requirements. Finally, this neural feature field is rendered to achieve 3D target construction in virtual space. Considering the numerous problems associated with using only a limited set of images of targets in real-world scenes, a virtual neural feature field generation technology is introduced. Specifically, this virtual neural feature field is formed by fusing feature information from a standard neural feature field and the real-world image set. It possesses both the standard 3D feature information for constructing this type of target and the real 3D feature information for constructing the target in this scene. By analyzing the relationship between the two types of 3D feature information—those stored in the standard neural feature field and those extracted from the real target image set—it is possible to customize the extraction, storage, and fusion of the required 3D feature information according to the needs of the construction task, generating a new virtual neural feature field and improving construction speed and efficiency. Finally, by rendering the newly generated virtual neural feature field, the specified target can be visualized quickly, accurately, and efficiently in virtual 3D space, thereby solving a series of problems in traditional target 3D construction tasks.

[0007] Based on the above inventive concept, the present invention adopts the following technical solution:

[0008] A target 3D construction and real scale estimation system based on neural feature fields is proposed. It can construct real targets in virtual space and estimate their real scale. The method in the construction process adopts key technologies of neural feature field generation, relies on multi-view acquisition and other methods to obtain real target image sets and corresponding 3D feature information, and fuses these feature information with the standard neural feature fields of the target category to generate new virtual neural feature fields. After rendering, the 3D construction and real scale estimation of the target in virtual space are realized. The system mainly consists of a multi-view image acquisition unit, a feature information extraction unit, a target recognition unit, a standard neural feature field registration unit, a virtual neural feature field generation unit, a visualization rendering and 3D construction unit, and a target true scale estimation unit. Its basic features are: the multi-view image acquisition unit and the feature information extraction unit are connected via wired or wireless means; the feature information extraction unit and the target recognition unit are connected via wired means; the target recognition unit and the standard neural feature field registration unit are connected via wired means; the standard neural feature field registration unit and the virtual neural feature field generation unit are connected via wired means; the virtual neural feature field generation unit and the visualization rendering and 3D construction unit are connected via wired means; and the visualization rendering and 3D construction unit and the target true scale estimation unit are connected via wired means.

[0009] The aforementioned multi-view image acquisition unit includes: an image acquisition module, an image preprocessing module, and a data transmission module. The image acquisition module is connected to the image preprocessing module; the image preprocessing module is connected to both the image acquisition module and the data transmission module; the data transmission module is connected to both the feature information extraction unit and the target true scale estimation unit. The image acquisition module captures multi-view image information of the real target, including a multi-view image set and the shooting pose of the acquisition nodes. The image preprocessing module corrects and restores each image in the captured multi-view image set and unifies its scale to ensure that images obtained at a single acquisition node are not affected by the camera's tilt angle, and images obtained at multiple acquisition nodes are not affected by different shooting scales. The data transmission module uploads the preprocessed multi-view image set S and other relevant information.

[0010] The aforementioned feature information extraction unit includes: a target recognition feature information extraction module and a three-dimensional feature information extraction module. The target recognition feature information extraction module is connected to the data transmission module and the target recognition unit. The three-dimensional feature information extraction module is connected to the data transmission module, the target recognition unit, and the standard neural feature field registration unit. The target recognition feature information extraction module extracts features from the uploaded multi-view image set. The image features extracted in this operation are only applicable to subsequent target detection and target recognition functions. The three-dimensional feature information extraction module extracts three-dimensional features from the uploaded multi-view image set. Specifically, based on the target category information output by the target recognition unit, it crops the region where the target is located in each real-world image from the multi-view image set S and constructs the target multi-view image set S. Z Then for S Z Three-dimensional feature information is extracted, including three-dimensional shape features and three-dimensional appearance features.

[0011] The aforementioned target recognition unit includes: a target detection module, a target recognition module, and a multi-category target scale information storage module. The target detection module is connected to the target recognition feature information extraction module and the target recognition module. The target recognition module is connected to the target detection module, the 3D feature information extraction module, and the standard neural feature field registration unit. The multi-category target scale information storage module is connected to the target scale estimation unit. The target detection module detects the targets to be constructed in the real image set transmitted in the data and obtains target category information. The target recognition module classifies the obtained target category information to obtain detailed category information of the targets to be constructed. The multi-category target scale information storage module stores scale information of multiple categories of targets, including scale information, conventional size information, etc.

[0012] The aforementioned standard neural feature field registration unit includes: a standard neural feature field storage module, a standard neural feature field loading module, and a three-dimensional feature information and standard neural feature field registration module. The standard neural feature field storage module is connected to the standard neural feature field loading module; the standard neural feature field loading module is connected to the standard neural feature field storage module, the target recognition module, and the three-dimensional feature information and standard neural feature field registration module; the standard neural feature field storage module stores standard neural feature fields for multiple categories of targets, forming a multi-category target standard neural feature field library; the standard neural feature field loading module constructs detailed category information of the target as needed based on the target recognition module, performs a corresponding search in the multi-category target standard neural feature field library, and loads the standard neural feature field of the target; the three-dimensional feature information and standard neural feature field registration module registers the three-dimensional feature information of the target to be constructed with the standard neural feature field of the target category, specifically, finding the neuron combination that stores the corresponding three-dimensional feature information in the standard neural feature field, and registering the three-dimensional feature information extracted from the real target according to the storage method of the combination, forming a new neuron combination with real three-dimensional feature information.

[0013] The aforementioned virtual neural feature field generation unit includes: a fusion module for 3D feature information and standard neural feature fields, a virtual neural feature field self-verification module, and a virtual neural feature field storage module. The fusion module is connected to both the registration module and the self-verification module. The self-verification module is also connected to both the fusion module and the storage module. The storage module is connected to both the self-verification module and the visualization rendering and 3D construction unit. The fusion module fuses the newly registered neuron combination with real 3D feature information with the corresponding part of the standard neural feature field, generating a new neural feature field called a virtual neural feature field. This virtual neural feature field possesses both the 3D feature information of the real target and other overall 3D feature information of the target category. The self-verification module verifies the fused virtual neural feature field to ensure it can complete the subsequent 3D target construction task. The storage module stores the verified virtual neural feature field of the target.

[0014] The aforementioned visualization rendering and 3D construction unit includes: a visualization rendering module and a target 3D construction module. The visualization rendering module is connected to the virtual neural feature field storage module and the target 3D construction module. The target 3D construction module is connected to the visualization rendering module and the target real scale estimation unit. The visualization rendering module uses classic rendering techniques to complete the rendering of the virtual neural feature field. The target 3D construction module visualizes the rendering result of the target in virtual space, thus completing the target 3D construction task.

[0015] The aforementioned target true scale estimation unit includes: a target scale information loading module, an environmental scale information loading module, and a true scale information calculation module. The target scale information loading module is connected to a multi-category target scale information storage module and an environmental scale information loading module. The environmental scale information loading module is connected to a data transmission module, a target true scale information loading module, and a true scale information calculation module. The true scale information calculation module is connected to a target 3D construction module and an environmental scale information loading module. The target scale information loading module loads commonly used scale information and scale information of the target category from the multi-category target scale information storage module. The environmental scale information loading module loads environmental information existing in the real atlas from the data transmission module. The true scale information calculation module integrates the rendered scale information obtained from the target 3D construction module, the target's commonly used scale information, and the possible environmental scale information to calculate the target's scale in the real atlas environment. This scale information has a certain degree of reliability and realism and can be regarded as completing the estimation of the target's true scale to some extent.

[0016] A method for target 3D construction and real scale estimation based on neural feature fields, using the above-mentioned system, is characterized by the following workflow: 1) multi-view image acquisition process; 2) feature information extraction process; 3) standard neural feature field registration process; 4) virtual neural feature field generation process; 5) target real scale estimation process.

[0017] The above multi-view image acquisition process is as follows: First, for a single image acquisition node, there is a certain shooting range as the camera rotates; for multiple image acquisition nodes, coordinated shooting can also obtain a larger image acquisition range. Then, the image acquisition node module performs holistic acquisition of the environment where the target is located under limited viewing conditions, capturing multi-view image information of the real target. The image preprocessing module corrects and restores distorted images based on the actual situation of the multi-view captured images. Next, the corrected multi-view images undergo scale unification, mainly to avoid excessive errors in subsequent 3D target construction. Finally, if the preprocessed multi-view image set meets the design requirements, the data is distributed to other modules through the data transmission module; if the preprocessing results do not meet the requirements, the images are corrected, restored, and scaled again.

[0018] The above operation steps involve feature information extraction as follows: A multi-view image set is used as input data, loaded as two data streams into the target recognition feature information extraction module and the 3D feature information extraction module within the feature information extraction unit. First, the data passes through the target detection module to extract target-related image features, and then through the target recognition module to obtain the target's specific category information. If specific information about the target in the multi-view image set can be obtained, the target region of each image in the multi-view image set S is cropped. This step is mainly to reduce the complexity of subsequent 3D information extraction by manually removing the influence of the target's environment as the background. If the target information cannot be detected and identified, the detection and recognition operations are repeated. Finally, the cropped multi-view image set S... Z Perform 3D feature information extraction.

[0019] The above operation steps are as follows: First, the standard neural feature field loading module loads the specific target category information output by the target recognition module and performs a corresponding retrieval in the standard neural feature field storage module based on this information. If the retrieval is successful, the corresponding standard neural feature field for the target category is obtained from the standard neural feature field library; if no associated standard neural feature field can be found, the retrieval is repeated. Next, the combination of neurons storing 3D feature information is queried in the standard neural feature field. Typically, this combination of neurons stores the 3D feature information such as the shape and appearance of the target category. If the query is successful, subsequent operations are performed; if the query fails, the query is repeated. Finally, the relevant 3D feature information obtained from the 3D feature information extraction module is encoded according to the queried neuron combination. The encoding method must ensure that the encoded result can replace the corresponding neuron combination in the standard neural feature field.

[0020] The above-described virtual neural feature field generation process is as follows: First, based on the initial registration of the neuron combinations that store 3D feature information in the standard neural feature field, the target's true 3D feature information is extracted from multi-view images and weighted according to the registered neuron combinations to obtain new neuron combinations that store true 3D feature information. Since the weight storage format is consistent with that of the corresponding neuron combinations in the standard neural feature field, calculations can be performed under the same scalar. Then, the new neuron combination is fused with the corresponding neuron combination in the standard neural feature field. The 3D feature information fusion module with the standard neural feature field is adaptive, ensuring that the fused neuron combination can still complete the 3D construction task of the target. Simultaneously, a preliminary self-verification module for the fused neuron combination helps complete local verification. If the local self-verification is successful, the fused neuron combination with the completed local self-verification is loaded into the standard neural feature field, completing the virtual neural feature field generation; if the local self-verification fails, the adaptive fusion operation is repeated. Finally, the virtual neural feature field is self-verified and stored in the virtual neural feature field storage module for rendering, ensuring that the generated virtual neural feature field can complete the overall 3D construction task of the target. If the overall self-verification is successful, the virtual neural feature field will have the general three-dimensional feature information of the target category in the standard neural feature field and the real three-dimensional feature information of the target in the real scene; if the overall self-verification fails, the virtual neural feature field needs to be generated again starting from the weight saving operation.

[0021] The above-described target real-scale estimation process is as follows: First, the target scale information constructed in virtual space by the target 3D construction module and the general scale information of this type of target output by the target scale information loading module are loaded to calculate the target's virtual scale information in virtual space. Specifically, the target 3D construction module constructs a series of proportional information about the target scale in virtual space, but it cannot infer the actual scale of the target in the real environment. However, by associating with the general scale information of this type of target, the actual scale of the target can be estimated to a certain extent. Then, the target's scale information in virtual space is verified through a scale judgment operation. If the virtual scale information output is deemed reasonable, it is output normally; if it is deemed unreasonable, the virtual scale information of the target in virtual space is recalculated. Finally, the reasonably determined target virtual scale information and the auxiliary scale information captured in the real environment by the environmental scale information loading module are correlated, and the real scale information calculation module calculates and estimates the target's real-scale information in the real environment. Similarly, if the real scale information output is deemed reasonable, it is output normally; if it is deemed unreasonable, the real scale information of the target in the real environment is recalculated.

[0022] Compared with the prior art, the present invention has the following obvious and prominent substantive features and significant advantages:

[0023] 1. This invention not only greatly reduces the demand for 3D construction equipment and helps promote the miniaturization and simplification of construction equipment in VR, AR and autonomous driving fields, but also, by combining with deep learning ideas, completes the synthesis of new views, effectively solving problems such as feature loss and target occlusion in traditional 3D construction;

[0024] 2. The "virtual neural feature field" proposed in this invention will help users selectively extract three-dimensional feature information according to the needs of actual scenarios, greatly reducing the complexity of construction, saving the resource consumption of repeated construction, and realizing a faster, more accurate, and more personalized target three-dimensional construction capability; at the same time, combined with the real scale estimation unit, it can transmit the target scale information in the virtual space to the real environment, helping to obtain richer environmental perception. Attached Figure Description

[0025] Figure 1 This is a system block diagram of one embodiment of the present invention.

[0026] Figure 2 yes Figure 1 Example of a multi-view image acquisition flowchart.

[0027] Figure 3 yes Figure 1 Example feature information extraction flowchart.

[0028] Figure 4 yes Figure 1 Example of a standard neural feature field registration flowchart.

[0029] Figure 5 yes Figure 1 Example flowchart of virtual neural feature field generation.

[0030] Figure 6 yes Figure 1 Example of a flowchart for target true scale estimation. Detailed Implementation

[0031] Preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings:

[0032] Example 1:

[0033] See Figure 1A target 3D construction and true scale estimation system based on neural feature fields includes a multi-view image acquisition unit 1, a feature information extraction unit 2, a target recognition unit 3, a standard neural feature field registration unit 4, a virtual neural feature field generation unit 5, a visualization rendering and 3D construction unit 6, and a target true scale estimation unit 7. The multi-view image acquisition unit 1 and the feature information extraction unit 2 are connected via wired or wireless means; the feature information extraction unit 2 and the target recognition unit 3 are connected via wired means; the target recognition unit 3 and the standard neural feature field registration unit 4 are connected via wired means; the standard neural feature field registration unit 4 and the virtual neural feature field generation unit 5 are connected via wired means; the virtual neural feature field generation unit 5 and the visualization rendering and 3D construction unit 6 are connected via wired means; and the visualization rendering and 3D construction unit 6 and the target true scale estimation unit 7 are connected via wired means.

[0034] The system of this invention not only significantly reduces the demand for target 3D construction equipment but also enables the synthesis of new views, greatly reducing construction complexity and saving resources from repetitive construction. It possesses faster, more accurate, and more personalized target 3D construction capabilities and the ability to perceive true-scale information. The system of this invention has a simple structure, is easy to operate, and has superior performance. It can adapt to various target 3D construction tasks, has strong practical value, and a wide range of applications.

[0035] Example 2:

[0036] This example is basically the same as Example 1, with the following differences:

[0037] See Figure 1 The multi-view image acquisition unit 1 includes: an image acquisition module 1.1, an image preprocessing module 1.2, and a data transmission module 1.3. The image acquisition module 1.1 is connected to the image preprocessing module 1.2; the image preprocessing module 1.2 is connected to the image acquisition module 1.1 and the data transmission module 1.3; the data transmission module 1.3 is connected to the image preprocessing module 1.2, a target recognition feature information extraction module 2.1, a three-dimensional feature information extraction module 2.2, and an environmental scale information loading module 7.2. The image acquisition module 1.1 captures multi-view image information of the real target, including a multi-view image set S and the shooting pose P of the acquisition node; the image preprocessing module 1.2 corrects and restores each image s in the captured multi-view image set S and unifies its scale to ensure that the photos obtained on a single acquisition node are not affected by the camera tilt angle, and the photos obtained on multiple acquisition nodes are not affected by different shooting scales; the data transmission module 1.3 uploads the preprocessed multi-view image set S. Z And other relevant information.

[0038] See Figure 1 The feature information extraction unit 2 includes: a target recognition feature information extraction module 2.1 and a three-dimensional feature information extraction module 2.2. The target recognition feature information extraction module 2.1 is connected to the data transmission module 1.3 and the target detection module 3.1. The three-dimensional feature information extraction module 2.2 is connected to the data transmission module 1.3, the target recognition module 3.2, and the three-dimensional feature information and standard neural feature field registration module 3.3. The target recognition feature information extraction module 2.1 performs feature extraction on the uploaded multi-view image set. The image features extracted in this operation are only applicable to subsequent target detection and target recognition functions. The three-dimensional feature information extraction module 2.2 performs three-dimensional feature extraction on the uploaded multi-view image set. Specifically, based on the target category information output by the target recognition unit, it crops the area where the target is located in each real environment image from the multi-view image set S and constructs the target multi-view image set S. Z Then for S Z Three-dimensional feature information is extracted, including three-dimensional shape features and three-dimensional appearance features.

[0039] See Figure 1 The target recognition unit 3 includes: a target detection module 3.1, a target recognition module 3.2, and a multi-category target scale information storage module 3.3. The target detection module 3.1 is connected to the target recognition feature information extraction module 2.1 and the target recognition module 3.2; the target recognition module 3.2 is connected to the target detection module 3.1, the three-dimensional feature information extraction module 2.2, and the standard neural feature field loading module 4.2; the multi-category target scale information storage module 3.3 is connected to the target scale information loading module 7.1. The target detection module 3.1 detects the targets to be constructed in the real image set transmitted by data and obtains the target category information; the target recognition module 3.2 classifies the obtained target category information to obtain detailed category information of the targets to be constructed; the multi-category target scale information storage module 3.3 stores the scale information of multiple categories of targets, including proportion information, conventional size information, etc.

[0040] See Figure 1The standard neural feature field registration unit 4 includes: a standard neural feature field storage module 4.1, a standard neural feature field loading module 4.2, and a three-dimensional feature information and standard neural feature field registration module 4.3. The standard neural feature field storage module 4.1 is connected to the standard neural feature field loading module 4.2; the standard neural feature field loading module 4.2 is connected to the standard neural feature field storage module 4.3, the target recognition module 3.2, and the three-dimensional feature information and standard neural feature field registration module 4.3; the standard neural feature field storage module 4.1 stores standard neural feature fields for multiple categories of targets, forming a standard neural feature field library for multiple categories of targets; the above... The standard neural feature field loading module 4.2, based on the detailed category information of the target to be constructed obtained from the target recognition module 3.2, performs a corresponding search in the standard neural feature field library for multi-category targets and loads the standard neural feature field of the target. The above-mentioned three-dimensional feature information and standard neural feature field registration module 4.3 registers the three-dimensional feature information of the target to be constructed with the standard neural feature field of the target category. Specifically, it finds the neuron combination that stores the corresponding three-dimensional feature information in the standard neural feature field and registers the three-dimensional feature information extracted from the real target according to the storage method of the combination to form a new neuron combination with real three-dimensional feature information.

[0041] See Figure 1 The virtual neural feature field generation unit 5 includes: a 3D feature information and standard neural feature field fusion module 5.1, a virtual neural feature field self-verification module 5.2, and a virtual neural feature field storage module 5.3. The 3D feature information and standard neural feature field fusion module 5.1 is connected to the 3D feature information and standard neural feature field registration module 4.3 and the virtual neural feature field self-verification module 5.2; the virtual neural feature field self-verification module 5.2 is connected to the 3D feature information and standard neural feature field fusion module 5.1 and the virtual neural feature field storage module 5.3; and the virtual neural feature field storage module 5.3 is connected to the virtual neural feature field self-verification module 5.2 and the visualization rendering module. Block 6.1; The above-mentioned 3D feature information and standard neural feature field fusion module 5.1 fuses the registered new neuron combination with real 3D feature information with the corresponding part of the standard neural feature field to generate a new neural feature field called a virtual neural feature field. The virtual neural feature field has both the 3D feature information of the real target and other overall 3D feature information of the target of this category; The above-mentioned virtual neural feature field self-verification module 5.2 verifies the fused virtual neural feature field to ensure that the neural feature field can complete the subsequent target 3D construction task; The above-mentioned virtual neural feature field storage module 5.3 stores the verified virtual neural feature field of the target.

[0042] See Figure 1The 3D spatial visualization rendering unit 6 includes: a visualization rendering module 6.1 and a target 3D construction module 6.2. The visualization rendering module 6.1 is connected to the virtual neural feature field storage module 5.3 and the target 3D construction module 6.2; the target 3D construction module 6.2 is connected to the visualization rendering module 6.1 and the real scale information calculation module 7.3; the visualization rendering module 6.1 uses classic rendering technology to complete the rendering of the virtual neural feature field; the target 3D construction module 6.2 visualizes the rendering result of the target in virtual space, completing the target 3D construction task.

[0043] See Figure 1 The target true scale estimation unit 7 includes: a target scale information loading module 7.1, an environmental scale information loading module 7.2, and a true scale information calculation module 7.3. The target scale information loading module 7.1 is connected to the multi-category target scale information storage module 3.3 and the environmental scale information loading module 7.2; the environmental scale information loading module 7.2 is connected to the data transmission module 1.3, the target true scale information loading module 7.1, and the true scale information calculation module 7.3; the true scale information calculation module 7.3 is connected to the target 3D construction module 6.2 and the environmental scale information loading module 7.2. The target scale information loading module 7.1 loads commonly used scale information of the target category from the multi-category target scale information storage module 3.3; the environmental scale information loading module 7.2 loads environmental information existing in the real map atlas from the data transmission module 1.3; the real scale information calculation module 7.3 integrates the rendered scale information obtained from the target 3D construction module 6.2, the target's commonly used scale information, and the possible environmental scale information to calculate the target's scale in the real map atlas environment. This scale information has a certain degree of reliability and authenticity, and can be regarded as completing the estimation of the target's real scale to some extent.

[0044] Example 3:

[0045] This method for target 3D construction and real scale estimation based on neural feature fields adopts the above operation, including the following operation process: 1) multi-view image acquisition process; 2) feature information extraction process; 3) standard neural feature field registration process; 4) virtual neural feature field generation process; 5) target real scale estimation process.

[0046] Example 4:

[0047] This example is basically the same as Example 3, with the following differences:

[0048] See Figure 2The multi-view image acquisition process is as follows: First, for a single image acquisition node, the shooting range is limited by the camera's rotation; for multiple image acquisition nodes, coordinated shooting can achieve a larger image acquisition range. Then, image acquisition node module 1.1 performs holistic acquisition of the environment where the target is located under limited viewing conditions, capturing multi-view image information of the real target. Image preprocessing module 1.2 corrects and restores distorted images based on the actual situation of the multi-view captured images. Next, the corrected multi-view images undergo scale unification, primarily to avoid excessive errors in subsequent 3D target construction. Finally, if the preprocessed multi-view image set meets the design requirements, the data is distributed to other modules via data transmission module 1.3; if the preprocessing results do not meet the requirements, the images are corrected, restored, and scaled again.

[0049] See Figure 3 The operation steps and feature information extraction process are as follows: A multi-view image set is used as input data, loaded as two data streams into the target recognition feature information extraction module 2.1 and the 3D feature information extraction module 2.2 within the feature information extraction unit. First, the data passes through the target detection module 3.1 to extract target-related image features, and then through the target recognition module 3.2 to obtain the specific category information of the target. If specific information about the target in the multi-view image set can be obtained, the target region of each image in the multi-view image set is cropped. This step is mainly to reduce the complexity of subsequent 3D information extraction by manually removing the influence of the target's environment as the background. If the target information cannot be detected and identified, the detection and recognition operations are repeated. Finally, 3D feature information is extracted from the cropped multi-view image set.

[0050] See Figure 4 The standard neural feature field registration process is as follows: First, the standard neural feature field loading module 4.2 loads the specific target category information output by the target recognition module 3.2, and performs a corresponding retrieval in the standard neural feature field storage module 4.1 based on this information. If the retrieval is successful, the standard neural feature field corresponding to the target category is obtained from the standard neural feature field library; if no associated standard neural feature field can be found, the retrieval is repeated. Then, the combination of neurons storing 3D feature information is queried in the standard neural feature field. This combination typically stores the 3D feature information such as the shape and appearance of the target category. If the query is successful, subsequent operations are performed; if the query fails, the query is repeated. Finally, the relevant 3D feature information obtained from the 3D feature information extraction module 4.3 is encoded according to the queried neuron combination. The encoding method must ensure that the encoding result can replace the corresponding neuron combination in the standard neural feature field.

[0051] See Figure 5 The virtual neural feature field generation process is as follows: First, based on the initial registration of the neuron combination that stores 3D feature information in the standard neural feature field, the target's real 3D feature information is extracted from the multi-view images and weighted according to the registered neuron combination to obtain a new neuron combination that stores real 3D feature information. Since the weight storage format is consistent with the weight storage format of this part of the neuron combination in the standard neural feature field, it can perform calculations under the same scalar. Then, the new neuron combination is fused with the corresponding neuron combination in the standard neural feature field. The 3D feature information and standard neural feature field fusion module 5.1 is adaptive, ensuring that the fused neuron combination can still complete the 3D construction task of the target. At the same time, a preliminary self-verification module 5.2 for the fused neuron combination can also help complete the local verification work. If the local self-verification is successful, the fused neuron combination that has completed the local self-verification is loaded into the standard neural feature field to complete the virtual neural feature field generation; if the local self-verification fails, the adaptive fusion operation is repeated. Finally, the virtual neural feature field is self-verified and saved in the virtual neural feature field storage module 5.3 for rendering, ensuring that the generated virtual neural feature field can complete the 3D construction task of the target as a whole. If the overall self-verification is successful, the virtual neural feature field will have the general 3D feature information of the target category in the standard neural feature field and the real 3D feature information of the target in the real scene; if the overall self-verification fails, the virtual neural feature field needs to be generated again starting from the weight saving operation.

[0052] See Figure 6The operation steps for estimating the target's true scale are as follows: First, load the target scale information constructed in virtual space by the target 3D construction module 6.2 and the general scale information of this type of target output by the target scale information loading module 7.1, and calculate the target's virtual scale information in virtual space. Specifically, the target 3D construction module 6.2 constructs a series of proportional information about the target scale in virtual space, but it cannot predict the actual scale of the target in the real environment. However, by associating with the general scale information of this type of target, the actual scale of the target can be estimated to a certain extent. Then, the scale information of the target in virtual space is verified through a scale judgment operation. If the virtual scale information output is deemed reasonable, it is output normally; if it is deemed unreasonable, the virtual scale information of the target in virtual space is recalculated. Finally, the reasonably determined target virtual scale information and the auxiliary scale information captured in the real environment by the environmental scale information loading module 7.2 are associated, and the true scale information calculation module 7.3 calculates and estimates the target's true scale information in the real environment. Similarly, if the output of the true scale information is deemed reasonable, then the output is processed normally; if it is deemed unreasonable, then the true scale information of the target in the real environment is recalculated.

[0053] The system and method described in the above embodiments of the present invention not only significantly reduce the demand for target 3D construction equipment but also enable the synthesis of new views, greatly reducing construction complexity and saving resources consumed by repeated construction. It possesses faster, more accurate, and more personalized target 3D construction capabilities and the ability to perceive true-scale information. The system of the present invention has a simple structure, is easy to operate, and has superior performance. It can adapt to various target 3D construction tasks, has strong practical value, and a wide range of applications.

[0054] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made according to the purpose of the invention. Any changes, modifications, substitutions, combinations or simplifications made based on the spirit and principle of the technical solution of the present invention shall be equivalent substitutions. As long as they meet the purpose of the invention and do not deviate from the technical principle and inventive concept of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A target 3D construction and real scale estimation system based on neural feature fields, comprising a multi-view image acquisition unit (1), a feature information extraction unit (2), a target recognition unit (3), a standard neural feature field registration unit (4), a virtual neural feature field generation unit (5), a visualization rendering and 3D construction unit (6), and a target real scale estimation unit (7), characterized in that: The multi-view image acquisition unit (1) and the feature information extraction unit (2) are connected by wired or wireless means; the feature information extraction unit (2) and the target recognition unit (3) are connected by wired means; the target recognition unit (3) and the standard neural feature field registration unit (4) are connected by wired means; the standard neural feature field registration unit (4) and the virtual neural feature field generation unit (5) are connected by wired means; the virtual neural feature field generation unit (5) and the visualization rendering and 3D construction unit (6) are connected by wired means; the visualization rendering and 3D construction unit (6) and the target real scale estimation unit (7) are connected by wired means. The multi-view image acquisition unit (1) includes an image acquisition module (1.1), an image preprocessing module (1.2), and a data transmission module (1.3). The image acquisition module (1.1) captures multi-view image information of the real target, including a multi-view image set and the shooting pose of the acquisition node. The image preprocessing module (1.2) corrects, restores, and scales each image in the captured multi-view image set. The data transmission module (1.3) uploads the preprocessed multi-view image set. S Other relevant information; The feature information extraction unit (2) includes a target recognition feature information extraction module (2.1) and a three-dimensional feature information extraction module (2.2). The multi-view image set is loaded as data input into the target recognition feature information extraction module (2.1) and the three-dimensional feature information extraction module (2.2) in the feature information extraction unit for feature extraction. The target recognition unit (3) includes a target detection module (3.1), a target recognition module (3.2), and a multi-category target scale information storage module (3.3). The target detection module (3.1) detects the targets to be constructed in the real image set transmitted in the data, and then obtains the specific category information of the target through the target recognition module (3.2) and stores it in the storage module (3.3). The standard neural feature field registration unit (4) includes a standard neural feature field storage module (4.1), a standard neural feature field loading module (4.2), and a three-dimensional feature information and standard neural feature field registration module (4.3). The standard neural feature field loading module (4.2) needs to load the specific category information of the target output by the target recognition module (3.2) and perform corresponding retrieval in the standard neural feature field storage module (4.1) according to the information. The three-dimensional feature information and standard neural feature field registration module (4.3) registers the three-dimensional feature information of the target to be constructed with the standard neural feature field of the target category. The virtual neural feature field generation unit (5) includes a three-dimensional feature information and standard neural feature field fusion module (5.1), a virtual neural feature field self-verification module (5.2), and a virtual neural feature field storage module (5.3). The three-dimensional feature information and standard neural feature field fusion module (5.1) fuses the newly registered neuron combination with real three-dimensional feature information with the corresponding part of the standard neural feature field. The virtual neural feature field self-verification module (5.2) verifies the fused virtual neural feature field and saves it in the virtual neural feature field storage module (5.3). The visualization rendering and 3D construction unit (6) includes a visualization rendering module (6.1) and a target 3D construction module (6.2) to render the virtual neural feature field; The target 3D construction module (6.2) visualizes the rendering result of the target in virtual space; The target true scale estimation unit (7) includes a target scale information loading module (7.1), an environmental scale information loading module (7.2), and a true scale information calculation module (7.3). The target scale information loading module (7.1) loads the commonly used proportion information and scale information of the target of the category from the multi-category target scale information storage module (3.3). The environmental scale information loading module (7.2) loads environmental information existing in the real map set from the data transmission module (1.3); the real scale information calculation module (7.3) calculates and estimates the real scale information of the target in the real environment.

2. The neural feature field based object 3D reconstruction and real scale estimation system of claim 1, wherein: The image acquisition module (1.1) is connected to the data transmission module (1.3) via the image preprocessing module (1.2); the data transmission module (1.3) is connected to the feature information extraction unit (2) and the target real scale estimation unit (7); the image acquisition module (1.1) captures multi-view image information of the real target, including a multi-view image set and the shooting pose of the acquisition node; the image preprocessing module (1.2) corrects and restores each image in the captured multi-view image set and unifies the scale, so as to ensure that the photos obtained on a single acquisition node are not affected by the camera tilt angle, and the photos obtained on multiple acquisition nodes are not affected by different shooting scales.

3. The neural feature field based object 3D reconstruction and real scale estimation system of claim 2, wherein: The target recognition feature information extraction module (2.1) is connected to the data transmission module (1.3) and the target recognition unit (3); the three-dimensional feature information extraction module (2.2) is connected to the data transmission module (1.3), the target recognition unit (3), and the standard neural feature field registration unit (4); the target recognition feature information extraction module (2.1) performs feature extraction on the uploaded multi-view image set, and the image features extracted by this operation are only applicable to subsequent target detection and target recognition functions; the three-dimensional feature information extraction module (2.2) performs three-dimensional feature extraction on the uploaded multi-view image set, specifically, based on the target category information output by the target recognition unit, extracting features from the multi-view image set. S The target area is cropped from each real-world image and combined to form a multi-view image set of the target. S z Then on S z Three-dimensional feature information is extracted, including three-dimensional shape features and three-dimensional appearance features.

4. The neural feature field based object 3D reconstruction and real scale estimation system of claim 3, wherein: The target detection module (3.1) is connected to the target recognition feature information extraction module (2.1) and the target recognition module (3.2); The target recognition module (3.2) is connected to the target detection module (3.1), the three-dimensional feature information extraction module (2.2), and the standard neural feature field registration unit (4); the multi-category target scale information storage module (3.3) is connected to the target scale estimation unit (7); the target detection module (3.1) detects the targets to be constructed in the real image set transmitted in the data and obtains the target category information; the target recognition module (3.2) classifies the obtained target category information and obtains the detailed category information of the target to be constructed; the multi-category target scale information storage module (3.3) stores the scale information of multiple categories of targets, including scale information and conventional size information.

5. The neural feature field based object 3D reconstruction and real scale estimation system of claim 4, wherein: The standard neural feature field storage module (4.1) is connected to the standard neural feature field loading module (4.2); the standard neural feature field loading module (4.2) is connected to the standard neural feature field storage module (4.1), the target recognition module (3.2), and the three-dimensional feature information and standard neural feature field registration module (4.3); the standard neural feature field storage module (4.1) stores standard neural feature fields of multiple categories of targets, forming a standard neural feature field library for multiple categories of targets; the standard neural feature field loading module (4.2) retrieves the corresponding target in the standard neural feature field library for multiple categories of targets based on the detailed category information of the target to be constructed obtained by the target recognition module (3.2), and loads the standard neural feature field of the target; the three-dimensional feature information and standard neural feature field registration module (4.3) registers the three-dimensional feature information of the target to be constructed with the standard neural feature field of the target category, specifically, finding the neuron combination that stores the corresponding three-dimensional feature information in the standard neural feature field, and registering the three-dimensional feature information extracted from the real target according to the storage method of the combination, forming a new neuron combination with real three-dimensional feature information.

6. The neural feature field based object 3D reconstruction and real scale estimation system of claim 5, wherein: The 3D feature information and standard neural feature field fusion module (5.1) is connected to the 3D feature information and standard neural feature field registration module (4.3) and the virtual neural feature field self-verification module (5.2); the virtual neural feature field self-verification module (5.2) is connected to the 3D feature information and standard neural feature field fusion module (5.1) and the virtual neural feature field storage module (5.3); the virtual neural feature field storage module (5.3) is connected to the virtual neural feature field self-verification module (5.2) and the visualization rendering and 3D construction unit (6); the 3D feature information and standard neural feature field fusion module (5.1) The newly registered neuron combination with real three-dimensional feature information is fused with the corresponding part of the standard neural feature field to generate a new neural feature field called a virtual neural feature field. The virtual neural feature field has both the three-dimensional feature information of the real target and other overall three-dimensional feature information of the target of this category. The virtual neural feature field self-verification module (5.2) verifies the fused virtual neural feature field to ensure that the neural feature field can complete the subsequent target three-dimensional construction task. The virtual neural feature field storage module (5.3) stores the verified virtual neural feature field of the target.

7. The neural feature field based object 3D reconstruction and real scale estimation system of claim 6, wherein: The visualization rendering module (6.1) is connected to the virtual neural feature field storage module (5.3) and the target 3D construction module (6.2); The target 3D construction module (6.2) is connected to the visualization rendering module (6.1) and the target real scale estimation unit (7); the visualization rendering module (6.1) uses classic rendering technology to complete the rendering of the virtual neural feature field; The target 3D construction module (6.2) visualizes the rendering result of the target in virtual space, thus completing the target 3D construction task.

8. The neural feature field based object 3D reconstruction and real scale estimation system of claim 7, wherein: The target scale information loading module (7.1) is connected to the multi-category target scale information storage module (3.3) and the environmental scale information loading module (7.2); the environmental scale information loading module (7.2) is connected to the data transmission module (1.3), the target real scale information loading module (7.1), and the real scale information calculation module (7.3); the real scale information calculation module (7.3) is connected to the target 3D construction module (6.2) and the environmental scale information loading module (7.2); the target scale information loading module (7.1) loads the commonly used proportion information and scale information of the target of the category from the multi-category target scale information storage module (3.3); The environmental scale information loading module (7.2) loads environmental information existing in the real map set from the data transmission module (1.3); The real scale information calculation module (7.3) integrates the rendered scale information obtained from the target 3D construction module (6.2), the target's common scale information, and the possible environmental scale information to calculate the target's scale in the real atlas environment. This scale information has a certain degree of reliability and authenticity, and can be regarded as an estimation of the target's real scale to some extent. 9.A method of object 3D construction and real scale estimation based on neural feature field, operating with the system of object 3D construction and real scale estimation based on neural feature field of claim 1, characterized in that The workflow includes: (1) Multi-view image acquisition process; (2) Feature information extraction process; (3) Standard neural feature field registration process; (4) Virtual neural feature field generation process; (5) Target true scale estimation process.

10. The neural feature field based object 3D construction and real scale estimation method according to claim 9, characterized in that: In step (1), the multi-view image acquisition process includes the following steps: for a single image acquisition node, as the camera rotates, there will be a certain angle of shooting range; for multiple image acquisition nodes, collaborative shooting can also obtain a larger image acquisition range; then, the image acquisition node module performs overall acquisition of the environment where the target is located under limited viewing conditions, capturing multi-view image information of the real target; the image preprocessing module (1.2) corrects and restores the distorted images according to the actual situation of the multi-view captured images; then, the scale of the corrected multi-view images is unified, mainly to avoid excessive errors in the subsequent three-dimensional construction of the target; finally, if the preprocessed multi-view image set meets the design requirements, the data is distributed to other modules through the data transmission module (1.3); If the preprocessing results do not meet the requirements, the image will be corrected, restored, and scaled again. Alternatively, in step (2), the feature information extraction process includes using a multi-view image set as data input, which is loaded as two data streams into the target recognition feature information extraction module (2.1) and the three-dimensional feature information extraction module (2.2) in the feature information extraction unit, respectively; first, the data passes through the target detection module (3.1) to extract the target-related image features, and then the target recognition module (3.2) obtains the specific category information of the target; if the multi-view image set can be obtained... S Regarding the specific information about the target, the target region is cropped from each image in the multi-view image set. This step is mainly to reduce the complexity of subsequent 3D information extraction and manually remove the influence of the target's environment as the construction background. If the target information cannot be detected and recognized, the detection and recognition operation is re-performed; finally, the multi-view picture set after cropping S z Three-dimensional feature information extraction is performed; Alternatively, in step (3), the standard neural feature field registration process includes the standard neural feature field loading module (4.2) loading the target specific category information output by the target recognition module (3.2), and performing a corresponding retrieval in the standard neural feature field storage module (4.1) based on this information; if the retrieval is completed, the standard neural feature field corresponding to the target category is obtained from the standard neural feature field library; if the associated standard neural feature field cannot be retrieved, the retrieval is performed again; then, the combination of neurons that stores three-dimensional feature information is queried in the standard neural feature field, which usually stores the three-dimensional feature information such as the shape and appearance of the target of this category. If the query is successful, proceed with the subsequent operations; if the query fails, perform the query operation again; finally, encode the relevant three-dimensional feature information obtained in the three-dimensional feature information extraction module (2.2) according to the neuron combination obtained from the query. The encoding method must ensure that the encoding result can replace the corresponding neuron combination in the standard neural feature field. Alternatively, in step (4), the virtual neural feature field generation process includes, based on the initial registration of the neuron combination that stores 3D feature information in the standard neural feature field, extracting the target's real 3D feature information from the multi-view image set and storing it with weights according to the registered neuron combination to obtain a new neuron combination that stores real 3D feature information. Since the weight storage form is consistent with the weight storage form of this part of the neuron combination in the standard neural feature field, it can perform calculations under the same scalar. Then, the new neuron combination is fused with the corresponding neuron combination in the standard neural feature field. The 3D feature information and standard neural feature field fusion module (5.1) is adaptive and can ensure that the fused neuron combination can still complete the 3D construction task of the target. At the same time, a fusion neuron combination is used to fuse the neuron combination. The initial self-verification module (5.2) of the meta-combination can also help complete the local verification work; if the local self-verification is successful, the fusion neuron combination that has completed the local self-verification is loaded into the standard neural feature field to complete the virtual neural feature field generation; if the local self-verification fails, the adaptive fusion operation is repeated; finally, the virtual neural feature field is self-verified and saved in the virtual neural feature field storage module (5.3) to wait for rendering, ensuring that the generated virtual neural feature field can complete the three-dimensional construction task of the target in an overall manner; if the overall self-verification is successful, the virtual neural feature field will have the general three-dimensional feature information of the target of this category in the standard neural feature field and the real three-dimensional feature information of the target in the real scene; if the overall self-verification fails, the virtual neural feature field generation needs to be repeated from the weight saving operation; Alternatively, in step (5), the target real scale estimation process operation steps include loading the target scale information that has been constructed in the virtual space output by the target 3D construction module (6.2) and the general scale information of the target of this category output by the target scale information loading module (7.1), and calculating the target virtual scale information of the target in the virtual space; Specifically, the target 3D construction module (6.2) constructs a series of scale information about the target in the virtual space, but it cannot predict the actual scale of the target in the real environment. However, by associating the general scale information of this type of target, it can predict the actual scale of the target to a certain extent. Then, the scale information of the target in the virtual space is verified through a scale determination operation; If the output virtual scale information is deemed reasonable, then the output will proceed normally. If the determination is unreasonable, the virtual scale information of the target in the virtual space is recalculated; finally, the virtual scale information of the target determined to be reasonable and the auxiliary scale information in the real environment captured by the environmental scale information loading module (7.2) are correlated, and the real scale information calculation module (7.3) calculates and estimates the real scale information of the target in the real environment. Similarly, if the output's true scale information is deemed reasonable, then the output will proceed normally. If the determination is unreasonable, the true scale information of the target in the real environment will be recalculated.