Field forestry survey data acquisition method and acquisition equipment thereof

By using a multimodal data acquisition and fusion diagnostic model, the problem of separating the measurement of tree geometric parameters from their physiological health status in existing technologies has been solved, enabling efficient and accurate data acquisition and analysis of field forestry surveys.

CN121789060APending Publication Date: 2026-04-03ZHEJIANG HENGYA FORESTRY SURVEY & DESIGN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing field forestry survey techniques cannot achieve high-precision three-dimensional reconstruction and multi-dimensional physiological diagnosis in a single operation, resulting in the separation of tree geometric parameter measurement and physiological health status assessment, low operational efficiency, and limited information dimensions.

Method used

A multimodal diagnostic probe is used to simultaneously acquire RGB image data, multispectral image data, and three-dimensional depth data. Combined with a neural radiation field model and a multimodal fusion diagnostic model, the geometric parameters and health status of forest trees can be acquired and analyzed simultaneously.

Benefits of technology

It achieves seamless integration of tree geometry parameters and health status, improves operational efficiency, provides real-time high-precision measurement and diagnostic reports, and the diagnostic results are objective and reliable, capable of detecting early signs of disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789060A_ABST
    Figure CN121789060A_ABST
Patent Text Reader

Abstract

The invention discloses a field forestry survey data acquisition method and acquisition equipment thereof, and the method comprises the steps: controlling a multi-mode diagnosis probe to scan around the trunk of a target tree, and synchronously collecting RGB image data, multispectral image data and three-dimensional depth data; running a neural radiation field model on the computing terminal, reconstructing a three-dimensional model of the trunk by using the RGB image data and the three-dimensional depth data, and computing geometric parameters; meanwhile, geometric structure features, texture color features and spectral response features are extracted in parallel based on the multi-modal data, and the features are input into the multi-modal fusion diagnosis model to obtain health state information of the tree. Through software and hardware cooperation, accurate measurement of forest geometric parameters and quantitative evaluation of physiological health conditions are simultaneously realized in one operation process, the problems of function splitting and low efficiency in the prior art are solved, and an efficient and comprehensive data acquisition scheme is provided for accurate forestry management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intersection between computer technology and forestry resource monitoring technology, and more specifically, to a data acquisition method and data acquisition equipment, particularly a method and equipment for acquiring data from a field forestry survey. Background Technology

[0002] As a crucial component of Earth's ecosystem, the health of forest resources and their sustainable development have a profound impact on the global environment and economy. Forestry surveys are fundamental work for the systematic inventory, monitoring, and assessment of forest resources. One of their core tasks is the parameter measurement and status assessment of individual trees (i.e., single trees) within forest land. Traditional single-tree survey indicators mainly include geometric parameters, such as the diameter at breast height (DBH) of the trunk at a specific height (usually 1.3 meters above ground), as well as the overall height and location information of the tree. Simultaneously, assessing the health status of trees, such as the presence of pests and diseases, nutritional stress, or physical damage, is also an important aspect of forestry surveys. However, existing field forestry survey data collection technologies still face a series of technical challenges in terms of integration, efficiency, information dimensions, and cost-effectiveness.

[0003] Currently, the main technical methods for collecting data in forestry field surveys fall into the following categories. The first category is the traditional manual contact measurement method. Surveyors typically carry tools such as diameter-at-breast height (DBH) measuring tapes, circumference measuring tapes, and laser altimeters, venturing deep into the forest to manually measure and visually observe each tree to be surveyed. While this method is inexpensive in terms of tools, its inherent drawbacks are significant: low operational efficiency, especially in areas with complex terrain and high tree density, resulting in high labor and time costs; measurement accuracy is easily affected by the operator's subjective judgment, reading habits, and fatigue, making it difficult to guarantee data consistency and reliability; and the information obtained is extremely limited, mainly confined to a few basic geometric dimensions. The assessment of tree health relies entirely on the surveyor's personal experience for qualitative descriptions, lacking objective and quantifiable data support, and failing to meet the data refinement requirements of precision forestry management.

[0004] The second category is survey methods based on three-dimensional laser scanning technology. With the development of sensor technology, terrestrial laser scanning (TLS) has been introduced into the field of forestry surveys. By deploying scanning stations within the sample plot, the laser scanner can emit laser pulses and receive echoes, thereby acquiring a massive amount of three-dimensional coordinate points on the surface of objects within the forest, forming high-precision point cloud data. Through complex post-processing of the point cloud data, such as point cloud registration, denoising, segmentation, and model fitting, fine three-dimensional structural parameters such as diameter at breast height (DBH), tree height, and crown volume of individual trees can be extracted. While this technology offers high measurement accuracy, it also faces significant technical obstacles in large-scale applications. Firstly, high-performance ground-based laser scanners are expensive and typically bulky, resulting in low portability and deployment efficiency in the field. Secondly, data acquisition and processing are separated spatially and temporally. Field operations only acquire the raw point cloud; extensive parameter extraction requires indoor use by professionals with high-performance computers and specialized software, leading to a lengthy workflow and hindering on-site verification and immediate feedback of data quality. Thirdly, lidar acquires purely geometric scene information, failing to perceive object color, texture, or spectral information reflecting plant physiological states. Therefore, this technology is primarily suited for geometric parameter measurement, but it is ineffective for health diagnosis tasks such as identifying early lesions or color anomalies on bark surfaces, or assessing internal physiological stress in plants.

[0005] The third category is survey methods based on ordinary image or visual technologies. In recent years, benefiting from the widespread use of smartphones and the development of computer vision algorithms, methods that utilize mobile terminal cameras, combined with Structure from Motion (SfM) or Simultaneous Localization and Mapping (SLAM) technologies, to perform 3D reconstruction and measure tree parameters have received widespread attention. The operator takes a series of photos or a video with overlapping perspectives around the tree. The algorithm matches corresponding feature points between images, calculates the camera's motion trajectory, and reconstructs a sparse or dense 3D point cloud of the scene, from which parameters such as diameter at breast height (DBH) can be measured. This type of method has significant advantages in terms of equipment cost and portability. However, its technical bottlenecks are also prominent: the accuracy and stability of the reconstruction largely depend on the uniformity of ambient lighting, the richness of the bark texture, and the standardization of the shooting operation. In scenes with excessively strong or dim lighting, or with simple or highly repetitive bark textures, the failure rate of feature point extraction and matching is high, easily leading to reconstruction failure or large geometric distortions in the model, thus affecting measurement accuracy. More importantly, standard RGB cameras can only capture information in the visible light band. While they can identify some diseases that have already shown obvious changes in color and morphology, they cannot effectively detect and assess many physiological stresses in their early stages, which only show abnormal spectral responses in non-visible light bands (such as the near-infrared band). In summary, existing field forestry survey techniques generally suffer from difficulties in accurately measuring tree geometric parameters and quantitatively diagnosing their physiological health status, resulting in a disconnect between technical methods and workflows. Even combining the aforementioned technologies fails to provide an integrated solution that can achieve high-precision 3D reconstruction and multi-dimensional physiological diagnosis in real time during a single, continuous field survey. Surveyors either use one type of equipment focused on measuring geometric dimensions or rely on another method (usually subjective experience) to assess health status, lacking a technical solution that can efficiently, accurately, and cost-effectively integrate these two core survey tasks into a single workflow. Therefore, how to provide a solution that integrates multi-dimensional information perception capabilities, can quickly complete data collection and analysis in the field, and simultaneously outputs high-precision geometric parameters and quantitative health reports is a problem that urgently needs to be solved in the field of forestry survey technology. Summary of the Invention

[0006] The purpose of this invention is to provide a method and equipment for collecting data in forestry field surveys, so as to solve the technical problems of separation of tree geometric parameter measurement and physiological health diagnosis, low work efficiency and single information dimension in the prior art.

[0007] To achieve the above objectives, the present invention provides a method for collecting data in a field forestry survey, which may include the following steps: The multimodal diagnostic probe is controlled to scan around the tree intervention position of the target tree to simultaneously acquire and output multimodal data containing at least RGB image data, multispectral image data, and three-dimensional depth data; The system receives the multimodal data, runs a neural radiation field model on a preset computing terminal, uses the RGB image data and the three-dimensional depth data as input to reconstruct a three-dimensional model of the tree intervention setting location, and calculates the geometric parameters of the target tree based on the three-dimensional model. Based on the multimodal data, geometric structure features, texture color features, and spectral response features are extracted in parallel, and the geometric structure features, texture color features, and spectral response features are input into a preset multimodal fusion diagnostic model to obtain and output the health status information of the target tree.

[0008] Optionally, the multimodal data may further include inertial measurement unit (IMU) data. Before reconstructing the 3D model, the method may further include: determining the camera pose corresponding to each frame of the RGB image data using a visual-inertial simultaneous localization and mapping (VIMMA) algorithm, based on the RGB image data and the IMU data.

[0009] Optionally, when reconstructing the three-dimensional model, the RGB image data can be used as the photometric constraint of the neural radiation field model, and the three-dimensional depth data can be used as the geometric constraint of the neural radiation field model. The neural radiation field model can be trained by jointly optimizing the photometric loss function and the geometric loss function.

[0010] Optionally, when extracting the features, the geometric structure features may be extracted based on the curvature or normal vector information of the surface of the three-dimensional model; the texture color features may be extracted based on the RGB image data through a preset convolutional neural network; and the spectral response features may be vegetation indices calculated based on the multispectral image data.

[0011] Optionally, the multimodal fusion diagnostic model can be a fusion model based on the Transformer architecture, which can perform fusion processing on the geometric structural features, texture color features and spectral response features through the self-attention mechanism and cross-attention mechanism of the fusion model to obtain the health status information.

[0012] Another aspect of the present invention provides a field forestry survey data acquisition device, which, in order to implement the above method, may include: A multimodal diagnostic probe is configured to scan around the tree intervention location of the target tree, and simultaneously acquire and output multimodal data including at least RGB image data, multispectral image data, and three-dimensional depth data. The processing module, connected to the multimodal diagnostic probe, is configured to: receive the multimodal data, run a neural radiation field model, reconstruct a three-dimensional model of the tree intervention location using the RGB image data and the three-dimensional depth data as input, and calculate the geometric parameters of the target tree based on the three-dimensional model; and extract geometric structural features, texture color features, and spectral response features in parallel based on the multimodal data, and input the geometric structural features, texture color features, and spectral response features into a preset multimodal fusion diagnostic model to obtain and output the health status information of the target tree.

[0013] Optionally, the multimodal diagnostic probe can also be configured to acquire inertial measurement unit (IMU) data. The processing module can also be configured to: before reconstructing the 3D model, determine the camera pose corresponding to each frame of the RGB image data based on the RGB image data and the IMU data, using a visual-inertial simultaneous localization and mapping (VIM) algorithm.

[0014] Optionally, when the processing module runs the neural radiation field model, it can be specifically configured to: use the RGB image data as the photometric constraint of the neural radiation field model, use the three-dimensional depth data as the geometric constraint of the neural radiation field model, and train the neural radiation field model by jointly optimizing the photometric loss function and the geometric loss function.

[0015] Optionally, when extracting the features, the processing module can be specifically configured to: extract the geometric structure features based on the curvature or normal vector information of the surface of the three-dimensional model; extract the texture color features based on the RGB image data through a preset convolutional neural network; and calculate the vegetation index based on the multispectral image data as the spectral response feature.

[0016] Optionally, the multimodal fusion diagnostic model in the processing module can be a fusion model based on the Transformer architecture. Specifically, the processing module can be configured to perform fusion processing on the geometric structural features, texture color features, and spectral response features through the self-attention mechanism and cross-attention mechanism of the fusion model to obtain the health status information.

[0017] The technical solution provided by this invention, through the design of a multimodal diagnostic probe integrating multiple sensors and combined with corresponding processing algorithms, enables the simultaneous acquisition and comprehensive analysis of the geometric parameters and health status of forest trees within a continuous workflow. Compared with existing technologies, this invention has the following beneficial effects: First, it boasts high functional integration and rich information dimensions. This invention seamlessly integrates high-precision geometric parameter measurement with multi-dimensional physiological health diagnosis, enabling the output of comprehensive information—which traditionally requires multiple acquisitions and devices—in a single data collection, thus enhancing the comprehensiveness of single-mode information acquisition.

[0018] Secondly, the operation mode is highly efficient and real-time. Through lightweight hardware design and efficient front-end processing algorithms, this invention completes data collection, processing, and result presentation in the field. Operators can obtain measurement and diagnostic reports instantly and assess and confirm data quality on the spot. This avoids invalid data collection and rework that may occur due to the separation of field and office work in traditional technologies, and significantly improves the overall operational efficiency of forestry surveys.

[0019] Third, the diagnostic results are objective and accurate. The health diagnosis method proposed in this invention integrates three-dimensional microscopic geometric features reflecting physical damage, texture and color features reflecting surface lesions, and spectral response features reflecting internal physiological stress. Through cross-validation of multimodal data, its diagnostic results are more objective and reliable than traditional methods that rely on a single information source or human experience, and it is capable of detecting early lesion signs that are not visible to the naked eye. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the system architecture of a field forestry survey data acquisition device provided in an embodiment of the present invention.

[0022] Figure 2 for Figure 1 The diagram shows the hardware structure of the multimodal diagnostic probe in the device.

[0023] Figure 3 This is a flowchart illustrating a method for collecting field forestry survey data, as provided in an embodiment of the present invention.

[0024] Figure 4This is a schematic diagram illustrating the principle of three-dimensional reconstruction and parameter extraction based on neural radiation fields in an embodiment of the present invention.

[0025] Figure 5 This is a schematic diagram of the network structure of the multimodal fusion diagnostic model in an embodiment of the present invention.

[0026] Figure 6 This is an example diagram of the user interface for displaying results on a smart terminal in an embodiment of the present invention. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0028] Example 1 This invention provides a method for collecting data during a forestry field survey. The method can be implemented using the forestry field survey data collection device provided in this invention, which may include hardware and / or software. (See also...) Figure 1 The device may include a multimodal diagnostic probe and a smart terminal that communicate via a data interface. The multimodal diagnostic probe is responsible for data sensing and acquisition in the field, while the smart terminal, such as a smartphone, tablet, or dedicated industrial handheld device, serves as the core for computing and interaction, running the corresponding application and performing subsequent data processing, analysis, and result presentation. The following will refer to... Figure 3 The method flowchart, and combined with Figure 1 , Figure 2 , Figure 4 , Figure 5 and Figure 6 The diagram illustrates the device structure and operating principle, and provides a detailed explanation of the specific steps of the method.

[0029] like Figure 3 As shown, the method provided in this embodiment of the invention may include the following steps: Step S301: Task Initialization and Equipment Preparation. Before conducting field data collection operations, the operator first establishes physical connections and starts the software. Specifically, the operator uses a standard data communication cable, such as a USB Type-C cable supporting high-speed data transmission and power delivery protocols (e.g., USB Power Delivery), to physically connect the interface of the multimodal diagnostic probe to the corresponding interface of the smart terminal. After establishing the physical connection, the operator starts the pre-installed field forestry survey application on the smart terminal.

[0030] Upon startup, the application automatically executes a series of initialization procedures. The program detects the connection status with the multimodal diagnostic probe and confirms the probe's model, firmware version, and the online status of each sensor module via a preset handshake protocol. Further, the program reads the calibration parameters of each sensor stored internally within the probe. These parameters are fundamental for subsequent precise calculations by the algorithm, such as the relative pose and intrinsic parameters between the RGB image acquisition unit and the multispectral image acquisition unit, and the coordinate system transformation matrix between the 3D depth sensing unit and the RGB image acquisition unit. After confirming that the device is functioning correctly and the calibration parameters have been successfully loaded, the application enters the task creation interface. On this interface, operators can input or select relevant metadata for the current survey task, such as the plot number, survey date, operator information, and the target tree number or species. This information will be associated with and stored with the subsequently collected data for easy data management and traceability.

[0031] Step S302: Surround Scanning and Multimodal Data Acquisition. The multimodal diagnostic probe is controlled to scan around the tree intervention location of the target tree, synchronously acquiring and outputting multimodal data including at least RGB image data, multispectral image data, and three-dimensional depth data. Specifically, after task initialization, the operator holds the multimodal diagnostic probe and, following the real-time visual guidance provided by the application interface on the smart terminal, aligns it with the tree intervention location of the target tree, for example, at breast height (approximately 1.3 meters above the ground), and clicks the "Start Acquisition" button on the interface to trigger the data acquisition process.

[0032] At this point, the smart terminal sends a start acquisition command to the central control and interface unit of the multimodal diagnostic probe. The central control unit then activates all sensor modules inside the probe simultaneously through a hardware synchronization triggering mechanism. Specifically, the RGB image acquisition unit begins to capture a continuous color video stream at a preset frame rate (e.g., 30 frames per second) and resolution (e.g., 3840x2160 pixels); the multispectral image acquisition unit simultaneously acquires image sequences in the red and near-infrared bands; the three-dimensional depth sensing unit (e.g., an active structured light module) begins to project an coded infrared speckle pattern and capture the deformed pattern to calculate the depth image in real time; simultaneously, the inertial measurement unit (IMU) begins to output triaxial acceleration and triaxial angular velocity data at a frequency much higher than the image frame rate (e.g., 200 Hz).

[0033] The central control and interface unit assigns a high-precision timestamp to each frame of data from different sensors. This timestamp is derived from a unified hardware clock, ensuring strict time synchronization of all data streams. Subsequently, this multimodal data with synchronized timestamps is packaged and transmitted in real time to the smart terminal via a USB-C interface.

[0034] During data acquisition, operators need to follow the guidance of the application interface, holding the probe and moving it around the tree trunk at a relatively steady speed. To ensure data quality, the application provides real-time acquisition assistance functions. For example, the interface can display video footage captured by the RGB camera in real time, overlaid with a virtual distance indicator. Color changes (e.g., green for moderate distance, yellow for too close or too far, red for beyond the effective range) prompt the operator to maintain the working distance between the probe and the tree trunk surface within an optimal range, such as 0.4 meters to 1.2 meters. Simultaneously, the program can analyze the displacement between IMU data or image frames to determine the operator's movement speed and guide the operator through interface prompts (e.g., a speed bar) to avoid motion blur caused by moving too fast or redundant data caused by moving too slowly. Operators need to complete at least a 360-degree full coverage scan of the tree trunk's breast height section to ensure the integrity of subsequent 3D reconstruction.

[0035] Step S303: Complete Acquisition and Data Preprocessing. After the operator confirms that the target area has been fully scanned, click the "Complete Acquisition" button on the application interface. The smart terminal sends a stop command to the multimodal diagnostic probe, and the probe stops acquiring data and transmitting data. At this point, the physical acquisition operation in the field is complete. Subsequently, the application on the smart terminal will structure and organize all synchronous data streams acquired in this task, including RGB video files, multispectral image sequences, depth data sequences, IMU data records, and metadata files, forming a complete data packet and assigning it a unique identifier for storage. This data packet will serve as the input for all subsequent algorithmic processing.

[0036] Step S304: Fast Camera Pose Estimation. This step is an optional optimization step. In a preferred embodiment, to provide high-quality initial camera pose values ​​for subsequent neural radiation field reconstruction, the system first executes a lightweight fast camera pose estimation algorithm. Specifically, this algorithm is a visual-inertial simultaneous localization and mapping (VI-SLAM) algorithm. It uses the RGB image data acquired in step S302 and high-frequency inertial measurement unit (IMU) data as joint inputs. A specific implementation detail of this algorithm can be broken down into the following process: First, feature point extraction and tracking are performed. For each new RGB image frame received in the data stream, the processing module quickly extracts a set of salient corner features from the image. These features can be extracted using computationally efficient algorithms suitable for real-time processing on mobile devices, such as FAST (Features from Accelerated Segment Test) or ORB (Oriented FAST and Rotated BRIEF). After extracting features in a new frame, the system tracks and matches them with feature points from the previous frame. This tracking process can utilize methods such as KLT (Kanade-Lucas-Tomasi) optical flow to predict the motion of feature points between frames, or, when using algorithms like ORB, by matching feature descriptors. This process establishes geometric constraints between successive camera views.

[0037] Next, IMU pre-integration is performed. During the time interval between two consecutive image frames (e.g., approximately 33.3 milliseconds for a video stream of 30 frames per second), the IMU outputs a series of high-frequency acceleration and angular velocity measurements. The algorithm numerically integrates these IMU measurements in the sensor's body frame. This pre-integration step calculates the changes in relative attitude, velocity, and position during this brief time interval. This calculation is formulated such that the final pre-integrated term depends only on the IMU biases (gyroscope and accelerometer biases) at the start and end of this time interval, making it a self-contained relative motion factor that can be used as a measurement in subsequent optimization steps.

[0038] Next, graph-based state estimation is performed. The algorithm maintains a sliding window that stores a subset of recent keyframes and their associated state variables. Within this window, the state vector to be optimized for each keyframe typically includes its six-DOF pose (position and orientation in world coordinates), velocity, and IMU gyroscope and accelerometer biases. Simultaneously, the spatial positions of the stably tracked and triangulated 3D feature points are also included as variables in the optimization problem. The system then constructs a factor graph and a corresponding nonlinear least-squares optimization problem. The cost function to be minimized mainly consists of two types of error terms (residuals): Visual reprojection error: This error term measures the difference between the 2D coordinates of a 3D map point reprojected onto the image, given the current camera pose estimate and 3D point position estimate, and the true 2D coordinates of the feature point originally observed in the image. Minimizing this error forces the estimated camera pose and map point position to be geometrically consistent with the visual observation.

[0039] IMU pre-integration error: This error term measures the difference between the relative motion (changes in pose, velocity, and bias) calculated from the IMU pre-integration factor and the relative motion derived from the state estimates of two corresponding keyframes in the sliding window. Minimizing this error ensures that the estimated motion trajectory conforms to the physical motion measured by the IMU.

[0040] By using a nonlinear optimization solver, such as one employing the Levenberg-Marquardt algorithm (e.g., Ceres Solver or g2o), the system iteratively adjusts the state variables (pose, velocity, bias, and the position of 3D points) to minimize the sum of all these error terms. This process produces a high-precision, locally consistent motion trajectory. Before entering this main optimization loop, an initialization procedure is typically required to guide the system in estimating the initial metric scale, gravity direction, and initial IMU bias values.

[0041] The final output of the VI-SLAM step is a series of precise six-DOF poses, each corresponding to a frame of RGB image. Compared to pure visual SLAM, this pose information has higher robustness and accuracy, and can better cope with conditions such as rapid movement, weak texture, or changes in lighting. Especially when dealing with special scenes such as rapid operator movement or scanning bark surfaces with simple or repetitive textures, this pose information provides reliable prior conditions for the next step of neural radiation field reconstruction.

[0042] Step S305: Based on the three-dimensional reconstruction and geometric parameter calculation of the neural radiation field, the multimodal data is received, and the neural radiation field model is run on a preset computing terminal. Using the RGB image data and the three-dimensional depth data as input, a three-dimensional model of the tree intervention setting position is reconstructed, and the geometric parameters of the target tree are calculated based on the three-dimensional model. Specifically, refer to... Figure 4 This step receives preprocessed multimodal data as input and reconstructs the 3D geometry and appearance of the tree trunk using an optimized, lightweight Neural Radiance Fields (NeRF) model. The core of this NeRF model is to implicitly represent the 3D scene as a continuous function that maps the coordinates (x, y, z) of a 3D point and a 2D viewing direction (θ, φ) to the point's volume density (σ) and color (c).

[0043] like Figure 4As shown, the lightweight neural radiation field model comprises a multi-resolution hash encoder and a small MLP decoder. The three-dimensional spatial point coordinates x are used as input and are first fed into the multi-resolution hash encoder. This encoder outputs an aggregated feature vector, which, along with the observation direction (encoded by a spherical harmonic function), is input into the small MLP decoder. The MLP decoder has two branches: one outputs the volume density σ, and the other outputs the color c. The entire model is trained end-to-end by jointly optimizing the photometric loss and geometric loss.

[0044] First, the principles and structure of this neural radiation field model are explained. The fundamental goal of the model is to learn a continuous five-dimensional scene function F_Θ:(x, d)→(c, σ), which is approximated by a neural network containing learnable parameters Θ. The input to this function is the coordinates of a three-dimensional point x = (x, y, z) and a two-dimensional viewing direction d = (θ, φ). The output is the color c = (r, g, b) and volume density σ of that point under that viewing direction. The volume density σ is a non-negative scalar, which can be intuitively understood as the differential probability that light is blocked at that point.

[0045] To implement this model on smart terminals with limited computing resources, this embodiment has made the following design and optimization to its network structure and workflow: First, a hybrid network structure with multi-resolution hash encoding is adopted to achieve lightweighting.

[0046] Reference Figure 4 This lightweight NeRF model is not a single, large multilayer perceptron (MLP), but rather employs a hybrid, explicit-implicit combined representation. The structure mainly consists of two parts: a multi-resolution hash encoder and a small MLP decoder.

[0047] Multi-resolution hash encoder: This encoder explicitly discretizes the 3D space into L layers of feature grids with different resolutions (e.g., L=16). The resolution N_l of each grid layer increases exponentially from a coarser N_min (e.g., 16) to a finer N_max (e.g., 512). Each vertex of each grid layer stores a learnable, low-dimensional feature vector (e.g., F=2-dimensional). To control storage overhead while maintaining high resolution, this embodiment uses a hash function to index the feature vectors. Specifically, for each fine-grained virtual grid layer, the integer coordinates of its vertices are mapped by a hash function to an entry in a hash table of size T (e.g., T=2^19). This hash table stores all learnable feature vectors.

[0048] When a 3D point coordinate x is input, for each resolution l, the system first calculates the coordinates of the eight vertices of the grid cell containing that point. Then, a hash function is applied to each of these eight vertex coordinates to retrieve their corresponding feature vectors from the hash table. Next, through linear interpolation, based on the relative position of point x within the grid cell, the interpolation feature f_l(x) of point x at that resolution level is calculated. Finally, the interpolation features {f_l(x)}_(l=1...L) calculated from all L grid layers are concatenated with the input coordinate x itself to form an aggregated feature vector.

[0049] Small MLP Decoder: This decoder is a simple, fully connected neural network with few parameters. For example, it can be a small MLP with two hidden layers and 64 neurons per layer. It receives the aggregated feature vector output by the hash encoder as input, processes it through several layers, and outputs the volume density σ of the spatial point. Then, it concatenates the activation values ​​of the intermediate layers with the observation direction d (usually also encoded in some way, such as spherical harmonic function encoding), and inputs this into the color head of the MLP, finally outputting the direction-dependent color c.

[0050] Through this hybrid structure, most of the high-frequency details of the scene are stored in an explicit feature grid that can be efficiently queried and interpolated, while the small MLP is only responsible for decoding physical properties from these local features. This design reduces the reliance on large neural networks, achieving lightweight models and fast computation.

[0051] Second, the training and joint optimization process based on the principle of volume rendering.

[0052] The training objective of the model is to adjust the parameters Θ (mainly the feature vectors in the hash table and the weights of the small MLP) so that the image rendered from any viewpoint is as consistent as possible with the real-world image. This process is based on the physical principles of volume rendering. For a given camera ray r(t) = o + td (where o is the camera origin, d is the unit direction vector, and t is the step size), the pixel color C(r) that it ultimately represents on the image sensor can be calculated using the following integral formula: C(r) = ∫[t_n to t_f] T(t) * σ(r(t)) * c(r(t), d) dt Where T(t) = exp(-∫[t_n to t] σ(r(s)) ds) is the cumulative transmittance from the near end t_n to point t, representing the probability that the light reaches point t without being blocked.

[0053] In practice, this continuous integral is discretized into a summation form. The system samples N points {t_i} hierarchically along the ray between the near end t_n and the far end t_f. For each sampled point, its volume density σ_i and color c_i are obtained by querying the NeRF network F_Θ. Then, the discretized pixel color Ĉ(r) is calculated numerically.

[0054] This embodiment also jointly optimizes the loss function. The training process is not merely about minimizing the difference between the rendered color and the real color.

[0055] Photometric loss: Calculate the mean square error between the rendered pixel color Ĉ(r) and the pixel color C_gt(r) sampled from the real RGB image, i.e., L_rgb = ||Ĉ(r) - C_gt(r)||_2^2.

[0056] Geometric Loss: The 3D depth data D_gt(r) from the 3D depth sensing unit 130 is used as a strong geometric prior. The system also calculates the expected depth of the ray D̂(r) = ∫[t_n to t_f] T(t) * σ(r(t)) * t dt using volume rendering principles. Then, the loss between the rendered depth and the actual measured depth is calculated, for example, L_depth =||D̂(r) - D_gt(r)||_1.

[0057] Ultimately, the total loss function is defined as a weighted sum of these two terms: L_total = L_rgb + λ * L_depth, where λ is a weighted hyperparameter used to balance appearance and geometric constraints. This total loss is minimized using stochastic gradient descent via an optimizer such as Adam, with gradients backpropagated and network parameters Θ updated. The introduction of geometric loss provides direct supervision of the optimization process regarding the location of scene surfaces, significantly constraining the understanding space, thereby accelerating convergence and improving the geometric accuracy of the final model.

[0058] Third, extract explicit geometric parameters from the trained implicit model.

[0059] Once the network training converges, the model (i.e., the hash table and the weights Θ of the MLP) constitutes a neural asset capable of providing a high-fidelity, continuous three-dimensional representation of the scanned tree trunk. To extract engineering-usable geometric parameters, such as diameter at breast height (DBH), this embodiment employs the following steps: 1. Surface Point Extraction: In the calibrated 3D space, a horizontal cross-section representing the height at breast height (e.g., z = 1.3m) is defined. Within this plane, a set of rays is densely projected outwards in a 360-degree arc from an approximate trunk center location. For each ray r, the algorithm does not require full color rendering but focuses on finding its intersection with the implicit surface. This is achieved by fine-stepping sampling along the ray and querying the NeRF network to obtain the volume density σ of each sampling point. When the cumulative transmittance T(t) first drops to a certain threshold (e.g., 0.5), the position r(t) of that point is considered a high-precision intersection of the ray and the surface.

[0060] 2. Robust Fitting and Parameter Calculation: Repeating the above process yields thousands or even tens of thousands of 3D surface points on the diameter at breast height (DBH) section, accurate to the sub-pixel level. Considering the potential for local irregularities on the trunk surface (such as small bumps or depressions), directly calculating the average diameter of these points is not robust. Therefore, this embodiment employs a robust geometric fitting algorithm, such as the Random Sample Consensus (RANSAC) algorithm. This algorithm can robustly estimate the model parameters from data containing a large number of outliers (i.e., points that do not conform to the main shape). In this scenario, the algorithm repeatedly randomly selects a minimum set of points (e.g., 3 points) to define a circle, and then calculates how many other points (“inside points”) in the dataset fit this circle. Ultimately, the circle with the most “inside point” support is considered the best fit. The diameter of this best-fit circle is output as the DBH value of the trunk.

[0061] Once the neural network training converges, it faithfully represents the 3D geometry and surface texture of the scanned tree trunk as a continuous function. Next, the automatic extraction of geometric parameters begins. To calculate the diameter at breast height (DBH), the algorithm first defines a horizontal cross-section parallel to the ground in the reconstructed 3D space, representing the DBH position. Then, starting from an estimated trunk center point within this cross-section, the algorithm densely projects virtual rays in a 360-degree direction. For each ray, the algorithm steps along the ray direction, queries the NeRF network to obtain the volume density value of each sampling point, and uses numerical integration to find the position where the cumulative transparency decays to a certain threshold (e.g., 0.5). This position is then identified as a surface point on the trunk. In this way, the algorithm can extract tens of thousands of precise 3D surface points on the DBH cross-section with sub-pixel accuracy. Finally, to eliminate the effects of noise and local unevenness, the algorithm employs a robust fitting algorithm, such as random sampling consistency, to fit an optimal 2D circle from these surface points. The diameter of this fitted circle is then output as the final target tree DBH value.

[0062] Steps S306 and S307: Multimodal feature extraction and fusion diagnosis. Based on the multimodal data, geometric structure features, texture color features, and spectral response features are extracted in parallel. The geometric structure features, texture color features, and spectral response features are then input into a preset multimodal fusion diagnosis model to obtain and output the health status information of the target tree.

[0063] The purpose of these two steps is to extract discriminative, multi-dimensional features from multimodal data for subsequent health diagnosis, and to intelligently fuse them to make a final assessment.

[0064] Step S306 specifically includes the parallel extraction of three different types of features.

[0065] First, the extraction of geometric structural features. The data source for these features is the high-fidelity 3D surface represented by the NeRF model trained in step S305. The algorithm calculates the normal vector (i.e., the density gradient direction) and differential geometric quantities such as Gaussian curvature and mean curvature at any point on the surface by differentiating the implicit surface function (volume density field σ(x,y,z)) defined by NeRF. These geometric quantities can finely characterize the microscopic morphology of tree bark. For example, the curvature of a healthy tree bark surface is usually gradual; a wormhole will appear as a local extremum of negative Gaussian curvature, and a crack will appear as a linear region where the normal vector changes drastically. By analyzing the statistical distribution of these geometric quantities, or through template matching, anomaly detection, and other methods, the algorithm can identify and quantify these microstructural anomalies, forming a geometric feature vector describing the degree of physical damage.

[0066] Second, the extraction of texture and color features. The data source for these features is the RGB image data collected in step S302. This embodiment uses a pre-trained lightweight Convolutional Neural Network (CNN), such as the MobileNet or EfficientNet series models. This model has learned rich general visual feature extraction capabilities on large image datasets. To adapt to the specific task of this invention, a dedicated dataset containing various bark lesions, resin exudation, mold, color anomalies, etc., can be used to fine-tune the pre-trained model. In actual extraction, the RGB image patch of the collected tree trunk region is input into the fine-tuned CNN model, and the activation values ​​of its middle layer or penultimate layer are extracted as a high-dimensional texture feature vector that can characterize the texture and color information of the image patch.

[0067] Third, the extraction of spectral response features. The data source for these features is the multispectral image data acquired by the multispectral image acquisition unit 120. First, using pre-calibrated camera parameters, the multispectral image is precisely spatially registered with the RGB image to ensure that pixels in different bands correspond one-to-one. Then, for each pixel, the Normalized Difference Vegetation Index (NDVI) is calculated using its reflectance values ​​in the near-infrared (NIR) and red bands. The formula is: NDVI = (NIR - Red) / (NIR + Red). NDVI is a recognized indicator reflecting the photosynthetic efficiency and cell health of plants. Healthy, vigorous plant tissues strongly reflect near-infrared light and strongly absorb red light, thus having a high NDVI value. However, when plants suffer from pests, diseases, water or nutrient stress, their chlorophyll content decreases, and cell structure is damaged, leading to increased reflection of red light and decreased reflection of near-infrared light, resulting in a significant decrease in the NDVI value. The algorithm calculates the average, variance, entropy, and histogram of NDVI values ​​within the scanned tree trunk area, forming a spectral feature vector that reflects the physiological health status of the tree.

[0068] Step S307 is responsible for intelligently fusing the three heterogeneous features extracted in the previous step and making a final health assessment. (Refer to...) Figure 5 Since the intensity and pattern of different diseases or stress states vary across different modalities, simply concatenating feature vectors is unlikely to achieve ideal results. Therefore, this embodiment employs a multimodal fusion diagnostic model, whose network structure can be based on the Transformer architecture.

[0069] First, the input, network structure, and working principle of this multimodal fusion Transformer model are explained in detail. For example... Figure 5 As shown, the input to the multimodal fusion diagnostic model consists of three modal feature vectors: geometric feature vector V_geom, texture feature vector V_tex, and spectral response feature vector V_spec. These three vectors are first mapped to a unified-dimensional embedding space through their respective linear projection layers, and then superimposed with their corresponding learnable modality type embeddings to form the input sequence X_input. This sequence is then fed into a backbone network composed of N stacked Transformer encoder layers. Each encoder layer includes a multi-head self-attention mechanism and a feed-forward network, achieving intra-modal feature recalibration through self-attention and inter-modal information interaction through cross-attention. Finally, the fused features are processed by a global aggregation and classification head to output a health index and disease probability.

[0070] First, the processing and embedding of the input data. The model receives three feature vectors from step S306 as input: a geometric feature vector (V_geom), a texture color feature vector (V_tex), and a spectral response feature vector (V_spec). These three vectors have different dimensions and semantics, therefore, they need to undergo unified preprocessing and embedding operations before being input into the Transformer backbone network: 1. Modal Feature Embedding: First, the three original feature vectors, which may have different dimensions, are each passed through a modality-specific linear projection layer (i.e., a fully connected layer) to map them into a unified, predefined high-dimensional embedding space, resulting in embedding vectors E_geom, E_tex, and E_spec with dimension d_model. This step makes features from different modalities comparable and interactive in the same mathematical space.

[0071] 2. Modality Type Embedding: To enable the model to distinguish the source modalities of different input vectors, this embodiment introduces modality type embedding. The system creates three learnable embedding vectors M_geom, M_tex, and M_spec, which are added to the feature embedding vectors after linear projection. That is, the final input embedding sequence is X_input = {E_geom + M_geom, E_tex + M_tex, E_spec + M_spec}. In this way, even if the features of different modalities are numerically similar, the model can distinguish them by their attached modality labels.

[0072] Second, the Transformer encoder structure based on the multi-head attention mechanism. The processed input embedding sequence X_input is fed into a Transformer encoder backbone network consisting of N (e.g., N=4) identical encoder layers stacked together. Each encoder layer mainly consists of two core sub-layers: a multi-head self-attention sub-layer and a feed-forward network sub-layer.

[0073] The multi-head self-attention sublayer works as follows: For an input sequence X = {x_1, x_2, x_3} (representing the embedding vectors for geometry, texture, and spectrum, respectively), the self-attention mechanism computes three new vectors for each input vector x_i: a query vector (Query, q_i), a key vector (Key, k_i), and a value vector (Value, v_i). These three are obtained by multiplying x_i by three different, learnable weight matrices (W_q, W_k, W_v).

[0074] Next, to compute the first output vector y_1 (corresponding to the updated representation of the geometric modality), the system calculates the dot product of q_1 with all key vectors {k_1, k_2, k_3} to obtain attention scores. These scores represent the degree of "attention" the geometric modality (query q_1) should have towards all other modalities (including itself). These scores are then scaled (divided by the square root of the key vector dimension) and SoftMax normalized to become a set of weights. Finally, these weights are weighted and summed with all value vectors {v_1, v_2, v_3} to obtain the final output y_1.

[0075] This process can be intuitively understood as follows: the new representation y_1 of the geometric modality is formed by dynamically weighting and aggregating information from itself (v_1), the texture modality (v_2), and the spectral modality (v_3) according to the attention weights of each of the geometric modalities. For example, if the model learns during training that the NDVI value in the spectral features is particularly important when the geometric features are represented by small indentations, then the dot product between q_geom and k_spec will become very large when calculating y_geom, thus giving the spectral value vector v_spec a high weight.

[0076] The multi-head mechanism refers to the process being executed in parallel h times (e.g., h=8), each time using a different set of weight matrices (W_q, W_k, W_v). This allows the model to focus on different aspects of information simultaneously in different representation subspaces. The outputs of all h heads are concatenated and then subjected to a linear projection to obtain the final output of the multi-head self-attention sublayer. This layer simultaneously achieves feature recalibration within a modality (self-attention) and information exchange between modalities (cross-attention).

[0077] Feedforward Neural Network Sublayer: The output of the multi-head self-attention sublayer passes through a residual connection and layer normalization before being fed into a simple feedforward neural network. This network typically consists of two linear layers and a ReLU activation function, used to perform non-linear transformations on the fused features, increasing the model's expressive power. The output of this sublayer also passes through a residual connection and layer normalization.

[0078] The entire Transformer encoder uses N layers stacked together to achieve full and repeated deep fusion of feature information from the three modalities.

[0079] Third, the generation of diagnostic results. After processing through N encoder layers, the model obtains a set of context-aware high-level feature representations {z_geom, z_tex, z_spec} that integrate all modal information. To make the final diagnosis, this embodiment processes these outputs in the following way: 1. Global feature aggregation: The three output vectors can be averaged or max-pooled, or a [CLS] label (a learnable vector specifically for classification) similar to that in BERT can be introduced to aggregate the information of the entire sequence, resulting in a single comprehensive feature vector Z_final that represents the overall health status.

[0080] 2. Classification Head: The synthesized feature vector Z_final is fed into a final classification head. This classification head can be a simple multilayer perceptron (MLP), whose structure can contain one or more hidden layers and a final output layer.

[0081] Health Index Output: A neuron in the output layer can be designed to output a continuous value from 0 to 1, which, after appropriate scaling, becomes a comprehensive health index ranging from 0 to 100. This neuron can use the Sigmoid activation function.

[0082] Disease probability output: The remaining neurons in the output layer, the number of which can correspond to the type of specific disease or stress state that needs to be diagnosed (e.g., longhorn beetle borer, rot, water stress, etc.). This part can use the SoftMax activation function to output a probability distribution representing the probability that the tree has each specific disease.

[0083] Preferably, the multimodal fusion diagnostic model is trained as follows: First, a training set containing N samples is constructed. Each sample includes geometric structure feature vectors, texture color feature vectors, spectral response feature vectors, and corresponding health status labels (such as healthy, longhorn beetle borer, rot, etc.) collected from trees with known health status. Then, the model is trained end-to-end using a combination of cross-entropy loss function and mean squared error loss function until convergence.

[0084] Step S308: Result Generation and Display. This step integrates all the calculation and diagnostic results above and presents them to the operator in a user-friendly manner. (Refer to...) Figure 6 The application interface on the smart terminal generates an instant, visualized comprehensive report. This report may include: a clear, high-precision diameter at breast height (DBH) measurement; a quantitative comprehensive health index from 0 to 100, providing an intuitive assessment of the tree's overall health; and a detailed diagnostic list outlining specific potential risks identified by the model, such as "longhorn beetle borers," "canker," and "water stress," along with corresponding confidence levels or risk grades. Furthermore, the system can precisely mark the location of detected abnormal areas on a two-dimensional unfolded bark texture map or a three-dimensional model through highlighting, coloring, or overlaying labels, providing precise spatial guidance for subsequent forestry management decisions. Operators can view and save this report on-site or export it as a standard format file.

[0085] Step S309: End. After reviewing and saving the results, the operator can choose to measure the next tree or end the survey task.

[0086] Through the above steps, the method provided by the embodiments of the present invention can efficiently and accurately complete the measurement of geometric parameters and health status assessment of trees in the wild in a closed-loop operation process, providing a new technical approach for improving the efficiency and accuracy of forestry surveys.

[0087] Example 2 This invention also provides a field forestry survey data acquisition device. (See attached image.) Figure 1The device may include a multimodal diagnostic probe and a processing module. In one specific implementation of this embodiment, the processing module may be integrated within a smart terminal, implementing its functions through its built-in processor (e.g., CPU, GPU, or dedicated NPU) and software applications. The multimodal diagnostic probe 100 communicates with the smart terminal containing the processing module via wired or wireless means. The various components of the device and their functions will be described in detail below.

[0088] The multimodal diagnostic probe is the front-end data acquisition unit of the device of this invention. Its core function is to scan around the tree intervention design position (e.g., at breast height) of the target tree, and simultaneously acquire and output multimodal data containing at least RGB image data, multispectral image data, and three-dimensional depth data. (Refer to...) Figure 2 To achieve this function, the probe can compactly integrate the following core sensor units and control units in its physical structure: The RGB image acquisition unit is used to acquire high-resolution, true-color image information of the target tree trunk surface. This image information serves as the foundation for subsequent 3D reconstruction to generate realistic texture models, and also as the data source for extracting texture and color features of visual anomalies such as lesions, resin exudation, and mold on the bark surface. This unit can be specifically implemented as an industrial-grade miniature camera module, containing a high-pixel CMOS or CCD image sensor, for example, with at least 12 megapixels and a global shutter function to reduce the rolling shutter effect that may occur during handheld scanning. A distortion-corrected wide-angle fixed-focus lens with a large field of view can be configured in front of this sensor to cover a wider area of ​​the tree trunk surface at a closer working distance. This unit needs to be able to output a high frame rate video stream (e.g., 4K resolution video at 30 frames per second) to ensure sufficient field-of-view overlap between adjacent frames, meeting the view density requirements of subsequent 3D reconstruction algorithms.

[0089] A multispectral image acquisition unit is used to acquire specific narrowband spectral information invisible to the human eye that reflects the physiological health of trees. It obtains key band data needed to calculate vegetation indices, enabling non-destructive detection of early physiological stress. This unit can employ a high-sensitivity monochrome CMOS sensor, with a miniature multi-channel bandpass filter array or a switchable filter wheel positioned in front of it. The filter array includes at least one channel with a center wavelength near the red absorption valley (e.g., 660 nm) and one channel with a center wavelength at the near-infrared high-reflectivity plateau (e.g., 850 nm). Alternatively, a dedicated multispectral image sensor can be used, with on-chip filters of different bands directly integrated into the pixels. To ensure data quality, the optical system of this unit needs to be coaxial or paraxial with the RGB image acquisition unit and undergo rigorous geometric calibration to facilitate subsequent pixel-level image registration.

[0090] A 3D depth sensing unit actively and non-contactly acquires the 3D geometric information of the tree trunk surface. Its output 3D depth data is a key geometric constraint for the subsequent rapid and accurate reconstruction of the neural radiation field model. This unit can employ an active light source. In a preferred embodiment, active structured light technology can be used. The unit may include an infrared laser emitter that projects a laser beam into a pre-coded, invisible infrared speckle pattern through a diffractive optical element (DOE); it also includes a dedicated infrared camera that receives this wavelength to capture the speckle pattern modulated by the tree trunk surface. By calculating the distortion of the speckle pattern and using triangulation principles, the depth value of each pixel within the field of view can be calculated in real time, forming a depth map. Alternatively, Time-of-Flight (ToF) technology can be used to calculate the distance by measuring the time difference between the emitted infrared light pulse and its return.

[0091] An inertial measurement unit (IMU), an optional auxiliary unit, senses the probe's own motion state during scanning, specifically angular velocity and linear acceleration. This high-frequency motion information is a crucial data source for robust and accurate camera pose estimation. The IMU can be a microelectromechanical system (MEMS) chip integrating a three-axis gyroscope and a three-axis accelerometer, capable of outputting six-DOF inertial measurement data at frequencies of several hundred hertz. The IMU should be rigidly mounted inside the probe as close as possible to the optical center to minimize errors caused by lever arm effects.

[0092] The central control and interface unit is responsible for receiving instructions from the smart terminal, controlling the synchronous operation of all the aforementioned sensor units, and performing timestamp alignment, packaging, and high-speed transmission of the acquired multi-source heterogeneous data. It can be implemented using a microcontroller (MCU) or a field-programmable gate array (FPGA). This unit connects to the synchronization pins of all image sensors via a hardware trigger signal line to ensure that all image frames begin exposure at the same time. It appends a high-precision timestamp from a unified hardware clock to each frame of acquired data (RGB image, multispectral image, depth map, IMU sample). Then, this unit aggregates these timestamped data streams into data packets conforming to a specific protocol (e.g., USB VideoClass) and transmits them at high speed to the smart terminal via the USB-C interface.

[0093] The processing module receives data from the multimodal diagnostic probe and executes a series of complex algorithms to ultimately output the geometric parameters and health status information required by the user. This module's functionality is achieved through a dedicated software application running on the smart terminal 200 and underlying hardware (CPU, GPU, NPU). The core functional configuration of this module is as follows: The data reception and pose estimation function is responsible for receiving multimodal data and preparing accurate camera poses for subsequent 3D reconstruction. The processing module receives data streams from the probe in real time via a USB driver and parses and aligns them according to timestamps. When an IMU is configured, the processing module initiates a vision-inertial SLAM algorithm. The software implementation of this algorithm can be an optimization or filtering-based framework that uses image features (such as ORB or FAST corner points) and the pre-integrated results of the IMU as observations to construct a factor graph or state equation. It then solves for the camera position in each frame of the image using nonlinear optimization (such as CeresSolver or g2o) or extended Kalman filtering (EKF) to form a complete motion trajectory.

[0094] The 3D reconstruction and geometric parameter calculation function is responsible for reconstructing a high-fidelity 3D model of the tree trunk using the received data and automatically calculating geometric parameters such as diameter at breast height (DBH). The software implementation of this function is a lightweight modified Neural Radiation Field (NeRF) model. The processing module (specifically, the GPU or NPU of the smart terminal) is responsible for the training and inference of this model. During the training phase, the processing module loads the RGB image and estimated camera pose into the GPU memory, and simultaneously loads depth data as additional input. The model's training code (e.g., an implementation based on PyTorch or TensorFlow Lite) constructs a joint loss function including photometric and geometric losses and iteratively optimizes it using a gradient descent optimizer (such as Adam). After training, the processing module executes the parameter extraction code logic: by defining cross-sections and projected rays in virtual space, it queries the trained NeRF network weights to obtain the surface point set; then, it calls the RANSAC algorithm implemented in a numerical computing library (such as Eigen or NumPy) to perform cylinder fitting and calculate the diameter.

[0095] The multimodal feature extraction and fusion diagnostic function is responsible for extracting deep information related to health status from multimodal data in parallel, performing intelligent fusion, and finally outputting quantitative diagnostic results. The software implementation of this function consists of several sub-modules. The geometric feature extraction sub-module needs to be able to automatically differentiate the NeRF implicit function, calculate the gradient and Hessian matrix of surface points, and thus obtain the normals and curvature. The texture feature extraction sub-module is implemented by loading a pre-trained and fine-tuned CNN model (e.g., stored in ONNX or TFLite format) and performing its forward propagation process. The spectral feature extraction sub-module is implemented by a series of image processing functions, including image registration (e.g., perspective transformation based on calibration parameters) and pixel-level NDVI calculation.

[0096] The implementation of multimodal fusion diagnostic functionality involves loading a pre-trained fusion model based on the Transformer architecture. The model's software code implements the computational logic for multi-head self-attention and cross-attention. The processing module takes the aforementioned three feature vectors as input, performs forward propagation of the fusion model, and then processes the output through a classification layer to ultimately obtain the health index and the probability of various diseases.

[0097] The human-computer interaction and results display function provides users with an interface to guide them through data collection and displays the final results in an intuitive and easy-to-understand manner. This function is implemented by the application's front-end UI framework (e.g., Kotlin / Java for Android or SwiftUI for iOS). It is responsible for drawing real-time guidance graphics during the data collection process (such as...). Figure 6(The user interface shown). After the background processing module calculates the results, this function dynamically updates the UI, displaying numerical values ​​(such as chest diameter and health index) in the text box and populating the list with diagnostic details. For visualization of the 3D model, an embedded 3D rendering engine (such as OpenGL ES or Vulkan) can be called to display the results rendered from the NeRF model, or an explicit mesh model extracted from NeRF using algorithms such as marching cubes, and specific areas can be colored or marked according to the diagnostic results.

[0098] Example 3 This embodiment describes a specific application scenario of using the field forestry survey data acquisition equipment provided by this invention to conduct a rapid survey of individual trees in an artificial pine forest sample plot. The main task of this survey is to measure the diameter at breast height (DBH) of trees with designated numbers within the sample plot and to quickly screen their health status (especially for early signs of longhorn beetle infestation).

[0099] I. Application Environment and Equipment Configuration Application Environment: An artificial Pinus tabulaeformis forest in North China. The trees in the sample plot are planted relatively neatly, with an age of approximately 20 years and an average diameter at breast height (DBH) of 20-30 cm. The understory environment is relatively simple, but uneven light and shadow variations occur in some areas due to canopy shading. Local longhorn beetle infestation is in its early stages; some trees may have subtle, barely perceptible boreholes and traces of resin exudation on the bark surface.

[0100] Multimodal diagnostic probe: The probe selected in this embodiment has the following specific hardware configuration: RGB image acquisition unit: uses Sony IMX586 sensor, set to 4K (3840x2160 pixels) video recording mode at 30 frames per second.

[0101] Multispectral image acquisition unit: It adopts a filter array-based scheme, integrating filter channels with center wavelengths of 660 nm (red light) and 850 nm (near infrared) on a monochrome CMOS sensor, and outputs multispectral image pairs with a resolution of 1280x720 pixels and a frame rate of 30fps.

[0102] 3D Depth Sensing Unit: Employs an Intel RealSense D435i depth camera module, which integrates an active infrared structured light depth engine and a Bosch BMI088 inertial measurement unit (IMU). The depth map output resolution is set to 848x480 pixels, with a frame rate of 30fps. The IMU output frequency is 200Hz.

[0103] Smart Terminal: A flagship Android smartphone equipped with a Qualcomm Snapdragon 8 Gen 2 mobile platform and 12GB of RAM was selected. This phone's processor integrates a high-performance CPU, Adreno GPU, and Hexagon NPU, providing ample computing power for the real-time calculations of this invention.

[0104] Algorithm model parameter settings: The Neural Radiation Field (NeRF) model employs a hash grid-based structure with a hash table size of 2^19, a feature dimension of 4, and 16 multi-resolution levels. In the joint optimization loss function, the weight λ of the geometric loss term is set to 0.1. The training iteration count is set to 15,000.

[0105] Multimodal fusion diagnostic model: The model has been pre-trained using multimodal data containing over 100,000 samples of various pine bark diseases (especially signs of longhorn beetle infestation at different stages). The Transformer encoder layer is set to 4 layers, and the number of heads for multi-head attention is set to 8.

[0106] II. On-site operation process Upon arrival at the sample plot, the investigators connected the diagnostic probes configured above to their smartphones using a USB-C cable. They then opened the Forestry Survey App, entered the sample plot number "PN-2024-01," and selected the "Start New Tree Survey" task.

[0107] The investigator approached a red pine tree numbered "T-054". Following the app's instructions, the investigator aimed the camera at the trunk 1.3 meters above the ground, where a real-time RGB video preview and distance indicator were displayed on the screen. The investigator clicked the start button and began walking around the trunk at a steady pace, keeping the distance indicator green. One lap took approximately 20 seconds. Throughout the walk, the phone interface remained smooth without any lag. After completing the lap, the investigator clicked the finish button.

[0108] The VI-SLAM algorithm first calculates the precise camera trajectory based on RGB video and IMU data within approximately 2 seconds. Then, a lightweight NeRF model begins training. Simultaneously with NeRF training, multimodal feature extraction and fusion diagnostic models are computed in parallel. Upon completion, the app automatically redirects to the results report page.

[0109] Reference Figure 6 The report page clearly presents the analysis results for tree number "T-054": Geometric parameters: The calculated diameter at breast height (DBH) is 26.8 cm. Simultaneously, a high-fidelity 3D textured model that can be freely rotated and scaled is displayed on the screen.

[0110] Health Diagnosis: The overall health index score is 65 / 100 (rated as "Medium Risk"). The diagnostic details list shows "Longhorn beetle boreholes: High risk (89%)" and "Moisture stress: No obvious signs." Simultaneously, three tiny areas are highlighted in red on the 3D model. By magnifying the model, investigators can clearly see that these areas are early holes where longhorn beetles lay eggs and larvae bore in; the bark texture and color around some of these holes show subtle differences from healthy areas. These signs would be easily overlooked by the naked eye during collection.

[0111] III. Comparative Analysis of Technical Effects To more intuitively demonstrate the advantages of this invention, it is compared with an existing solution based on traditional SfM (Structure of Motion) technology. The existing technology uses only a series of RGB photos taken by a smartphone's built-in camera. When processing tree trunk areas with significant light and shadow variations, the SfM algorithm faces difficulties in feature matching, potentially resulting in data gaps in the reconstructed 3D point cloud model in shadow areas, or geometric deformation due to incorrect matching. Furthermore, the generated textures are discrete and discontinuous.

[0112] The technical solution of this invention takes synchronously acquired RGB, multispectral, and depth data streams as input. Benefiting from the direct geometric constraints of the depth data and NeRF's continuous scene representation capabilities, the reconstructed 3D model is complete, continuous, and highly realistic, preserving the microscopic texture details of the bark, even at the boundaries of light and shadow, without geometric errors. More importantly, the output of this invention can also be overlaid with a health status heatmap on this high-fidelity model, visualizing invisible changes in NDVI values ​​or diagnosed disease risks (e.g., red representing high risk, green representing healthy), resulting in a higher information carrying capacity than existing technologies.

[0113] Furthermore, current technology can only provide a single measurement of diameter at breast height (DBH). It offers absolutely no information regarding the early risk of longhorn beetle infestation. Investigators still need to rely on visual inspection and experience for secondary assessment, which is inefficient and prone to missed detections.

[0114] This invention not only provides more accurate diameter at breast height (DBH) values ​​but also offers a complete set of quantitative health diagnostic reports. For early screening of longhorn beetle infestations, this invention can automatically and sensitively detect minute lesions and provide high-confidence risk alerts, with detection capabilities far exceeding those of the human eye. This allows forestry surveyors to shift their focus from tedious repetitive measurements and subjective observations to verifying high-risk trees and implementing precise control measures, thus improving the efficiency and accuracy of forest pest and disease monitoring and control.

[0115] Through the specific application of this embodiment, it can be seen that the field forestry survey data collection method and equipment provided by the present invention apply multimodal perception technology and artificial intelligence algorithms to actual forestry survey scenarios, which not only improves measurement accuracy and efficiency, but also achieves a qualitative leap in information dimension and diagnostic capabilities, demonstrating great practical application value and broad market prospects.

[0116] It should be noted that in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The above embodiments are merely illustrative of the technical solutions of the present invention and not intended to limit it; the present invention has been described in detail only with reference to preferred embodiments. Those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications and substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for collecting data during a field forestry survey, characterized in that, include: The multimodal diagnostic probe is controlled to scan around the tree intervention position of the target tree to simultaneously acquire and output multimodal data containing at least RGB image data, multispectral image data, and three-dimensional depth data; The system receives the multimodal data, runs a neural radiation field model on a preset computing terminal, uses the RGB image data and the three-dimensional depth data as input to reconstruct a three-dimensional model of the tree intervention setting location, and calculates the geometric parameters of the target tree based on the three-dimensional model. Based on the multimodal data, geometric structure features, texture color features, and spectral response features are extracted in parallel, and the geometric structure features, texture color features, and spectral response features are input into a preset multimodal fusion diagnostic model to obtain and output the health status information of the target tree.

2. The method according to claim 1, characterized in that, The multimodal data also includes inertial measurement unit data; Before reconstructing the 3D model, the following is also included: Based on the RGB image data and the inertial measurement unit data, the camera pose corresponding to each frame of the RGB image data is determined by a visual-inertial simultaneous localization and mapping (VIM) algorithm.

3. The method according to claim 1, characterized in that, When reconstructing the three-dimensional model, the RGB image data is used as the photometric constraint of the neural radiation field model, and the three-dimensional depth data is used as the geometric constraint of the neural radiation field model. The neural radiation field model is trained by jointly optimizing the photometric loss function and the geometric loss function.

4. The method according to claim 1, characterized in that, When extracting the features The geometric features are extracted based on the curvature or normal vector information of the surface of the three-dimensional model; The texture color features are extracted based on the RGB image data through a preset convolutional neural network; The spectral response characteristics are vegetation indices calculated based on the multispectral image data.

5. The method according to claim 4, characterized in that, The multimodal fusion diagnostic model is a fusion model based on the Transformer architecture. Through the self-attention mechanism and cross-attention mechanism of the fusion model, the geometric structural features, texture color features and spectral response features are fused to obtain the health status information.

6. A field forestry survey data acquisition device, used to implement the method described in any one of claims 1-5, characterized in that, include: A multimodal diagnostic probe is configured to scan around the tree intervention location of the target tree, and simultaneously acquire and output multimodal data including at least RGB image data, multispectral image data, and three-dimensional depth data. The processing module, which is connected to the multimodal diagnostic probe, is configured as follows: The system receives the multimodal data, runs the neural radiation field model, uses the RGB image data and the three-dimensional depth data as input to reconstruct a three-dimensional model of the tree intervention location, and calculates the geometric parameters of the target tree based on the three-dimensional model. Based on the multimodal data, geometric structure features, texture color features, and spectral response features are extracted in parallel, and the geometric structure features, texture color features, and spectral response features are input into a preset multimodal fusion diagnostic model to obtain and output the health status information of the target tree.

7. The device according to claim 6, characterized in that, The multimodal diagnostic probe is also configured to acquire inertial measurement unit data; The processing module is further configured to: Before reconstructing the 3D model, the camera pose corresponding to each frame of the RGB image data is determined based on the RGB image data and the inertial measurement unit data through a visual-inertial simultaneous localization and mapping (VIM) algorithm.

8. The device according to claim 6, characterized in that, The processing module is specifically configured as follows when running the neural radiation field model: The neural radiation field model is trained by using the RGB image data as the photometric constraint and the three-dimensional depth data as the geometric constraint, and by jointly optimizing the photometric loss function and the geometric loss function.

9. The device according to claim 6, characterized in that, The processing module is specifically configured as follows when extracting the features: The geometric features are extracted based on the curvature or normal vector information of the surface of the three-dimensional model; Based on the RGB image data, the texture color features are extracted using a preset convolutional neural network; The vegetation index is calculated based on the multispectral image data and used as the spectral response feature.

10. The device according to claim 9, characterized in that, The multimodal fusion diagnostic model in the processing module is a fusion model based on the Transformer architecture; the processing module is specifically configured to perform fusion processing on the geometric structural features, texture color features and spectral response features through the self-attention mechanism and cross-attention mechanism of the fusion model to obtain the health status information.