Three-dimensional reconstruction method and system for fruits and vegetables with high reflectance

Through the depth smoothing method and random diffusion sampling strategy, combined with adaptive scene center adjustment and neural network optimization, the lighting and perspective influence of high-reflectivity fruit and vegetable three-dimensional reconstruction is solved, and efficient and accurate fruit and vegetable surface reconstruction is achieved.

CN120298595APending Publication Date: 2025-07-11SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510433584.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction methods are susceptible to light conditions and sampling perspectives in high-reflectivity fruit and vegetable environments, resulting in poor reconstruction quality or even failure.

Method used

The depth smoothing method and random diffusion sampling strategy are adopted, combined with the adaptive scene center adjustment module, spherical harmonic function coding and hash coding, and the fruit and vegetable surface reconstruction is used using color neural network and TSDF network, and optimized through color rendering formulas and unbiased surface conversion formulas.

Benefits of technology

The stability and accuracy of three-dimensional reconstruction are improved, the high quality of fruit and vegetable surface reconstruction is ensured, the bright spots of high-reflectivity fruit and vegetable are overcome, and the high-effect vegetable phenotype data collection in complex agricultural scenarios is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298595A_ABST
    Figure CN120298595A_ABST
Patent Text Reader

Abstract

The invention discloses a high-reflectance fruit and vegetable three-dimensional reconstruction method and system, a random point and diffusion kernel sampling strategy is used, a sampling range consistent with that of a previous method can be realized, the stability of a small-range sampling area is ensured, optimization is realized through post-processing, a self-adaptive scene center adjustment module is introduced, and high-reflectance fruit and vegetable three-dimensional reconstruction is realized. The non-ideal state of the camera pose during sampling can be corrected in real time, it is ensured that fruits and vegetables to be reconstructed are in the center of a reconstruction scene, and the optimal reconstruction effect is achieved. A previous reconstruction method can only carry out sampling in a fixed sampling angle and range and is difficult to realize in a complex agricultural scene, and the sampling angle can influence the reconstruction quality. The adaptive scene center adjustment module optimizes the sampling angle, improves the reconstruction quality, designs geometric depth smoothing constraints, and ensures that in the reconstruction process, the model does not generate depth estimation errors due to bright spots on the surfaces of high-reflectance fruits and vegetables so as not to finally influence the model reconstruction to generate cavities through direct weighted smoothing on depth values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fruit and vegetable reconstruction, and relates to a three-dimensional reconstruction method and system for fruits and vegetables with high reflectivity. Background Art

[0002] Artificial intelligence has promoted the rapid development of precision agriculture, which is more efficient and productive than traditional production methods. In precision agriculture, measuring plant phenotypes is crucial. Phenotype analysis can quickly understand the traits of complex structural gene expressions and help understand the genetic characteristics of different plant functions. The development of intelligent devices and digital technologies has provided various new development directions for orchard intelligence. Using emerging intelligent devices and algorithms for phenotype acquisition provides a solution for realizing intelligent orchards. Two-dimensional RGB image data can be quickly obtained through a color camera, and at the same time, three-dimensional orchard phenotype data can be obtained at low cost and high efficiency using three-dimensional reconstruction methods, providing analysts with more dimensional analysis data and solutions for phenotypes.

[0003] Traditional three-dimensional reconstruction schemes are usually completed in a laboratory. Because under natural environmental lighting conditions, the reconstruction results are easily affected by lighting conditions and sampling perspectives. For example, for fruits and vegetables with smooth surfaces such as colored peppers, under natural light, there are high-reflectivity areas on the surface, resulting in the reconstruction model being unable to accurately estimate the surface position, leading to poor reconstruction quality or even failure. Reconstruction methods introducing deep learning can improve the stability and generalization of reconstruction, but these methods are often limited to general scenarios, lack applications in the plant environment, and lack solutions for the complexity and dynamics of the plant environment. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem in the prior art that the results of three-dimensional reconstruction are easily affected by lighting conditions and sampling perspectives, resulting in the reconstruction model being unable to accurately estimate the surface position, leading to poor reconstruction quality or even failure. A three-dimensional reconstruction method and system for fruits and vegetables with high reflectivity are provided. The reconstruction method incorporates a depth smoothing method and a random diffusion sampling strategy, eliminating the surface reconstruction defects caused by high reflectivity before, significantly improving the three-dimensional reconstruction effect, and testing on fruit and vegetable datasets such as colored peppers, pitayas, lychees, and tomatoes.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A three-dimensional reconstruction method for fruits and vegetables with high reflectivity, comprising the following steps:

[0007] Obtain fruit and vegetable images;

[0008] Sample the obtained fruit and vegetable images to obtain sampling points and calculate the camera pose;

[0009] Encode the sampling points, and predict the color values of fruits and vegetables based on the data after encoding the sampling points;

[0010] Encode the camera pose, and predict the surface position of fruits and vegetables based on the data after encoding the camera pose;

[0011] Based on the predicted color values of fruits and vegetables and the predicted surface positions of fruits and vegetables, perform surface reconstruction of fruits and vegetables.

[0012] A further improvement of the present invention lies in:

[0013] It also includes introducing a loss function to calculate color loss and geometric depth smoothing loss;

[0014] The calculation of the color loss includes:

[0015] Through ray tracing, calculate the color loss according to the predicted color value of the fruit and vegetable and the true color value to obtain the color loss;

[0016] The calculation of the geometric depth smoothing loss includes using a geometric depth smoothing algorithm to perform depth constraint:

[0017]

[0018] τ(i,j)=exp(-μ||C(i)-C(j)||1

[0019] where h is the depth value. When f(x) = 0, it is considered that the current is the object surface, which is estimated by calculating the termination depth. i is the pixel starting point corresponding to the sampling ray, is the nearby pixel region to be smoothed, C is the pixel color value corresponding on the image, and μ is a hyperparameter.

[0020] It also includes introducing an adaptive scene center adjustment module to perform adaptive scene center adjustment, including:

[0021] Create a linear equation system through the camera pose P and the line-of-sight direction d:

[0022]

[0023] where T represents matrix transpose, and k represents the kth camera position and line-of-sight direction;

[0024] Obtain the average center position of all cameras according to the linear equation system of the scene center:

[0025]

[0026] where center represents the three-dimensional reconstruction scene center, and N represents the number of cameras.

[0027] Encoding the sampling points includes inputting the camera pose into the viewing direction Encoding it into a low-dimensional vector through spherical harmonic functions;

[0028] Encoding the camera pose includes mapping the 3D coordinates into a feature vector through a multi-resolution hash table.

[0029] Construct a color neural network and a TSDF network, and introduce a color rendering formula and an unbiased surface transformation formula for calculating the predicted color value and surface position.

[0030] The color rendering formula includes:

[0031]

[0032] where p(t) is the point on the pixel ray, v is the ray direction, and w b (t) is the unbiased weight representing the SDF.

[0033] The unbiased surface transformation formula includes:

[0034]

[0035] where f(x) is the SDF value predicted by the MLP The function φ b (·) is the logistic density distribution, and the sigmoid function π(·) can control the SDF value range from -1 to 1.

[0036] A three-dimensional reconstruction system for high-reflectivity fruits and vegetables includes:

[0037] An image acquisition module for acquiring fruit and vegetable images;

[0038] An image processing module for sampling the acquired fruit and vegetable images, obtaining sampling points, and calculating the camera pose;

[0039] A prediction module for encoding the sampling points, predicting the color value of the fruits and vegetables based on the data after encoding the sampling points, encoding the camera pose, and predicting the surface position of the fruits and vegetables based on the data after encoding the camera pose;

[0040] A surface reconstruction module for performing surface reconstruction of the fruits and vegetables based on the predicted color value of the fruits and vegetables and the predicted surface position of the fruits and vegetables.

[0041] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of any method described in the present invention.

[0042] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the methods of the present invention are implemented.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] The present invention discloses a three-dimensional reconstruction method for fruits and vegetables with high reflectivity. Using a sampling strategy of random points plus a diffusion kernel, it can achieve the same sampling range as the previous method, while ensuring the stability of the small-range sampling area. Through post-processing optimization, an adaptive scene center adjustment module is introduced, which can correct the unsatisfactory camera pose during sampling in real time, ensure that the fruits and vegetables to be reconstructed are at the center of the reconstruction scene, and achieve the optimal reconstruction effect. The previous reconstruction method could only sample at fixed sampling angles and within a certain range, which was difficult to achieve in complex agricultural scenarios, and the sampling angle would affect the reconstruction quality. The adaptive scene center adjustment module optimizes the sampling angle and improves the reconstruction quality. A geometric depth smoothing constraint is designed. By directly performing weighted smoothing on the depth values, it is ensured that during the reconstruction process, the model will not produce incorrect depth estimates due to bright spots on the surface of high-reflectivity fruits and vegetables, ultimately affecting the appearance of holes in the model reconstruction.

[0045] Furthermore, the spherical harmonic function and hash coding used in the present invention, combined with an unbiased surface conversion formula, improve the efficiency of the reconstruction model while ensuring accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0047] Figure 1 It is a framework diagram of the method for high-reflectivity fruits and vegetables based on a neural implicit field of the invention;

[0048] Figure 2 It is a schematic diagram of the random diffusion kernel of the invention;

[0049] Figure 3 It is a schematic diagram of the geometric depth smoothing constraint of the invention;

[0050] Figure 4 It is a comparison of the reconstruction quality of multiple fruits and the normal map of the invention;

[0051] Figure 5 It is a schematic diagram of the automated sampling device and orchard scene of the model of the invention;

[0052] Figure 6 It is a flowchart for training the invention of the neural implicit field reconstruction method;

[0053] Figure 7 It is a structural diagram of the invention of the automated image acquisition system. Specific implementation manners

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0055] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0057] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed during use, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation to the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0058] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and it does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0059] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and defined, if the terms "set", "install", "connected", "connected" appear, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0060] The following further describes the present invention in detail with reference to the drawings:

[0061] See Figure 1 , the present invention discloses a three-dimensional reconstruction method for fruits and vegetables with high reflectivity, including:

[0062] Step 1: Obtain fruit and vegetable images through an automated image acquisition module;

[0063] The automated image acquisition module mainly consists of a crawler mobile chassis, a control system, a robotic arm, and a visible light camera, specifically including:

[0064] A crawler mobile chassis equipped with autonomous navigation and obstacle avoidance functions, capable of moving flexibly in complex agricultural scenarios;

[0065] A robotic arm equipped with six degrees of freedom, used to carry the visible light camera and adjust the position and angle of the visible light camera to ensure that the camera can take pictures of plants from multiple perspectives;

[0066] A visible light camera is equipped to capture high-resolution images of plants, while trying to maintain high-frame-rate video recording to ensure full-time coverage;

[0067] A control system is equipped to coordinate the actions of the mobile chassis, robotic arm, and camera to achieve automated data acquisition;

[0068] The automated work process of the image acquisition module includes the following steps:

[0069] Step 1.1: The system autonomously navigates to the target area according to the preset task path or the real-time planned route;

[0070] Step 1.2: After reaching the target position, the robotic arm adjusts the position and angle of the camera according to the preset path;

[0071] Step 1.3: The visible light camera takes pictures of fruits and vegetables from multiple angles to collect complete fruit and vegetable phenotype data for subsequent three-dimensional reconstruction algorithm processing.

[0072] Step 2: The three-dimensional reconstruction module

[0073] Step 2.1:

[0074] Use the sampling strategy of the random dot diffusion kernel RDS (Random Dilation Sampling) to randomly select points for the fruit and vegetable colors in the two-dimensional image, and add the diffusion kernel for sampling;

[0075] Specifically: Before the model is trained, it is necessary to sample the color image data collected by the visible light camera; the processing method uses the strategy of the random dot diffusion kernel; the random points are a set of random center points selected from all pixel points; the diffusion kernel is to select 3*3 pixel points in the area near the center point. To expand the sampling range, the sampled pixel points are extended so that each pixel is spaced 1 to 2 pixel points.

[0076] Step 2.2:

[0077] Use the camera internal parameter and pose estimation algorithm based on SFM (structure from motion) to calculate the camera pose of each perspective;

[0078] After obtaining the color image of the fruit and vegetable through the camera, use the camera internal parameter and pose estimation algorithm based on SFM to calculate the camera pose of each perspective. The parameters describe the geometric characteristics and distortion conditions of the camera;

[0079] The camera internal parameter estimation uses the structure from motion recovery SFM algorithm to estimate the internal parameter matrix information of the camera;

[0080] The camera pose estimation algorithm uses the COLMAP algorithm to calculate the camera pose of each color image. The representation form of the pose is the rotation matrix q(w, x, y, z) and the translation matrix (tx, ty, tz).

[0081] Step 2.3:

[0082] Use the adaptive scene center adjustment module ASM (Adaptive Scene Module) to calibrate the scene center, so that the reconstruction model always keeps the fruit and vegetable at the scene center during reconstruction;

[0083] Use the adaptive scene center adjustment module ASM to calibrate the scene center. Specifically: The adaptive scene center adjustment module includes a linear equation system of the scene center, an average center position equation, and a scene size estimation algorithm;

[0084] The linear equation system of the scene center is: Create a linear equation system with the camera pose P and the line of sight direction d:

[0085]

[0086] where \(T\) represents matrix transpose, and \(k\) represents the \(k\)-th camera position and line-of-sight direction;

[0087] The equation for the average center position is as follows: the average center position of all cameras is obtained from the linear equations of the scene center, and the equation is as follows:

[0088]

[0089] where \(center\) represents the 3D reconstruction scene center, and \(N\) represents the number of cameras;

[0090] The scene size estimation algorithm is as follows: calculate the average distance between the camera center pose and the scene center to estimate the size of the scene. The reconstruction radius can be controlled using the scene size, thereby obtaining the best reconstruction quality.

[0091] Step 2.4:

[0092] Encode the camera pose and the 3D coordinates of the sampling points based on Spherical Harmonics Encoding and Hash Table Encoding respectively.

[0093] The Spherical Harmonics Encoding is a set of orthogonal basis functions defined on the sphere. The input camera pose viewing direction \((\theta,\varphi)\) is spherically encoded into a low-dimensional vector, which is combined with the spatial position features and then input into the color neural network for predicting color values;

[0094] The Hash Encoding divides the space through a multi-resolution grid, and maps the 3D coordinates of the sampling points into feature vectors through a multi-resolution hash table, which can greatly compress the size of the neural network and increase the training speed.

[0095] Step 2.5:

[0096] Based on a multi-layer perceptron, establish a color network and a TSDF network, and save the color weights and TSDF weight values (weights for surface reconstruction) of the reconstruction model respectively. During the training process, backpropagation is performed by calculating the loss functions of color and geometric constraints.

[0097] A multi-layer perceptron is a fully connected neural network with a simple structure;

[0098] The color network is a lightweight neural network constructed based on a multi-layer perceptron, which is specifically used to store the color weights during training;

[0099] The TSDF network is a lightweight neural network constructed based on a multi-layer perceptron, which is specifically used to store the weights of the truncated signed distance function during training;

[0100] The truncated signed distance function can accurately represent the surface position in the reconstruction stage.

[0101] Step 2.6:

[0102] By emitting rays and following the ray tracing principle, using the color rendering formula and an unbiased surface transformation formula, high-fidelity neural implicit surface reconstruction is achieved.

[0103] The ray tracing principle calculates the pixel color value by emitting rays from the observation point and sampling the volume data along the ray path. The formula for ray tracing is:

[0104] r = o + td

[0105] where r is the ray, o is the ray origin, t is the sampling point position, and d is the viewing direction;

[0106] The specific color rendering formula is:

[0107]

[0108] where p(t) is the point on the pixel ray, v is the ray direction, and w b (t) is the unbiased weight representing the SDF;

[0109] The unbiased surface transformation formula:

[0110]

[0111] where f(x) is the SDF value predicted by the MLP function φ b (·) is the logistic density distribution, and the sigmoid function π(·) can control the SDF value range from -1 to 1.

[0112] It also includes introducing a loss function:

[0113] Through the camera pose encoding of the captured image and the 3D coordinate encoding of the sampling point, backpropagation is performed using the color constraint and geometric depth smoothing constraint loss, and the network weights are updated;

[0114] The color constraint calculates the loss between the color value obtained by ray tracing and according to the color rendering formula and the true color value. The calculation formula is as follows:

[0115]

[0116] where C pred (r) is the color value extracted from the color network based on the camera pose and the 3D coordinates of the sampling point and calculated using the color rendering formula based on ray tracing, C gt(r) is the true color value in pixel coordinates, and finally the color loss is obtained through weighted calculation;

[0117] The geometric depth smoothing constraint is to deal with the bright spots generated by high-reflectivity fruits and vegetables, which affect the depth judgment of the model. Therefore, the geometric depth smoothing algorithm is used to constrain the depth to control distortion. The formula is as follows:

[0118]

[0119] τ(i,j)=exp(-μ||C(i)-C(j)||1

[0120] where h is the depth value. When f(x) = 0, it is considered that the current is the object surface, which can be estimated by calculating the termination depth. i is the pixel starting point corresponding to the sampling ray, is the nearby pixel area to be smoothed, C is the corresponding pixel color value on the image, and μ is a hyperparameter.

[0121] The present invention discloses specific embodiments:

[0122] Step 1:

[0123] Based on the automated workflow of the crawler mobile chassis and the robotic arm image acquisition module, the system will autonomously navigate to the target area according to the preset task path or the real-time planned route. After reaching the target position, the robotic arm adjusts the position and angle of the camera according to the preset path, and the visible light camera takes pictures of the fruits and vegetables from multiple angles to collect complete fruit and vegetable phenotype data for subsequent three-dimensional reconstruction algorithm processing.

[0124] Step 2:

[0125] For the high-quality surface reconstruction model based on the neural implicit field, it is necessary to use a consumer-grade device (such as a Gopro action camera) to record videos of the fruits and vegetables or scenes to be reconstructed from multiple angles, ensuring as stable a perspective as possible without motion blur or perspective occlusion. Using a high-frame-rate camera can avoid motion blur caused by too fast movement. It is recommended that the frame rate of the shooting device be 30 - 120 frames. After completing the acquisition of the original data, it is necessary to use frame extraction software to extract the image sequence at a fixed frame rate to generate multi-view images. Comparing the method of video recording and frame extraction with single-image shooting can ensure the stability of parameters such as the camera aperture and shutter, and at the same time will not produce overexposed or underexposed images. After completing the frame extraction, use colmap to perform structure from motion, extract the sparse point cloud and camera poses, and the preliminary processing of the original data is completed.

[0126] Step 3: After obtaining the preliminary multi-view image camera poses, preprocessing of the sampling information will be carried out.

[0127] First, randomly select a center point among all pixels, and at the same time select 8 surrounding pixels at a fixed distance interval (the diffusion radius can be adjusted according to the scale to be smoothed). Sample the color values of each selected pixel. Through the sampling method of ray marching (r = o + td), record the ray starting point (related to the camera pose) and the three-dimensional point coordinates in space. During the sampling process, to accelerate the sampling efficiency, first divide the scene space into 1283 voxels, each voxel is marked as 0 (blank) or 1 (with an object), forming a binary voxel field. At the same time, sample at equal intervals only within the voxels marked as 1, skipping the blank areas, significantly reducing the number of sampling points. Secondly, dynamically adjust the sampling step size according to the distance from the camera, with dense sampling near and sparse sampling far. Finally, adaptively adjust the scene center according to all the camera poses to ensure that the reconstructed object is at the scene center to obtain the best reconstruction quality. The adaptive scene center calculation consists of the following: The camera pose P and the line-of-sight direction d create a linear equation system:

[0128]

[0129] where T represents matrix transpose, and k represents the k-th camera position and line-of-sight direction;

[0130] The average center position equation is: Obtain the average center position of all cameras according to the linear equation system of the scene center. The equation is as follows:

[0131]

[0132] where center represents the three-dimensional reconstruction scene center, and N represents the number of cameras.

[0133] Step 4: Perform spherical harmonic function encoding and hash encoding on the camera pose and three-dimensional points respectively.

[0134] The camera pose inputs the viewing direction Encoded into a low-dimensional vector through the spherical harmonic function, combined with the spatial position features and then input into the color neural network to predict the color value (RGB). The spherical harmonic coefficients can efficiently represent complex lighting changes and reduce the network calculation burden.

[0135] The hash encoding divides the space through a multi-resolution grid, maps the 3D coordinates to a feature vector through a multi-resolution hash table, replacing the traditional MLP position encoding. The training speed is greatly improved, and at the same time, it supports high-resolution details (such as textures). After encoding, it is input into the TSDF network together with the scene center value to predict the TSDF value (used to calculate the surface position).

[0136] Step 5: Based on the ray tracing principle, using the color rendering formula and the unbiased surface conversion formula, predict the color value and the surface position (TSDF value) from the color network and the TSDF network to achieve high-fidelity neural implicit surface reconstruction. The ray tracing principle calculates the pixel color value by emitting rays from the observation point and sampling the volume data along the ray path. The formula for ray tracing is: r = o + td

[0137] where: r is the ray, o is the ray origin, t is the sampling point position, and d is the viewing direction;

[0138] The specific color rendering formula is:

[0139]

[0140] where: p(t) is the point on the pixel ray, v is the ray direction, and w b (t) is the unbiased weight representing the SDF;

[0141] The unbiased surface conversion formula

[0142]

[0143] where: f(x) is the SDF value predicted by the MLP The function φ b (·) is the logistic density distribution, and the sigmoid function π(·) can control the SDF value range from -1 to 1.

[0144] Step 6:

[0145] After the calculation in Step 4, the accurate surface values and corresponding color values of the entire 3D model can be obtained, and ray tracing can be performed in reverse to find the color values of the corresponding pixels for calculating the loss function and backpropagation.

[0146] The loss function is a quantitative metric for measuring the difference between the model prediction result and the true value, and its purpose is to provide an optimization direction for backpropagation. The main losses used in this model are the color loss and the geometric depth smoothness loss; the color constraint is the color value obtained by ray tracing and calculated according to the color rendering formula, and the loss calculation is performed with the true color value. The calculation formula is as follows:

[0147]

[0148] where: C pred (r) is the color value extracted from the color network based on the camera pose and the 3D coordinates of the sampling point and calculated using the color rendering formula based on ray tracing, and C gt (r) is the true color value at the pixel coordinates, and finally, the weighted calculation is performed to obtain the color loss;

[0149] The geometric depth smoothing constraint is to address the bright spots generated by highly reflective fruits and vegetables, which affect the depth judgment of the model. Therefore, the geometric depth smoothing algorithm is used to constrain the depth to control distortion. The formula is as follows:

[0150]

[0151] τ(i,j)=exp(-μ||C(i)-C(j)||1

[0152] Where: h is the depth value. When f(x) = 0, it is considered that the current is the object surface, which can be estimated by calculating the termination depth. i is the pixel starting point corresponding to the sampling ray, is the nearby pixel area to be smoothed, C is the corresponding pixel color value on the image, and μ is a hyperparameter.

[0153] After calculating the loss function, the gradient is calculated layer by layer through the chain rule for backpropagation, and the Adam optimizer is used to update the parameters. The number of training rounds can be adjusted according to the complexity of the scene and the level of detail to be reconstructed. The training is completed when the convergence criterion is met, and the corresponding fruit and vegetable mesh model is output. Color meshes, colorless meshes, and normal meshes can be output. At the same time, the training effect and rendering can be viewed in real time.

[0154] The solution disclosed in this embodiment has the following advantages:

[0155] First, the present invention can collect phenotypic data using a visible light camera and convert two-dimensional data into three-dimensional data using a three-dimensional reconstruction model based on a neural implicit field. This method has the advantages of high speed and high accuracy compared with traditional phenotypic acquisition. At the same time, it does not need to rely on precision instruments to collect phenotypes, has a very large cost advantage, and can make phenotypic collection more flexible and overcome complex plant scenes.

[0156] Second, the sampling strategy of the neural implicit field model is improved. The default sampling randomly selects among all pixels, lacking stability when dealing with extreme changes (such as highly reflective fruits). The present invention uses a sampling strategy of random points plus a diffusion kernel, which can achieve the same sampling range as the previous method, while ensuring the stability of the small-range sampling area and realizing optimization through post-processing.

[0157] Third, an adaptive scene center adjustment module is introduced, which can correct the unsatisfactory state of the camera pose during sampling in real time, ensure that the fruits and vegetables to be reconstructed are at the center of the reconstruction scene, and achieve the optimal reconstruction effect. Previous reconstruction methods could only sample at fixed sampling angles and within a certain range, which was difficult to achieve in complex agricultural scenarios. The sampling angle would affect the reconstruction quality. The adaptive scene center adjustment module optimizes the sampling angle and improves the reconstruction quality.

[0158] Fourth, a geometric depth smoothing constraint is designed. By directly performing weighted smoothing on the depth values, it is ensured that during the reconstruction process, the model will not produce incorrect depth estimates due to bright spots on the surface of highly reflective fruits and vegetables, which ultimately affects the appearance of holes in the model reconstruction.

[0159] Fifth, the spherical harmonic function and hash encoding used in the present invention, combined with an unbiased surface conversion formula, improve the efficiency of the reconstruction model while ensuring accuracy. This technical solution has been tested in scenarios of multiple fruits, and the results show that it is superior to previous methods in terms of both reconstruction quality and time.

[0160] Sixth, an automated fruit and vegetable phenotype acquisition system based on a tracked mobile chassis and a robotic arm is applied. By mounting a visible light camera on the robotic arm, it is possible to collect fruit and vegetable phenotype data according to a preset trajectory or a real-time trajectory.

[0161] The schematic diagram of the terminal device provided by an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0162] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention.

[0163] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.

[0164] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0165] The memory can be used to store the computer program and / or modules. By running or executing the computer program and / or modules stored in the memory, and invoking the data stored in the memory, the processor realizes various functions of the terminal device.

[0166] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, Read-Only Memory (ROM), Random Access Memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0167] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A three-dimensional reconstruction method for highly reflective fruits and vegetables, characterized in that, It includes the following steps: Obtain fruit and vegetable images; Sample the obtained fruit and vegetable images, obtain sampling points, and calculate the camera pose; Encode the sampling points, and predict the color values of the fruits and vegetables based on the data after encoding the sampling points; Encode the camera pose, and predict the surface position of the fruits and vegetables based on the data after encoding the camera pose; Based on the predicted color values of the fruits and vegetables and the predicted surface positions of the fruits and vegetables, perform surface reconstruction of the fruits and vegetables.

2. A three-dimensional reconstruction method for fruits and vegetables with high reflectivity according to claim 1, wherein, It also includes introducing a loss function to calculate the color loss and the geometric depth smoothing loss; The calculation of the color loss includes: Through ray tracing, calculate the color loss according to the predicted color values and the true color values of the fruits and vegetables to obtain the color loss; The calculation of the geometric depth smoothing loss includes using a geometric depth smoothing algorithm for depth constraint: τ(i,j)=exp(-μ∥C(i)-C(j)∥1 Among them, h is the depth value. When f(x) = 0, it is considered that the current position is the object surface, which is estimated by calculating the termination depth. i is the pixel starting point corresponding to the sampling ray. is the nearby pixel area to be smoothed, C is the corresponding pixel color value on the image, and μ is the hyperparameter.

3. A three-dimensional reconstruction method for fruits and vegetables with high reflectivity according to claim 1, characterized in that, It also includes introducing an adaptive scene center adjustment module for adaptive scene center adjustment, including: Create a linear equation system through the camera pose P and the line-of-sight direction d: where, T represents matrix transpose, and k represents the k-th camera position and line-of-sight direction; Obtain the average center position of all cameras according to the linear equation system of the scene center: where, center represents the 3D reconstruction scene center, and N represents the number of cameras.

4. A three-dimensional reconstruction method for fruits and vegetables with high reflectivity according to claim 1, characterized in that Encoding the sampling points includes inputting the camera pose into the viewing direction and encoding it into a low-dimensional vector by spherical harmonic functions; The encoding of the camera pose includes mapping 3D coordinates to feature vectors through a multi-resolution hash table.

5. A three-dimensional reconstruction method for fruits and vegetables with high reflectivity according to claim 4, characterized in that Construct a color neural network and a TSDF network, and introduce a color rendering formula and an unbiased surface conversion formula for calculating the predicted color values and surface positions.

6. A three-dimensional reconstruction method for high-reflectivity fruits and vegetables according to claim 5, characterized in that, The color rendering formula includes: where p(t) is the point on the pixel ray, b is the ray direction, and w b (t) is the unbiased weight representing the SDF.

7. A three-dimensional reconstruction method for high-reflectivity fruits and vegetables according to claim 5, characterized in that The unbiased surface conversion formula includes: where f(x) is the SDF value predicted by the MLP function φ b (·) is the logical density distribution, and the sigmoid function π(·) can control the SDF value range from -1 to 1.

8. A three-dimensional reconstruction system for high-reflectivity fruits and vegetables, characterized in that, It includes: An image acquisition module for obtaining fruit and vegetable images; An image processing module for sampling the obtained fruit and vegetable images, obtaining sampling points, and calculating the camera pose; A prediction module for encoding the sampling points, predicting the color values of the fruits and vegetables based on the data after encoding the sampling points, encoding the camera pose, and predicting the surface positions of the fruits and vegetables based on the data after encoding the camera pose; A surface reconstruction module for performing surface reconstruction of the fruits and vegetables based on the predicted color values of the fruits and vegetables and the predicted surface positions of the fruits and vegetables.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.