Fruit tree branch sample disease detection image identification and detection method and system

By using multimodal data fusion technology, combining RGB, near-infrared and depth images, efficient and accurate detection of fruit tree diseases has been achieved, solving the problems of insufficient detection accuracy and adaptability in existing technologies, and improving the accuracy and efficiency of disease identification.

CN121788985APending Publication Date: 2026-04-03INST OF TROPICAL & SUBTROPICAL CASH CROP YUNNAN ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing fruit tree disease detection methods suffer from low detection accuracy, low efficiency, and insufficient adaptability. In particular, they are difficult to identify complex disease types, leading to false detections or missed detections.

Method used

Multimodal data fusion technology is used to combine high-resolution RGB images, near-infrared images and depth images. Through grayscale filtering, curvature feature extraction, multi-scale texture consistency analysis and historical database matching, a fused coding vector is generated for disease identification.

Benefits of technology

It improves the accuracy and robustness of fruit tree disease detection, enabling rapid identification of different types of diseases, reducing information loss, and enhancing the comprehensiveness and distinguishing ability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121788985A_ABST
    Figure CN121788985A_ABST
Patent Text Reader

Abstract

The invention discloses a fruit tree branch sample disease detection image identification and detection method and system, and the method comprises the steps: synchronously collecting a high-resolution RGB image, a near-infrared image and a depth image of the surface of a target fruit tree branch, and carrying out the preprocessing of the high-resolution RGB image, and obtaining a to-be-detected image; carrying out graying filtering on pixel points of the to-be-detected image to obtain a grayscale image; selecting a local window for the grayscale image to carry out window sliding to calculate a curvature characteristic value of a central pixel point; extracting a fusion coding vector by using the three-dimensional initial feature vector set, the near-infrared image and the depth image; and constructing a historical fusion coding vector database, and matching the fusion coding vector with the historical fusion coding vector database to obtain a fruit tree branch and trunk disease detection result. Different image analyses are performed by combining a high-resolution RGB image, a near-infrared image and a depth image, and fused features are extracted to understand different disease detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of diseases, and more particularly to a method and system for image recognition and detection of diseases in fruit tree branch and trunk samples. Background Technology

[0002] With the development of agricultural modernization and intelligentization, fruit tree disease detection has gradually become an indispensable part of fruit tree management and production. Diseases not only affect the growth of fruit trees and fruit quality, but can also lead to large-scale yield losses. Therefore, how to efficiently and accurately detect diseases on fruit tree branches and trunks has become an important issue in agricultural research and industrial applications.

[0003] Traditional methods for detecting fruit tree diseases typically rely on manual visual inspection or the use of simple sensor devices (such as temperature and humidity sensors). These methods have limitations in terms of detection accuracy, efficiency, and coverage. Manual inspection is not only time-consuming and prone to subjective errors, but also makes it difficult to conduct comprehensive and timely monitoring of fruit trees. Meanwhile, the accuracy and adaptability of sensor devices are often insufficient to cope with complex natural environments and diverse types of diseases.

[0004] To address these challenges, researchers have gradually introduced computer vision, image processing techniques, and multimodal data fusion technologies. Utilizing information from various sensors, including high-resolution RGB images, near-infrared images, and depth images, and combining this with machine learning and image analysis algorithms, they have achieved more accurate and automated detection of fruit tree diseases. RGB images provide information about the surface color of fruit trees, near-infrared images reveal the physiological state of plants by reflecting the optical properties of their internal tissues, and depth images provide spatial information about the plant's surface morphology.

[0005] Currently, several disease detection methods based on RGB images have been proposed, such as identifying surface diseases of fruit trees through color feature extraction and texture analysis. However, single image features are often insufficient to accurately identify complex diseases, leading to false positives or false negatives. Therefore, combining multimodal data, especially near-infrared and depth images, can effectively improve the robustness and accuracy of disease detection. Summary of the Invention

[0006] The purpose of this invention is to provide a method and system for image recognition and detection of diseases in fruit tree branch and trunk samples, which solves the above-mentioned technical problems pointed out in the prior art.

[0007] This invention provides a method for image recognition and detection of diseases in fruit tree branch and trunk samples, comprising the following steps: High-resolution RGB images, near-infrared images, and depth images of the target fruit tree branch and trunk surface are acquired simultaneously. The high-resolution RGB images are preprocessed to obtain the image to be detected. The pixels of the image to be detected are filtered to grayscale to obtain a grayscale image; a local window is selected in the grayscale image and a sliding window is used to calculate the curvature feature value of the center pixel; the pixels in the grayscale image are encoded with a defined multi-scale radius to obtain a texture consistency index; the curvature feature value and the texture consistency index are concatenated to obtain a three-dimensional initial feature vector set; the three-dimensional initial feature vector set is used to extract and fuse a coded vector with the near-infrared image and the depth image. The fused coding vector includes color high-order moment features, near-infrared features, and depth features; A historical fusion coding vector database is constructed, and the fusion coding vectors are matched with the historical fusion coding vector database to obtain the detection results of fruit tree branch diseases.

[0008] Accordingly, the present invention also proposes an image recognition and detection system for disease inspection of fruit tree branch and trunk samples, comprising: an acquisition module; an analysis module; and an identification module; The acquisition module is used to acquire high-resolution RGB images, near-infrared images, and depth images of the target fruit tree branch surface simultaneously, and to preprocess the high-resolution RGB images to obtain the image to be detected. The analysis module is used to perform grayscale filtering on the pixels of the image to be detected to obtain a grayscale image; select a local window in the grayscale image and perform sliding window calculation on the curvature feature value of the center pixel; encode the pixels in the grayscale image by defining a multi-scale radius to obtain a texture consistency index; concatenate the curvature feature value and the texture consistency index to obtain a three-dimensional initial feature vector set; and use the three-dimensional initial feature vector set with the near-infrared image and the depth image to extract and fuse the encoded vector. The identification module is used to construct a historical fusion coding vector database, and to match the fusion coding vector with the historical fusion coding vector database to obtain the detection results of fruit tree branch diseases.

[0009] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages: Analysis of the above-mentioned method and system for identifying and detecting fruit tree branch and trunk disease samples provided by the present invention shows that, in specific applications, RGB images are preprocessed, converted into grayscale images, and filtered to remove environmental interference and noise. This solution extracts curvature feature values ​​by performing local window sliding calculations on the grayscale image, which can effectively capture changes in the surface morphology of fruit trees caused by diseases, such as bark cracks or surface protrusions and deformations. The extraction of this feature provides important surface morphological information for disease identification. Furthermore, the extraction of multi-scale texture consistency indices can analyze the texture features of fruit tree surfaces at different scales, thereby revealing potential structural changes in diseases. Different types of diseases usually exhibit different changes in texture, and multi-scale analysis can capture disease manifestations of different degrees, further enhancing the reliability and accuracy of disease detection. Furthermore, the generation of fusion coding vectors combines color higher-order moment features, near-infrared features, and depth features from RGB images, near-infrared images, and depth images, which can comprehensively characterize the health status of fruit trees; the fusion of multimodal features helps to extract disease information from multiple dimensions, reduce the information loss that may be caused by a single image source, and improve the comprehensiveness and discriminative ability of detection. Furthermore, this scheme uses historical data matching, combining existing labeled data with the feature vectors of the current image to quickly identify disease types and improve detection speed and accuracy. The historical database matching process calculates the matching degree between fused encoding vectors to determine the presence of diseases and identify disease types. Through the integration of these innovative technologies, this method achieves efficient, accurate, and robust fruit tree disease detection, and has broad practical application prospects. Attached Figure Description

[0010] Figure 1 This is a flowchart of the main process of an image recognition and detection method for disease testing of fruit tree branch and trunk samples, as described in Example 1. Figure 2 This is a schematic diagram of a fruit tree branch and trunk sample disease detection image recognition method according to Example 1. Figure 3 This is a flowchart of the three-dimensional initial feature vector set for a fruit tree branch and trunk sample disease inspection image recognition and detection method according to Example 1; Figure 4 This is a flowchart of the candidate disease region in an image recognition and detection method for disease inspection of fruit tree branch and trunk samples, as described in Example 1. Figure 5 This is a schematic diagram of multi-dimensional sampling points for an image recognition and detection method for disease inspection of fruit tree branches and trunks, as described in Example 1. Figure 6 This is a schematic diagram of the retained sub-nodes in an image recognition and detection method for disease inspection of fruit tree branch and trunk samples according to Example 1. Figure 7 This is a flowchart illustrating the acquisition and fusion encoding vector of a method for detecting and identifying diseases in fruit tree branch and trunk samples, as described in Example 1. Figure 8 This is a flowchart of an image recognition and detection system for disease testing of fruit tree branch and trunk samples, as described in Example 2. Labels: Acquisition module 10; Analysis module 20; Identification module 30. Detailed Implementation

[0011] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings.

[0013] Example 1 like Figure 1 , 2 As shown in the figure, this invention provides a method for image recognition and detection of diseases in fruit tree branch and trunk samples, including the following steps: S10: Simultaneously acquire high-resolution RGB images, near-infrared images, and depth images of the target fruit tree branch and trunk surface; preprocess the high-resolution RGB images to obtain the image to be detected. S20: Filter the pixels of the image to be detected to obtain a grayscale image; select a local window in the grayscale image and perform a sliding window calculation on the curvature feature value of the center pixel; encode the pixels in the grayscale image using a defined multi-scale radius to obtain a texture consistency index; concatenate the curvature feature value and the texture consistency index to obtain a three-dimensional initial feature vector set; use the three-dimensional initial feature vector set with the near-infrared image and the depth image to extract and fuse the encoded vector. The fused coding vector includes color high-order moment features, near-infrared features, and depth features; It should be noted that converting RGB images to grayscale and then filtering them simplifies image processing and reduces the interference of color information on subsequent analysis. Filtering operations (such as Gaussian filtering) can effectively remove noise. The curvature feature value of the center pixel of each local region is calculated using a sliding window. This curvature feature can reflect changes in the shape of the image surface. Diseases often cause changes in the surface structure of branches and trunks (such as protrusion deformation, bark cracks, bark peeling, etc.). In the specific operation process, multi-scale analysis of the image and extraction of texture consistency indicators at different scales help to reflect potential structural changes in the image. Different diseases may manifest as changes in surface texture. The curvature feature and texture consistency indicator can extract subtle changes in the image, helping to identify whether there are structural changes on the plant surface caused by diseases. Furthermore, the multi-scale extraction of texture features helps to capture disease manifestations of different degrees, enhancing the robustness of disease detection. S30: Construct a historical fusion coding vector database, match the fusion coding vector with the historical fusion coding vector database, and obtain the detection results of fruit tree branch diseases; It should be noted that features extracted from RGB, near-infrared, and depth images (such as color higher-order moment features, near-infrared features, and depth features) are fused and encoded. Multimodal feature fusion integrates useful information from different sources, enhancing the distinguishability of features. A feature database containing historical data is created for matching with the current fused encoding vector. In this way, existing disease data can be used to determine the current health status of fruit trees. By fusing multiple features (color, near-infrared, and depth), this step can more comprehensively characterize the health status of fruit trees and improve the accuracy of disease detection. In specific operations, the matching of the historical database can be compared with existing labeled data (i.e., the specific matching is done by calculating the matching value between the two, and then judging according to the preset threshold to determine whether there is fruit tree disease), quickly identifying whether the current fruit tree has disease and the type of disease; see [link to relevant documentation]. Figure 2 , Figure 2 This diagram illustrates the branch cracking and shedding disease in plant surface structures, a typical branch disease problem. (See attached image.) Figure 2 Trunk rot of fruit trees is a fungal disease. The pathogen overwinters on diseased branches and trunks, and the following spring produces spores that are spread by wind and rain, entering through wounds, dead buds, and lenticels. This pathogen has the characteristic of latent infection; after invasion, it first grows on dead tissue in wounds, and then spreads to living tissue, even causing bark cracking or peeling, eventually forming a disease with typical characteristics that can be seen in RGB images, near-infrared images, and even depth images. Specifically, such as Figure 3 As shown, in step S20, the pixels of the image to be detected are filtered to grayscale to obtain a grayscale image; a local window is selected in the grayscale image to calculate the curvature feature value of the center pixel; the pixels in the grayscale image are encoded with a defined multi-scale radius to obtain a texture consistency index; the curvature feature value and the texture consistency index are concatenated to obtain a three-dimensional initial feature vector set; the three-dimensional initial feature vector set is used to extract and fuse the encoding vector with the near-infrared image and the depth image. The specific operation steps are as follows: S21: Convert the image to be detected to grayscale, align the grayscale image to be detected with the depth image, and make a sliding window on the image to be detected with the pixels of the depth image as the center, and calculate the depth gradient of the x-axis and y-axis for the pixels of each pair of sliding windows. The gradient magnitude of the pixel is calculated based on the depth gradient between the x-axis and y-axis (i.e., the larger the gradient magnitude, the more likely the pixel is located in the edge region of depth change (which may be a disease crack or the edge of a branch structure)). Set a baseline standard deviation for the sliding window; use this baseline standard deviation and the gradient magnitude of the pixel to calculate the dynamic standard deviation (i.e., based on the gradient magnitude). Calculate the dynamic standard deviation ;in, Expressed as the baseline standard deviation; Represented as the gradient magnitude of a pixel; It's an adjustment coefficient, representing the dynamic standard deviation in areas with large gradients (pixel edges). As the pixel size decreases, the weight distribution of pixels becomes more concentrated, thus avoiding smoothing across edges and preserving edge details; in flat areas (small gradient). Larger size, stronger smoothing effect); The grayscale standard deviation is calculated for the grayscale values ​​of the center pixel and its neighboring pixels within each sliding window. The grayscale standard deviation and dynamic standard deviation of the sliding window are used to perform bilateral filtering on the pixels of the image to be detected, resulting in smoothed grayscale pixels, which are then used as the smoothed grayscale image of the image to be detected (that is, the grayscale image of the image to be detected is smoothed by double filtering, and the smoothed grayscale pixels form a new grayscale image, which is the smoothed grayscale image of the image to be detected). It should be noted that converting the color image to be detected into a grayscale image simplifies subsequent processing and eliminates the interference of color on structural analysis. The above steps calculate the depth gradient of each pixel to obtain depth change information, revealing edge regions in the image. Diseases typically manifest as cracks or morphological variations, usually located in these edge regions. The dynamic standard deviation is calculated based on the gradient magnitude of each pixel, adjusting the filtering strength accordingly. Smoothing is weaker in edge regions to preserve details, while smoothing is stronger in flat regions to remove noise. The above steps combine the standard deviation of the grayscale image with the dynamic standard deviation of the depth image to perform bilateral filtering, which smooths the image while preserving edge features (such as cracks). The above steps, through joint analysis of depth gradient and grayscale values, are used to locate potential disease areas. Disease cracks and structural variations typically manifest as areas with large depth changes in the image; bilateral filtering effectively smooths non-edge areas while preserving disease-related details. S22: Select a local window from the depth image and slide the window. Within each local window, the center pixel is the local point set (that is, the set of pixels within the local window). The depth image is used to extract the depth value of each pixel in the local window (i.e., since the acquired depth image itself can provide known depth values ​​for the pixels, it will not be described in detail here). The MLS fitting method is used to assign weight values ​​to the neighboring pixels of the center pixel in the local window based on the depth value. The least squares method is used to calculate the local fitting surface for the weight values ​​of the center pixel and the neighboring pixels. Calculate the first and second partial derivatives of the center pixel of the fitted surface; construct the second basic form matrix corresponding to the center pixel using the first and second partial derivatives of the center pixel of the fitted surface. The second basic form matrix is ​​reduced in dimensionality by principal component analysis to obtain the first and second eigenvalues, which are used as principal curvatures (i.e., the dimensionality reduction of the matrix is ​​common knowledge and will not be elaborated further). The principal curvature (i.e., the principal curvature is the first principal curvature and the second principal curvature, obtained from the first feature value and the second feature value) is mapped to a preset fixed interval using shape index to obtain the curvature feature value of the center pixel of the local window (i.e., the curvature feature value of each pixel). It should be noted that the core of MLS mentioned above is for each neighborhood point Assign a weight This weight is related to To the center of the local window The distance is inversely proportional (e.g., using a Gaussian weighting function), making the fit focus more on the region near the center point. The above minimizes the weighted least squares error. Where i represents the center point of the local window. The index of the i-th pixel in the neighborhood; Let represent the actual depth value of the i-th pixel in the neighborhood (i.e., the known depth value provided by the depth image); wi represents the weight of the i-th pixel in the neighborhood; (xi, yi) represents the image plane coordinates of the i-th pixel in the neighborhood. The function is represented as a local surface function to be fitted. Fitting a local window yields the fitted surface. At the center point Calculate the first-order partial derivative at the point (That is, the calculation of the first-order partial derivative is the slope of the fitted surface in the x-axis and y-axis directions) and the second-order partial derivative. (That is, the second-order partial derivative calculates the rate of change of curvature of the fitted surface in the x-axis and y-axis directions, as well as the mixed curvature of the fitted surface.) The degree of distortion may represent the natural curvature of the fruit tree's surface, resembling a saddle shape, or the natural undulations of its texture; the Weingarten matrix (i.e., the second fundamental form matrix). It is The real symmetric matrix whose elements are determined by the first fundamental form (metric) and the second fundamental form (curvature) of the surface; Under parametric surfaces, the calculation formula can be simplified to: Let (That is, E represents the change in the x-direction; when moving along the x-direction on the fitted surface, the actual 3D distance is longer than the distance on the image plane; F represents the coupling in the x and y directions; the x and y directions of the fitted surface are not orthogonal; G represents the change in the y-direction, similar to EE, reflecting the stretching in the y-direction, where E, F, and G are coefficients of the first fundamental form); Let (That is, L represents the component of the normal curvature in the x-direction, and...) Proportional, but divided by Normalization is performed using the normalization factor to fit the curvature of the surface along the x-direction (projection onto the normal direction); M represents the mixed normal curvature component, and... Proportional to the degree of distortion of the fitted surface; N represents the component of the normal curvature in the y-direction, and... Proportional to the curvature of the fitted surface along the y-direction, where L, M, and N are the coefficients of the first fundamental form; output the Weingarten matrix. Finally, the center pixel is obtained. The corresponding Weingarten matrix (i.e., the second fundamental form matrix) ; Extract a scalar value from the Weingarten matrix to describe the "shape" of the local surface at that point, such as concave, convex, or saddle-shaped; for the matrix... Eigenvalue decomposition yields two eigenvalues. and The two feature values ​​obtained are the principal curvatures of the central pixel. ( )and (Right now The shape index mentioned above is a method for... A scalar descriptor mapped to a fixed interval is given by the formula: ,in, This is represented as a linear transformation (i.e., establishing mapping conditions); and This formula maps different shapes to Within a fixed range; the final output obtains the center pixel. curvature eigenvalues ; By extracting local curvature, the above steps enable the system to better identify morphological changes caused by diseases. For example, diseases such as bark cracking, rotting, peeling, or wilting often alter the surface morphology of plants, and changes in local curvature can effectively identify these lesions. S23: Define a multi-scale radius (i.e., ...) for the smoothed grayscale image of the image to be detected. R =3,5,7), collect n sampling points (e.g., 8) on the circumferential trajectory of the circle containing each multi-scale radius; The center gray value of the circle with the radius is compared with the gray values ​​of the n sampling points to obtain an n-bit binary number; The n-bit binary numbers are concatenated end to end, and the number of transitions for each adjacent binary number is counted. Set a preset uniform threshold; determine whether the number of transitions is less than the uniform threshold; If so, then the radius circle under the multi-scale radius is determined to be a uniform pattern (i.e., edge, spot, flat area). The n-bit binary number of the uniform pattern is rotated and normalized (i.e., rotation is to cyclically shift the order of the n-bit binary number, which is to perform displacement) to convert it into the smallest binary number. RIU-LBP encoding is performed on each sampling point in the multi-scale radius to construct a multi-scale histogram. The entropy value is calculated for each multi-scale histogram, and the entropy values ​​of all multi-scale histograms are weighted and fused to obtain the texture consistency index of each sampling point (that is, the texture consistency index of each pixel). It should be noted that a set of multi-scale radii is set for the smoothed grayscale image. ,For example (pixels), for each scale radius Define a circle with radius such that its circumference contains... 8 sampling points (e.g., 8 sampling points, where each sampling point is a pixel, and sampling points can be collected at equal intervals of 45°, 90°, 135°, etc., in the direction of the 8-neighborhood); for each scale radius center point The traditional circular LBP encoding at this scale is calculated using the sampling points (i.e., comparing the grayscale value of each sampling point in the neighborhood with the grayscale value of the center point; if the sampling point is greater than or equal to the center point, it is marked as 1; if the sampling point is less than the center point, it is marked as 0, resulting in an n-bit binary number (e.g., 8 bits when P=8)). The number of transitions from 0 to 1 or 1 to 0 in the n-bit binary number is counted (i.e., checking each pair of adjacent bits (including the last and first bits), incrementing the transition count by 1 each time a 0→1 or 1→0 change is found). If the number of transitions does not exceed 2, the pattern is considered "uniform"; otherwise, it is classified as "mixed". For each uniform pattern selected above, it is converted to the minimum value through cyclic shifting to achieve rotation invariance. Finally, for each sampling point... A RIU-LBP histogram is obtained at each scale. Its dimensions are ( (One uniform pattern + one mixed class); then further calculate the histogram for each scale. Entente value: ,in This is a normalized histogram; the entropy value characterizes the complexity and irregularity of the neighborhood texture of the sampling point (i.e., diseased areas are usually more irregular, and the entropy value may be higher); the multi-scale entropy values ​​are weighted and fused: Among them, weight Based on prior knowledge (e.g., larger scales may better capture macroscopic texture changes), output pixels. Texture consistency index ; The above steps calculate the entropy value of the histogram at each scale. The entropy value reflects the complexity and irregularity of the texture. Diseased areas usually exhibit higher entropy values ​​(complex and irregular textures). S24: Normalize and concatenate the gray values ​​of the smoothed grayscale pixels with the curvature feature values ​​and the texture consistency index to obtain a three-dimensional initial feature vector set (i.e., normalization makes the three features on a single dimension, so that they can be concatenated to form a three-dimensional feature vector set V(p)). It should be noted that the grayscale value, curvature feature value, and texture consistency index (i.e., feature value) are normalized to ensure that they are on the same scale, and then they are stitched together into a set of three-dimensional initial feature vectors. Feature stitching helps to comprehensively consider multiple features (such as color, shape, texture, etc.), so that the final feature vector more comprehensively and accurately reflects the health status of the fruit tree. The above method can more effectively identify different types of diseases. S25: Set parameter dimensions for the three-dimensional initial feature vector set, and obtain an initial parameter set for the three-dimensional initial feature vector set according to the parameter dimensions; use the octree algorithm and the initial parameter set to map each pixel in the image to be detected to the index space, and construct a five-dimensional space of the octree; filter the root nodes of the pixels in the five-dimensional space, and recursively split the root nodes to form child nodes; cluster all child nodes with a preset neighborhood radius for each pixel to obtain candidate disease regions; combine the features extracted from the near-infrared image and the depth image to obtain a fusion encoding vector for the candidate disease regions. It should be noted that the octree algorithm is used to map each pixel in the image and construct a five-dimensional space, which helps to quickly index and process high-dimensional data; the above method clusters all child nodes by setting a neighborhood radius to obtain candidate disease areas; the features extracted from near-infrared and depth images are fused to generate a fused encoding vector, further improving the accuracy of disease identification; the octree algorithm helps to improve processing efficiency, and the clustering algorithm can accurately locate disease areas; the above method, by fusing multiple image features, can more comprehensively identify potential disease areas and improve diagnostic accuracy. Specifically, such as Figure 4 As shown, in step S25, a parameter dimension is set for the three-dimensional initial feature vector set, and an initial parameter set for the three-dimensional initial feature vector set is obtained according to the parameter dimension; each pixel in the image to be detected is mapped to an index space using an octree algorithm and the initial parameter set to construct a five-dimensional space of the octree; the root node is selected from the pixels in the five-dimensional space, and the root node is recursively split to form child nodes; all child nodes are clustered by forming a boundary circle with a preset neighborhood radius for each pixel to obtain candidate disease regions; the fused encoding vector is obtained by combining the features extracted from the near-infrared image and the depth image with the candidate disease regions. The specific operation steps are as follows: S251: Set the parameter dimensions of the x-axis, y-axis, and z-axis for the three-dimensional initial feature vector set (i.e., the z-axis has a dimension of 3, corresponding to the grayscale, curvature, and texture three-dimensional feature vectors in step S24 above; at the same time, the x-axis is the neighborhood radius; the y-axis is the minimum number of points in the neighborhood; together they form a 5-dimensional parameter dimension). For each parameter dimension, a search range is set (i.e., a search range is set for each coordinate axis, such as the weight of the 3D feature vector of the z-axis being equal to 1) and a non-uniform search step size is used (i.e., if the x-axis is too large (neighborhood radius), all points will be grouped into one cluster; if the x-axis is too small (neighborhood radius), each point is noise, and the optimal value is usually in an "intermediate region"; for example, checking every 0.5 or even 0.1 units to find coarse or fine regions; the y-axis usually takes small integer values ​​(such as 3, 4, 5) for better results, and taking values ​​that are too large (such as above 15) can easily ignore small-scale real disease clusters; the z-axis makes the weights of grayscale, curvature, and texture equal, and the weights [0.33, 0.33, 0.33] mean that the three are equally important, which is a possible center point, but in reality, a certain feature may be extremely important (such as curvature [0.8, 0.1, 0.1]), or a certain feature may be completely useless (such as [0.49, 0.51, ...). [0.0], texture weight is 0); The Latin hypercube sampling method is used to divide the search range of the parameter dimensions into search intervals; for each search interval, random uniform sampling is performed according to the non-uniform search step size to obtain multi-dimensional sampling points (i.e., exactly one point in each x-axis interval, exactly one point in each y-axis interval, and one point in each z-axis interval, with at least one sampling point in each interval, such as...). Figure 5 (as shown) The coordinate position of each multi-dimensional sampling point is determined, and the initial parameter set of each search interval is determined based on the coordinate position of each multi-dimensional sampling point (that is, each search interval contains parameters of each dimension, forming an initial parameter set). It should be noted that the above-mentioned setting of the parameter dimensions of the feature vector provides diverse attributes for subsequent feature fusion. These dimensions include grayscale, curvature, texture (z-axis), neighborhood radius (x-axis), and minimum number of points in the neighborhood (y-axis). The selection of each dimension and the setting of the search range provide flexible adjustment space for subsequent feature extraction. By adjusting the feature weights and search range, it is possible to more accurately capture possible disease areas in the image (for example, a larger neighborhood radius may be suitable for identifying larger lesions, while a smaller neighborhood radius helps in the extraction of details). The non-uniform search step size and Latin hypercube sampling method make the selection of parameters more efficient, avoiding the negative impact of excessively large or small parameter values ​​on the feature extraction effect. Through the optimization of these parameters, it is possible to better identify disease areas of different scales and features, thereby improving the accuracy of disease identification. S252: An octree is used to perform a weighted transformation on the initial 3D feature vector of each pixel in the image to be detected, using the parameter dimension of the z-axis (i.e., since the parameter dimension of the z-axis is the feature weight vector of gray level, curvature, and texture, the feature weights can be used to perform a weighted transformation on the initial 3D feature vector (gray level, curvature, and texture) to obtain a normalized weighted feature vector that can form a five-dimensional space with the subsequent spatial coordinates, i.e., weighted gray level value, weighted curvature value, and weighted texture value). This is further combined with the x-axis and y-axis coordinate positions of each pixel in the initial parameter set for mapping. To the index space, construct the five-dimensional space of the octree (i.e., the coordinates of each pixel in the five-dimensional space = (weighted gray value, weighted curvature value, weighted texture value, spatial x-coordinate, spatial y-coordinate; that is, the weighted gray value, weighted curvature value, and weighted texture value are obtained by weighting the initial three-dimensional feature vector using the parameter dimension of the z-axis, while the spatial x-coordinate and spatial y-coordinate are the coordinate positions of the pixel in two-dimensional space, forming a five-dimensional pixel. The five-dimensional pixel is then mapped to the index space in the octree to obtain the five-dimensional space, which is also the index space)). The root node is the pixel with the minimum and maximum values ​​of the five-dimensional space (i.e., the minimum and maximum values ​​of the dimensions, for example: grayscale dimension range [0.1, 0.9], curvature dimension range [0.0, 0.8], texture dimension range [0.2, 0.95], x-coordinate range [0, 1920], y-coordinate range [0, 1080]). Set a maximum splitting depth threshold; simultaneously cut the five-dimensional space according to the root node of each coordinate axis to obtain subspaces (i.e., a traditional three-dimensional octree has 8 subspaces, while this is a five-dimensional "32-ary tree"); Create new child nodes for each subspace, and calculate the bounding box of each child node based on the bounding box of the root node (i.e., the bounding box of each child node is 1 / 32 of the bounding box of its parent node; the volume of the root node's bounding box). The bounding box of each child node is half the length of the root node in each dimension, so the volume of each child node is... The total volume of the 32 child nodes (which is exactly equal to the volume of the root node). Traverse all child nodes of each root node, allocate a subspace according to the coordinate value of each child node, and repeatedly filter the allocated subspace to the child nodes with the minimum and maximum values ​​of the dimensions as the recursive split of the root node (i.e., the subspace is still divided into 32 grandchild nodes in five dimensions). When the number of recursive splits is greater than or equal to the maximum split depth threshold, the recursive splitting stops, and the final child node is output as a leaf node (i.e., each node (whether an intermediate node or a leaf node) contains a bounding box (i.e., the minimum and maximum values ​​in five dimensions, defining the spatial range covered by the node), a list of data points (i.e., references (pointers or indices) of data points contained in the node), a child node pointer (i.e., if it is an intermediate node, it stores pointers to its child nodes), and node statistics (number of points, center point, whether it is a leaf node, etc.)). It should be noted that a weighted 3D initial feature vector is used to adjust the features of each pixel, and combined with its spatial location to generate 5D coordinates, thereby creating a 5D space in the index space. The octree algorithm described above can efficiently index and segment each pixel in the image. The weighted transformation utilizes the importance of different features (such as grayscale, curvature, and texture) to perform detailed spatial indexing of the image, enhancing the identification of disease areas in 5D space while improving computational efficiency. The octree algorithm recursively segments the space, making the processing of each region more precise and avoiding blindly processing the entire image area. The optimized 5D space makes the spatial representation of disease areas more explicit, helping to more accurately locate potential disease areas, especially in complex images. The segmentation and weighting described above can distinguish the contribution of different features to disease areas, improving the sensitivity of disease detection. S253: For each pixel (i.e., coordinates in five-dimensional space), form a boundary circle with a preset neighborhood radius; Determine whether the bounding box of the child node in each recursive split is inside the boundary circle; If the bounding box of a child node is within the boundary circle, then the child node and its corresponding recursively split grandchild nodes are retained (that is, it means that the child node is similar to the pixel, and if the child node is similar, then it also means that the split grandchild nodes of the child node are also similar, so they are retained). If the bounding box of a child node intersects with the boundary circle, then the final leaf node of that child node is directly used as a candidate node (i.e., the candidate node is used as the basis for finding more nodes that may be diseased in the future). For candidate nodes within the boundary circle, find their neighboring child nodes; Calculate the diagonal distance between the candidate node and the bounding boxes of all its adjacent child nodes; Determine whether the diagonal distance is less than a preset neighborhood radius; If so, count all adjacent child nodes and merge them into the boundary circle (that is, because these nodes are very small, even the farthest point is roughly on the order of the diagonal distance from the pixel, and the diagonal distance is much smaller than the preset neighborhood radius, so it is assumed that these points are all within the preset neighborhood radius, that is, within the boundary circle). Cluster all child nodes within the boundary circle to obtain candidate disease areas; It should be noted that the above-mentioned method uses a preset neighborhood radius and coordinates in five-dimensional space to perform boundary circular clustering, which is used to efficiently aggregate similar pixels, merge regions, and obtain candidate disease regions. The above-mentioned method uses a boundary circle formed by the neighborhood radius in five-dimensional space to effectively aggregate similar pixels and segment them using bounding boxes. The recursive splitting method helps to further refine the candidate regions and enhance the distinguishability of disease regions. The clustering process helps to discover more complex disease regions in the image, especially at edges or details. The above-mentioned method can extract clusters of disease regions by merging with neighboring regions, effectively reducing noise and highlighting key disease regions. S254: Perform binarization masking on each candidate disease area to obtain a binary mask area; convert the image to be detected and the binary mask area into the CIELAB color space to extract higher-order color moment features; extract near-infrared features from the binary mask area based on the near-infrared image of the image to be detected; extract depth features from the pixels of the binary mask area based on the depth image; and fuse the higher-order color moment features, near-infrared features, and depth features to obtain a fused encoding vector; It should be noted that within the candidate disease area, the identification of the disease area is enhanced by extracting higher-order moment features of color, near-infrared image features, and depth image features. These features are then fused to form a multimodal feature fusion encoding vector. Color, near-infrared, and depth features can describe the disease area from different perspectives. Higher-order moment of color can capture color distribution features, near-infrared images provide spectral information, and depth images reveal surface morphology. The fusion of these features can provide more comprehensive disease information, which helps improve the accuracy of diagnosis. By fusing different features, diseased and non-diseased areas can be distinguished more accurately. Especially under uneven lighting, occlusion, or other difficult conditions, near-infrared and depth features are particularly crucial and help identify diseases that are usually difficult to detect with a single feature. Specifically, such as Figure 4 As shown, in step S254, a binary mask is applied to each candidate disease region to obtain a binary mask region; the image to be detected and the binary mask region are converted into the CIELAB color space to extract higher-order color moment features; near-infrared features are extracted from the binary mask region based on the near-infrared image of the image to be detected; depth features are extracted from the pixels of the binary mask region based on the depth image; and the higher-order color moment features, near-infrared features, and depth features are fused to obtain a fused encoding vector. The specific operation steps are as follows: S2541: Construct a zero matrix for each candidate disease region with the same size as the image to be detected (i.e., since the candidate disease region is required to be constructed with the same size as the image to be detected, the zero matrix is ​​a binary image with holes and irregular boundaries). Morphological closing operations are performed on the zero matrix to expand and erode the candidate disease region, resulting in a smooth candidate disease region (i.e., morphological closing operations are common knowledge and will not be elaborated further). The smooth candidate disease region is binarized and masked to obtain a binary mask region; It should be noted that for each candidate lesion region, a zero-matrix with the same size as the image to be detected is first constructed, and morphological closing operation is performed. The boundaries of the lesion region are smoothed through dilation and erosion operations. The above-mentioned morphological closing operation removes holes and noise, and makes the lesion region more continuous and smooth, ultimately generating a smooth and regular binary mask region. The role of morphological closing operation in image processing is to eliminate noise and repair irregularly shaped regions, ensuring that the lesion region is well defined. The smoothing of candidate lesion regions can reduce errors caused by irregular lesion boundaries. The smoothed lesion region can avoid misidentification. Because the boundary is clear and there are no holes, the definition of the lesion region is more accurate, which helps to extract relevant features in the subsequent process and reduces false detections. S2542: Convert the image to be detected into the CIELAB color space and extract the LAB components (i.e., for disease detection, the L channel is usually greatly affected by light, while the a and b channels more stably reflect the color attributes of the object itself). Extract the a and b components corresponding to the CIELAB color space of the image to be detected from the binary mask region. Calculate the average gray level and standard deviation of the pixels in component a; The third and fourth moments are calculated using the average and standard deviation of pixel gray levels of the a-component, and are used as the skewness and kurtosis of the a-component. The same calculations as above are performed on the pixels of the b component to obtain the third and fourth moments, which are used as the skewness and kurtosis of the b component. The skewness and kurtosis of component a and component b are used as higher-order color moment features. It should be noted that the third moment is calculated for the average gray value μ_a and the standard deviation σ_a of the pixels in component a, with respect to these values. The formula is as follows: , of which molecules Retaining the sign, positive values ​​skew the distribution to the right, negative values ​​skew the distribution to the left, and the denominator... To eliminate the effects of scale and dimensions, average the result by dividing by N (i.e., the number of pixels); where This is represented as the a component value of the i-th pixel; This represents the average value of the a-component of all pixels in the CIELAB color space of the image to be detected within the binary mask region. A positive skewness of the a-component indicates a right-side long tail (more high a-value pixels, with the main peak concentrated on the right side; a right-side long tail (positive skewness) may indicate local erythema, with most areas being healthy (greenish), but disease causing a small number of pixels to turn red), while a negative value indicates a left-side long tail (i.e., the main peak is concentrated on the left side; a left-side long tail (negative skewness) may indicate chlorosis, with most areas fading to red / yellow due to disease, but a small number of healthy green pixels remaining). Disease may cause the color distribution to be biased to one side, mainly detecting the asymmetry of the component distribution. The kurtosis of the fourth-order moment a-component is calculated; a positive value indicates a more peaked distribution than normal (more concentrated color), while a negative value indicates a flatter distribution (more dispersed color), and disease may make the color distribution more dispersed. Specifically, the image to be detected is converted to the CIELAB color space, and the a and b components are extracted. The CIELAB space is separated into L (luminance) and a and b (color) channels. The a and b components have a relatively stable relationship with the color of the object itself. The gray average and standard deviation of the a and b components are calculated, and the third and fourth moments are further calculated to obtain skewness and kurtosis, which are used as higher-order moment features of color. Compared with the RGB space, CIELAB is closer to the perception of the human visual system, especially sensitive to color changes, and is suitable for color detection. The a and b channels are relatively stable and can better reflect the color changes of the diseased area. Skewness and kurtosis can reflect the morphological characteristics of color distribution, further helping to identify diseased areas with abnormal color. Skewness can reveal the asymmetry of color distribution, while kurtosis reflects the concentration of color distribution. S2543: Based on the near-infrared image of the image to be detected; The pixels adjacent to the outer edge of the binary mask area are extracted as the healthy background area; The near-infrared intensity values ​​of the corresponding pixels in the healthy background region are extracted from the near-infrared image and used as a health feature vector. Based on the near-infrared image, the near-infrared intensity values ​​of the corresponding pixels in the binary mask region are extracted and used as the mask feature vector. The absorption depth ratio (which reflects the degree of near-infrared light absorption in the diseased area relative to the background, calculated by subtracting the mask feature vector from the health feature vector and dividing by the health feature vector, i.e., the calculated absorption depth ratio between near-infrared intensity values) and the absorption width ratio (which reflects the spatial range or uniformity of the absorption effect in the diseased area, calculated by dividing the mask feature vector by the health feature vector) are calculated using the health feature vector and the mask feature vector. The absorption depth ratio and absorption width ratio are used as near-infrared features; It should be noted that, based on near-infrared images, the near-infrared intensity of candidate disease areas is extracted, and the absorption depth ratio (reflecting the degree of absorption of near-infrared light by the disease area) and absorption width ratio (reflecting the spatial range or uniformity of the absorption effect in the disease area) are calculated. Near-infrared imaging can deeply probe the internal structure of objects, and disease areas often differ from healthy areas in their near-infrared light absorption characteristics. The absorption depth ratio and absorption width ratio provide additional spectral information for the disease area, which can better reveal the nature of the disease. This step helps to capture the different absorption characteristics of near-infrared light in diseased areas; for some diseases (such as rot, pests, etc.), their absorption of near-infrared light may be different from that of healthy areas, thus helping to more accurately distinguish between healthy and diseased areas. S2544: Extract depth values ​​from the pixels in the healthy background region (i.e., obtain the pixel depth values ​​corresponding to the healthy background region based on the depth image, which will not be elaborated further); fit the depth values ​​of the pixels in the healthy background region with the coordinate positions of the pixels to obtain a low-order polynomial surface. The depth difference is calculated by using the depth values ​​of the pixels in the binary mask region and the depth values ​​of the pixels in the low-order polynomial surface (that is, a depth difference greater than 0 indicates a protrusion (which may be a gall or healing tissue), and a depth difference less than 0 indicates a depression (which may be an ulcer or rot); at the same time, the depth difference also represents the difference surface between the healthy area and the mask region (which may be a diseased area), that is, the surface that may be uneven). An ideal healthy branch surface model is pre-constructed; the depth difference is input into the ideal healthy branch surface model for output, and the depth features are obtained (that is, the depth features are mainly the extraction of the depth of the recessed pixels, such as the total volume of the recess, the average depth of the recess, and the depth distribution skewness). It should be noted that the depth values ​​of the healthy background area are extracted, a low-order polynomial surface is constructed, and the depth difference between the candidate diseased area and the healthy area is calculated. The depth difference reflects the difference between the surface of the diseased area and the healthy area, such as depressions (ulcers) or protrusions (insect damage, healing tissue). The depth image can reveal the geometry of the object's surface. Diseased areas usually have significant differences in surface morphology from healthy areas, such as depressions or protrusions. By comparing the depth differences, the disease type can be further distinguished. The low-order polynomial surface fitting can simulate the ideal surface of the healthy area, providing a benchmark for the depth difference calculation and enhancing the accuracy of the depth features. This step can reveal the morphological differences of the diseased area, especially for disease types such as ulcers (depressions) or insect damage (protrusions). The depth difference can effectively distinguish between diseased and healthy areas. S2545: Normalize and fuse the color high-order moment features, near-infrared features, and depth features to obtain the fused coding vector of the candidate disease region; It should be noted that the color high-order moment features, near-infrared features, and depth features are normalized and fused into a single fusion coding vector. This coding vector integrates the color, spectral, and depth information of the diseased area, providing a multi-dimensional feature representation. By fusing different types of features (color, near-infrared, and depth), the characteristics of the diseased area can be characterized from multiple dimensions, avoiding the limitations of single features and improving recognition accuracy. Normalization helps to unify feature values ​​at different scales, ensuring that the fused features are processed at the same scale, which helps to improve the fusion effect. The fused encoding vector in the above scheme can comprehensively represent the features of the diseased area, including color, spectrum and geometric shape; through this multimodal fusion, the accuracy of disease identification can be significantly improved, especially in complex backgrounds or when multiple diseases are mixed. Example 2 like Figure 5 As shown, the present invention also provides an image recognition and detection system for disease inspection of fruit tree branch and trunk samples, the system comprising: a data acquisition module 10; an analysis module 20; and a recognition module 30; The acquisition module 10 is used to acquire high-resolution RGB images, near-infrared images, and depth images of the target fruit tree branch surface simultaneously, and to preprocess the high-resolution RGB images to obtain the image to be detected. The analysis module 20 is used to perform grayscale filtering on the pixels of the image to be detected to obtain a grayscale image; select a local window in the grayscale image and perform sliding window calculation on the curvature feature value of the center pixel; encode the pixels in the grayscale image by defining a multi-scale radius to obtain a texture consistency index; concatenate the curvature feature value and the texture consistency index to obtain a three-dimensional initial feature vector set; and use the three-dimensional initial feature vector set with the near-infrared image and the depth image to extract and fuse the encoded vector. The identification module 30 is used to construct a historical fusion coding vector database, match the fusion coding vector with the historical fusion coding vector database, and obtain the fruit tree branch disease detection results. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; those skilled in the art can modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for image recognition and detection of diseases in fruit tree branch and trunk samples, characterized in that, The following steps are included: High-resolution RGB images, near-infrared images, and depth images of the target fruit tree branch and trunk surface are acquired simultaneously. The high-resolution RGB images are preprocessed to obtain the image to be detected. The pixels of the image to be detected are filtered to grayscale to obtain a grayscale image; a local window is selected in the grayscale image and a sliding window is used to calculate the curvature feature value of the center pixel. The pixels in the grayscale image are encoded with a defined multi-scale radius to obtain a texture consistency index; the curvature feature value and the texture consistency index are concatenated to obtain a three-dimensional initial feature vector set; the three-dimensional initial feature vector set is used to extract and fuse encoded vectors with the near-infrared image and the depth image. The fused coding vector includes color high-order moment features, near-infrared features, and depth features; A historical fusion coding vector database is constructed, and the fusion coding vectors are matched with the historical fusion coding vector database to obtain the detection results of fruit tree branch diseases.

2. The method for image recognition and detection of diseases in fruit tree branch and trunk samples according to claim 1, characterized in that, The pixels of the image to be detected are filtered to grayscale to obtain a grayscale image; a local window is selected in the grayscale image to calculate the curvature feature value of the center pixel using a sliding window method. The specific operation steps are as follows: The image to be detected is converted to grayscale, and the grayscale image to be detected is aligned with the depth image. A sliding window is formed on the image to be detected with the pixels of the depth image as the center, and the depth gradient of the x-axis and y-axis is calculated for the pixels of each pair of sliding windows. The gradient magnitude of the pixel is calculated based on the depth gradient of the x-axis and y-axis. Set a baseline standard deviation for the sliding window; calculate the dynamic standard deviation using the baseline standard deviation and the gradient magnitude of the pixel; calculate the grayscale standard deviation between the center pixel and its neighboring pixels within each sliding window; perform bilateral filtering on the pixels of the image to be detected using the grayscale standard deviation of the sliding window and the dynamic standard deviation to obtain smoothed grayscale pixels, which serve as the smoothed grayscale image of the image to be detected. The depth image is selected into a local window and a sliding window is used. Within each local window, the center pixel is used as the local point set. The depth value is extracted from each pixel within the local window using the depth image. The neighborhood pixels of the center pixel in the local window are assigned weight values ​​based on the depth values ​​using MLS fitting. The weight values ​​of the center pixel and the neighborhood pixels are calculated using the least squares method to create a locally fitted surface. Calculate the first and second partial derivatives of the center pixel of the fitted surface; construct the second basic form matrix corresponding to the center pixel using the first and second partial derivatives of the center pixel of the fitted surface. The second basic form matrix is ​​reduced in dimensionality by principal component analysis to obtain the first and second eigenvalues, which are used as principal curvatures. The principal curvatures are then mapped to a preset fixed interval using shape indexing to obtain the curvature feature value of the center pixel of the local window.

3. The method for image recognition and detection of diseases in fruit tree branch and trunk samples according to claim 2, characterized in that, The pixels in the grayscale image are encoded with defined multi-scale radii to obtain a texture consistency index; the curvature feature values ​​and the texture consistency index are concatenated to obtain a three-dimensional initial feature vector set. The specific operation steps are as follows: Define a multi-scale radius for the smoothed grayscale image of the image to be detected, and collect n sampling points on the circumferential trajectory of the circle containing each multi-scale radius; The center gray value of the circle with the radius is compared with the gray values ​​of the n sampling points to obtain an n-bit binary number; the n-bit binary number is concatenated end to end, and the number of transitions for each adjacent binary number is counted. Set a preset uniform threshold; determine whether the number of transitions is less than the uniform threshold; If so, then the radius circle under the multi-scale radius is determined to be a uniform pattern; The n-bit binary number of the uniform pattern is rotated and normalized to convert it into the smallest binary number. RIU-LBP encoding is performed on each sampling point in the multi-scale radius to construct a multi-scale histogram. The entropy value is calculated for each multi-scale histogram, and the entropy values ​​of all multi-scale histograms are weighted and fused to obtain the texture consistency index for each sampling point. The gray values ​​of the smoothed grayscale pixels are normalized and concatenated with the curvature feature values ​​and texture consistency index to obtain a set of three-dimensional initial feature vectors.

4. The method for image recognition and detection of diseases in fruit tree branch and trunk samples according to claim 1, characterized in that, The fusion encoding vector is extracted and fused using the aforementioned initial 3D feature vector set, near-infrared image, and depth image. The specific steps are as follows: A parameter dimension is set for the three-dimensional initial feature vector set, and an initial parameter set for the three-dimensional initial feature vector set is obtained according to the parameter dimension; the octree algorithm and the initial parameter set are used to map each pixel in the image to be detected to the index space, and a five-dimensional space of the octree is constructed. The root node is selected from the pixels in the five-dimensional space, and the root node is recursively split to form child nodes. For each pixel, a boundary circle is formed with a preset neighborhood radius, and all child nodes are clustered to obtain candidate disease regions. The fusion encoding vector is obtained by combining the features extracted from the near-infrared image and the depth image of the candidate disease regions.

5. The method for image recognition and detection of diseases in fruit tree branch and trunk samples according to claim 1, characterized in that, A parameter dimension is set for the three-dimensional initial feature vector set, and an initial parameter set is obtained based on the parameter dimension. Each pixel in the image to be detected is mapped to an index space using an octree algorithm and the initial parameter set, constructing a five-dimensional octree space. The pixels in the five-dimensional space are then filtered to select the root node, and the root node is recursively split to form child nodes. The specific operation steps are as follows: The x-axis, y-axis, and z-axis parameters are set for the three-dimensional initial feature vector set; the search range and non-uniform search step size are set for each parameter dimension. The search range of the parameter dimension is divided into search intervals using the Latin hypercube sampling method; each search interval is randomly and uniformly sampled according to the non-uniform search step size to obtain multi-dimensional sampling points; The coordinate position of each multi-dimensional sampling point is determined, and the initial parameter set of each search interval is determined based on the coordinate position of each multi-dimensional sampling point; The octree is used to perform a weighted transformation on the three-dimensional initial feature vector of each pixel in the image to be detected using the parameter dimension of the z-axis. Then, the coordinates of the x-axis and y-axis of each pixel in the initial parameter set are further combined to map to the index space, thus constructing the five-dimensional space of the octree. The root node is selected from the pixels with the minimum and maximum values ​​of the five-dimensional space. Set a maximum splitting depth threshold; simultaneously cut the five-dimensional space according to the root node of each coordinate axis to obtain subspaces; Create a new child node for each subspace, and calculate the bounding box of each child node based on the bounding box of the root node; Traverse all child nodes of each root node, allocate a subspace based on the coordinate value of each child node, and repeatedly filter the allocated subspaces to select the child nodes with the minimum and maximum dimensions as the root node for recursive splitting; when the number of recursive splits is greater than or equal to the maximum split depth threshold, stop the recursive splitting and output the final child node as the leaf node.

6. The method for image recognition and detection of diseases in fruit tree branch and trunk samples according to claim 5, characterized in that, For each pixel, a boundary circle is formed with a preset neighborhood radius. All child nodes are then clustered to obtain candidate disease regions. The specific operation steps are as follows: For each pixel, a boundary circle is formed with a preset neighborhood radius; Determine whether the bounding box of the child node in each recursive split is inside the boundary circle; If the bounding box of a child node is within the boundary circle, then the child node and its corresponding recursively split grandchild nodes are preserved. If the bounding box of a child node intersects with the boundary circle, then the final leaf node of that child node is directly used as a candidate node; for the candidate node within the boundary circle, the neighboring child nodes are searched. Calculate the diagonal distance between the candidate node and the bounding boxes of all its adjacent child nodes; Determine whether the diagonal distance is less than a preset neighborhood radius; If so, count all adjacent child nodes and merge them into the boundary circle; Cluster all child nodes within the boundary circle to obtain candidate disease areas.

7. The method for image recognition and detection of diseases in fruit tree branch and trunk samples according to claim 6, characterized in that, The fused encoding vector is obtained by combining the features extracted from the candidate disease region with the near-infrared image and the depth image. The specific operation steps are as follows: A binary mask is applied to each candidate disease region to obtain a binary mask region; the image to be detected and the binary mask region are converted into the CIELAB color space to extract higher-order color moment features; near-infrared features are extracted from the binary mask region based on the near-infrared image of the image to be detected; depth features are extracted from the pixels of the binary mask region based on the depth image; and the higher-order color moment features, near-infrared features, and depth features are fused to obtain a fused encoding vector.

8. The method for image recognition and detection of diseases in fruit tree branch and trunk samples according to claim 7, characterized in that, A binary mask is applied to each candidate disease region to obtain a binary mask region; the image to be detected and the binary mask region are converted into the CIELAB color space to extract higher-order color moment features. The specific operation steps are as follows: For each candidate disease region, a zero matrix with the same size as the image to be detected is constructed using pixels; morphological closing operations are performed on the zero matrix, and dilation and erosion operations are performed on the candidate disease regions to obtain smooth candidate disease regions; The smooth candidate disease region is binarized and masked to obtain a binary mask region; The image to be detected is converted to the CIELAB color space, and the LAB components are extracted. Extract the a and b components corresponding to the CIELAB color space of the image to be detected from the binary mask region. The average gray value and standard deviation of the pixels of the a component are calculated; the third moment and fourth moment are calculated using the calculated average gray value and standard deviation of the pixels of the a component, which are used as the skewness and kurtosis of the a component. The same calculations as above are performed on the pixels of the b component to obtain the third and fourth moments, which are used as the skewness and kurtosis of the b component. The skewness and kurtosis of component a and component b are used as higher-order color moment features.

9. The method for image recognition and detection of diseases in fruit tree branch and trunk samples according to claim 8, characterized in that, Near-infrared features are extracted from the binary mask region based on the near-infrared image of the image to be detected; depth features are extracted from the pixels of the binary mask region based on the depth image; and the higher-order color moment features, near-infrared features, and depth features are fused to obtain a fused encoding vector. The specific operation steps are as follows: Based on the near-infrared image of the image to be detected; The pixels adjacent to the outer edge of the binary mask region are extracted as a healthy background region; the near-infrared intensity values ​​of the corresponding pixels in the healthy background region are extracted based on the near-infrared image as a health feature vector; Based on the near-infrared image, the near-infrared intensity values ​​of the corresponding pixels in the binary mask region are extracted and used as the mask feature vector. The absorption depth ratio and absorption width ratio are calculated using the health feature vector and the mask feature vector; the absorption depth ratio and absorption width ratio are then used as near-infrared features. Depth values ​​are extracted from pixels in the healthy background region; the depth values ​​of pixels in the healthy background region are fitted with the coordinate positions of the pixels to obtain a low-order polynomial surface; the depth difference is calculated by using the depth values ​​of pixels in the binary mask region and the depth values ​​of pixels in the low-order polynomial surface. A surface model of an ideal healthy branch is pre-constructed; the depth difference is input into the surface model of the ideal healthy branch for output, and the depth feature is obtained. The color high-order moment features, near-infrared features, and depth features are normalized and fused to obtain the fused coding vector of the candidate disease region.

10. A system for image recognition and detection of diseases in fruit tree branch and trunk samples, characterized in that, include: Data acquisition module; Analysis module; Recognition module; The acquisition module is used to acquire high-resolution RGB images, near-infrared images, and depth images of the target fruit tree branch surface simultaneously, and to preprocess the high-resolution RGB images to obtain the image to be detected. The analysis module is used to perform grayscale filtering on the pixels of the image to be detected to obtain a grayscale image; and to select a local window in the grayscale image to perform sliding window calculation on the curvature feature value of the center pixel. The pixels in the grayscale image are encoded with a defined multi-scale radius to obtain a texture consistency index; the curvature feature value and the texture consistency index are concatenated to obtain a three-dimensional initial feature vector set. The 3D initial feature vector set is used to extract and fuse encoded vectors from near-infrared images and depth images; The identification module is used to construct a historical fusion coding vector database, and to match the fusion coding vector with the historical fusion coding vector database to obtain the detection results of fruit tree branch diseases.