Intelligent grain sampling method based on multi-scale fusion verification improved SGBM
Through improved SGBM algorithm and binocular vision technology, the scientificity and representativeness of intelligent food sampling are achieved, the problems of inefficiency and insufficient intelligence of traditional manual sampling methods are solved, and sampling accuracy and transparency are improved.
Patent Information
- Application Number
- CN202510488836.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
AI Technical Summary
The traditional manual sampling method has problems such as uneven distribution of sampling points, lack of standardization, low operation transparency and insufficient intelligent decision-making. The existing automation equipment cannot perceive changes in the grain pile shape in real time, resulting in unrepresentative sampling.
Using the improved SGBM algorithm based on multi-scale fusion verification, combined with binocular vision technology and YOLOv8-seg model, a parallax map is generated through image segmentation and stereo matching, and the three-dimensional spatial position of the sample points is calculated, and the sampler is driven to achieve fully automatic random sampling.
The scientificity and representativeness of bulk food sampling are achieved, the sampling efficiency and accuracy are improved, the risks of artificial loss and fraud are reduced, and the robustness and applicability of the algorithm are enhanced.
Smart Images

Figure CN120411208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for intelligent sampling of grain based on multi-scale fusion verification and improved SGBM. Background Art
[0002] Food security, as an important part of the national security system, the quality of its reserves is directly related to national strategic stability. In the grain purchase and storage link, sampling and testing is a key technical node to ensure the quality of grain, and the scientificity and representativeness of its sampling results directly affect the accuracy of quality assessment. The traditional manual sampling method mainly relies on the experience of operators and adopts the mode of sampling at fixed points with a probe rod, which has three major technical bottlenecks: First, manual operation is likely to lead to uneven distribution of sampling points, making it difficult to achieve full coverage detection of the three-dimensional space of the whole vehicle of grain pile, and problems such as deep mildew and pests are easily overlooked; Second, the operation process lacks standardization, and there are subjective deviations in parameters such as sampling depth and angle, resulting in insufficient comparability of detection data; Third, the transparency of the operation process is low, and there is a moral risk of artificially swapping samples or selective sampling. Although existing automated sampling equipment has partly solved the mechanization problem, there are still significant defects at the level of intelligent decision-making - it cannot perceive the morphological changes of the grain pile in real time, resulting in a lack of dynamic adaptability in the motion trajectory planning of the robotic arm. Especially when facing the morphological randomness characteristics of bulk grain piles, the fixed programmed operation mode is difficult to ensure the statistical representativeness of sampling. In recent years, the development of computer vision technology has provided a new path to solve this problem. Binocular vision technology locates bulk grain through a stereo matching algorithm and combines it with a sampling machine for fully automatic sampling, providing technical support for building a more reliable food security system. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for intelligent sampling of grain based on multi-scale fusion verification and improved SGBM to solve the problems of low efficiency and lack of objectivity in traditional manual sampling.
[0004] To achieve the above purpose, the solution of the present invention includes:
[0005] A method for intelligent sampling of grain based on multi-scale fusion verification and improved SGBM, comprising the following steps:
[0006] (1) Collect the left and right images of the bulk grain transport vehicle through a binocular camera;
[0007] (2) Use the YOLOv8-seg model to perform image segmentation on the collected images and extract the grain area in the carriage image;
[0008] (3) Divide the identified grain surface area to form multiple candidate sampling areas, and select multiple sampling areas from the candidate areas according to a preset random rule;
[0009] (4) Preprocess the extracted grain area image to optimize and prepare the image data;
[0010] (5) Use the improved SGBM algorithm to recover depth information from the left and right images and generate an optimized disparity map;
[0011] (6) Calculate the three-dimensional spatial positions of the sampling points in each random sampling area according to the optimized disparity map, convert the position information into control instructions, and drive the sampling machine to achieve fully automatic random sampling.
[0012] Further, in step (2), the YOLOv8-seg model is used to perform instance segmentation on the corrected left-eye image, extract the grain area in the carriage image, and combine the corner detection algorithm to extract the boundary feature points of the target area.
[0013] Further, in step (3), according to the national standard "Inspection Sampling and Subsampling Methods for Grain and Oil Seeds", the segmented grain area is divided to randomly generate multiple alternative sampling areas and sampling points.
[0014] Further, in step (4), the preprocessing of the extracted grain area image specifically includes image distortion correction, image registration, and noise removal.
[0015] Further, step (5) includes the following sub-steps:
[0016] S1: Perform image correction and noise suppression on the left and right images;
[0017] S2: Use the image block comparison method to calculate the matching cost of each pixel point;
[0018] S3: Use the dynamic programming method to aggregate the matching cost, and determine the disparity for each pixel point according to the minimum cost principle to generate a preliminary disparity map;
[0019] S4: Post-process the preliminary disparity map, including uniqueness detection, sub-pixel interpolation, left-right consistency detection, and connected region detection, to generate a preliminary optimized disparity map;
[0020] S5: Perform multi-scale fusion verification on the preliminary optimized disparity map, combine multi-scale strategies, disparity map fusion, and filtering collaborative verification to obtain the final optimized disparity map.
[0021] Further, the multi-scale fusion verification in step S5 includes the following steps:
[0022] S51: Resize the original image multiple times to construct a multi-scale image pyramid, and calculate the initial disparity map of the image at multiple resolutions respectively;
[0023] S52: Restore the disparity maps at different scales to the original resolution through interpolation, and perform weighted averaging on the disparity values of all scales at each pixel position to generate a unified mean disparity map;
[0024] S53: Perform optimization processing on the mean disparity map by introducing weighted least squares (WLS) filtering and guided filtering respectively to generate two sets of smoothed disparity maps.
[0025] S54: Perform a multi-scale collaborative verification operation on the two sets of smoothed disparity maps obtained in S53. Use the WLS-filtered disparity map as the main reference and the guided-filtered disparity map as the auxiliary verification source to perform consistency detection based on pixel-level difference analysis.
[0026] Further, in step S54, the WLS filtering result and the guided filtering result are collaboratively verified. Specifically, for each pixel point, obtain the disparity value after WLS filtering and the disparity value after guided filtering respectively, and calculate the difference between the two. If the difference is less than the set threshold and both of their disparity values are greater than 0, then take the average of the two as the final disparity value; otherwise, retain the WLS result as the final disparity value to generate an optimized final disparity map, as shown in the following formula:
[0027]
[0028] where D o represents the finally optimized disparity map, D wls (x, y) is the disparity value after WLS filtering, D g (x, y) is the disparity value after guided filtering, and T is the set threshold.
[0029] Further, the finally optimized disparity map is used to calculate the three-dimensional spatial positions of the sampling points in each randomly sampled area, and convert the three-dimensional coordinate information into instructions to drive the sampling machine to perform automated sampling operations.
[0030] The beneficial effects of the present disclosure are:
[0031] (1) The present invention combines binocular vision positioning and image segmentation technologies to realize the scientific selection of random sampling points for bulk grain, taking into account both the representativeness of samples (fully reflecting the quality level) and efficiency (reducing excessive sampling), reducing grain loss and the risk of human fraud. Through binocular vision technology, three-dimensional positioning of the on-vehicle grain pile image is performed to improve the spatial positioning accuracy of the sampling points, providing technical support for the automation and intelligence of grain sampling.
[0032] (2) Improve the quality of the disparity map: Through multi-scale fusion verification based on the image feature pyramid, the present disclosure effectively solves the problem of inaccurate disparity matching in the SGBM stereo matching algorithm in low-texture regions and cases of repetitive patterns. The mean disparity map generated by fusing different-level features extracted by the multi-scale strategy significantly reduces the discontinuous regions in the disparity map, thereby improving the quality of the disparity map.
[0033] (3) Enhance the accuracy of depth estimation: By introducing WLS filtering and guided filtering, the present disclosure further optimizes the smoothness and consistency of the disparity map, reducing the influence of noise and outliers on depth estimation. This method can more accurately estimate depth information when dealing with complex scenes, improving the overall accuracy of depth estimation.
[0034] (4) Improve the robustness of the algorithm: The combination of multi-scale fusion verification and filtering collaborative verification makes the algorithm show higher robustness when processing images under different textures and lighting conditions. This improvement enables the algorithm to maintain a high matching accuracy in various practical application scenarios, enhancing its applicability and reliability. Brief Description of the Drawings
[0035] Figure 1 is a flowchart of an intelligent grain sampling method based on improved SGBM with multi-scale fusion verification
[0036] Figure 2 is a flowchart of the improvement of the key steps of an intelligent grain sampling method based on improved SGBM with multi-scale fusion verification Detailed Embodiment
[0037] As Figure 1 shown, first, binocular cameras are used to synchronously collect the left-eye image and the right-eye image of the in-vehicle bulk grain area. Through the undistort() function of OpenCV or by combining initUndistortRectifyMap() and remap() functions, the radial distortion and tangential distortion of the lens are eliminated to complete distortion correction. Binocular stereo rectification is to convert the original non-coplanar row-aligned images into an ideal binocular system with coplanar row alignment, ensuring that the corresponding epipolar lines of the left and right images are horizontally aligned, providing geometric consistency constraints for subsequent stereo matching.
[0038] Subsequently, the YOLOv8seg model is used to perform instance segmentation on the rectified left-eye image. Its basic structure consists of three parts: The Backbone is based on CSPDarknet to achieve efficient feature extraction, the Neck performs multi-scale feature fusion through PANet, and the Head contains a detection head and a segmentation head, which respectively output the target bounding box, class confidence, and instance mask. After segmentation, the corner coordinates of the target area are extracted based on the Shi-Tomasi algorithm, providing spatial constraints for stereo matching.
[0039] Divide the target bulk grain area segmented from the network model to form multiple candidate sampling areas, and select multiple sampling areas from the candidate areas according to a preset random rule. Analyze the left and right views using the improved stereo matching algorithm of the present disclosure to generate a disparity map. A disparity map is a grayscale image, where each pixel point contains the distance information from the camera. It is a commonly used image representation method in computer vision for describing the three-dimensional structure of a scene.
[0040] Stereo matching is the key step in generating a disparity map. It is a technique for restoring the depth information of a real scene based on a planar image. The method is to find matching point pairs from two or more images of the same scene, and then calculate the depth of the spatial physical points corresponding to the point pairs according to the principle of triangulation. The present disclosure improves on the SGBM stereo matching algorithm, such as Figure 2 , and its calculation process mainly includes the following steps:
[0041] (1) Preprocessing, the purpose of which is to enhance the structural features of the image and extract gradient information, providing a robust feature input for cost calculation in stereo matching. The SGBM algorithm uses a horizontal Sobel operator to convolve the image to capture the gradient response in the horizontal direction, highlighting edges and texture details; subsequently, a mapping function is used to perform a non-linear mapping on the gradient magnitude to suppress noise interference and compress the dynamic range, finally generating an optimized feature map suitable for matching cost calculation. The mapping function is as follows:
[0042]
[0043] (2) Cost calculation, the core of matching cost calculation lies in evaluating the similarity between the pixel to be matched and the candidate pixels. This calculation usually includes two parts: one is the gradient information extracted based on image preprocessing, and the other is the gradient cost obtained through a sampling method. To improve the matching accuracy, the SAD (Sum of Absolute Differences) cost and the BT (Block Matching) method are combined, and the SAD-BT cost is calculated using neighborhood summation. This method incorporates local region information into the cost calculation, enhancing the accuracy and robustness of the matching.
[0044] (3) Dynamic programming. In traditional dynamic programming algorithms, the "trailing effect" in disparity optimization easily leads to incorrect matching in disparity mutation regions. The fundamental reason is that one-dimensional energy accumulation will spread the incorrect disparity on the path to subsequent regions. To solve this problem, the SGBM algorithm proposes a multi-direction one-dimensional path constraint strategy: by independently performing dynamic programming energy accumulation in multiple directions (such as horizontal, vertical, diagonal), constructing a global Markov energy equation, and superimposing the matching costs of paths in each direction as the total pixel cost, thereby dispersing the error impact of a single path. The final disparity is determined by directly selecting the minimum value from the total cost using the WTA (Winner Takes All) strategy. This method significantly suppresses the matching errors at disparity jumps through multi-path global optimization while retaining the efficiency advantages of dynamic programming.
[0045] (4) Post-processing of the disparity map. The post-processing of SGBM includes uniqueness detection, sub-pixel interpolation, left-right consistency detection, and connected region detection, aiming to optimize the quality of the disparity map. Through these steps, the spatial smoothness of the disparity map can be enhanced, and noise and discontinuities can be reduced. The smoothing process helps to eliminate local anomalies caused by matching errors or image noise, thereby improving the continuity and consistency of disparity values.
[0046] (5) Multi-scale fusion based on the feature pyramid to optimize the disparity. First, a multi-scale strategy is adopted to extract different-level features of the original disparity map and fuse them to calculate the mean disparity map. Subsequently, WLS filtering and guided filtering are introduced to optimize the fused disparity map. Finally, the pixel point differences are co-verified through filtering, and the optimal disparity set is selected as the final disparity map. The specific steps are as follows:
[0047] S1: Multi-scale strategy. It processes the image at different resolutions to extract different-level feature information, which is especially suitable for tasks with diverse features or scale changes. Among them, the image feature pyramid is a typical implementation method. It realizes the extraction and fusion of different-scale information by constructing multiple levels of feature maps. Its core idea is that images at different resolutions contain different levels of information: low-resolution images are helpful for capturing global information, while high-resolution images can retain more details. Therefore, in the multi-scale feature pyramid, each resolution level can provide different feature information. This disclosure uses the Gaussian pyramid model to optimize disparity calculation. The bottom layer of the Gaussian pyramid is the original image. Each time it goes up one layer, the image size is reduced by downsampling. Usually, the length and width of the image are reduced to half of the original. The common number of layers of the Gaussian pyramid is 3 - 6. Based on this, this disclosure performs two downsamplings on the original image to obtain feature maps with different resolutions and fuse this feature information to improve the accuracy of disparity calculation.
[0048] S2: Parallax map fusion. The core idea is to calculate the parallax map at different scales and fuse it through weighted average or mean processing to generate the final parallax map. This method can effectively smooth the parallax map and reduce local noise, thereby improving the accuracy of depth estimation.
[0049] Specifically, first, the original image is scaled twice to generate images with different resolutions, and the corresponding parallax maps are calculated at three resolutions respectively. Subsequently, the interpolation algorithm is used to restore the low-resolution parallax map to the original resolution. During this process, the interpolation operation can introduce smooth transitions between pixels, reduce sudden changes, thereby enhancing the continuity of the parallax map and reducing the noise impact that may be brought by low-resolution images. Finally, the parallax maps with different resolutions are weighted-averaged at each pixel position to fuse multi-scale information, eliminate local inconsistencies, and make the finally generated parallax map smoother and more consistent. The formula is described as follows:
[0050]
[0051] where D avg (i,j) represents the final parallax value of the fused parallax map, and N represents the number of parallax maps.
[0052] S3: Post-processing after fusion. Guided Filtering, Bilateral Filtering (BF), and Weighted Least Squares Filtering (WLS) are three common Edge-Preserving filtering methods. While smoothing the image and removing noise, they can effectively retain the edge information of the image, avoid edge blurring, thereby improving the image quality and clarity. In this disclosure, Guided Filtering and Weighted Least Squares Filtering are selected for processing.
[0053] Specifically, in the post-processing stage of the SGBM algorithm, first, multiple parallax maps are fused, and Weighted Least Squares Filtering (WLS) is introduced to further optimize the smoothness of the parallax map. Weighted Least Squares Filtering (WLS) keeps the structural information in the edge area while performing smoothing processing by constraining the output image to be as close as possible to the input image, thereby improving the quality of the parallax map and the accuracy of depth estimation. The loss function formula is described as follows:
[0054]
[0055] where the original image is g, p is the pixel point coordinate, the filtering result to be solved is u, a x 、a y are the weight matrices of the gradients in the x and y directions respectively. The first term of the function (u p -g P) represents the similarity between the input image and the output image. The second term is a regularization term. By minimizing the partial derivative, the smoothness of the output image is enhanced. a x,p (g), a y,p (g) is the weight coefficient, and the formula is described as follows:
[0056]
[0057] Where, l represents log, ε is to prevent the denominator from being zero, generally taking a very small value of 0.0001, and α represents the sensitivity of the gradient.
[0058] Furthermore, the present disclosure further applies guided filtering to the processed mean disparity map. Guided filtering filters the original image through a guiding image to achieve the effects of smoothing, denoising, and detail enhancement. This method can maintain the edge details of the image while smoothing large homogeneous areas, thereby improving the quality of the disparity map.
[0059] The core principle of guided filtering is to use the information of neighboring pixels at each pixel position to predict the value of the output pixel, thereby achieving a balance between smoothing and edge preservation. Specifically, this method assumes that the output image and the input image satisfy a linear relationship within a local window, and the formula is described as follows:
[0060]
[0061] Where, q is the value of the output pixel, I is the value of the input pixel, i and k are pixel indices, and a and b are the coefficients of this linear function when the window center is at k.
[0062] S4: Filtering collaborative verification. To improve the quality of the disparity map, reduce noise, and correct the mis-matched regions, the present disclosure introduces a filtering collaborative verification strategy. This method combines the results of WLS filtering and guided filtering, and through difference analysis, ensures that the optimization process can smooth the disparity map and avoid over-correction.
[0063] Specifically, first, the disparity values after WLS filtering and the disparity values after guided filtering are respectively obtained, and the difference between them is calculated. If the difference is less than the set threshold and the disparity values of both are greater than zero, it indicates that the matching results of this pixel point are relatively consistent. At this time, the average value of the two is taken for update to achieve the effects of smoothing and denoising. If the difference exceeds the threshold, it indicates that there is a large deviation between the two. At this time, the disparity value of WLS filtering is retained to avoid over-smoothing leading to the correction of matching errors, thereby ensuring the reliability of the disparity map. As shown in the following formula:
[0064]
[0065] Where, D o represents the finally optimized disparity map, D wls(x, y) is the disparity value after WLS filtering, D g (x, y) is the disparity value after guided filtering, and T is the set threshold.
[0066] Finally, the binocular vision stereo matching algorithm is used to calculate the disparity map, and the three-dimensional coordinates of each sampling point are deduced based on the disparity information, so as to control the grain sampler to accurately move to the specified sampling point position and complete the automated sampling operation.
[0067] To further verify the effectiveness of the optimized algorithm of the present disclosure, the present disclosure selects to measure the length and height of the simulated grain truck, compare the measurement data of the two algorithms with the true values, and then conduct error analysis. Since the matching effect of the stereo matching algorithm also varies at different distances, the size of the grain truck is measured in this experiment from about 1 m and 1.5 m away from the left camera respectively, and the measurement results are shown in the following table.
[0068]
[0069] The experimental results show that the SGBM algorithm based on multi-scale fusion is superior to the ordinary SGBM algorithm at both 1 m and 1.5 m heights. For the height of 1 m, the length and height of the truck measured by the multi-scale fusion algorithm are reduced by 1.62 and 1.79 percentage points respectively compared with the traditional SGBM algorithm. At the height of 1.5 m, the dimensions of the carriage measured by the multi-scale fusion algorithm are reduced by 1.5 and 1.35 percentage points respectively compared with the traditional SGBM algorithm. The overall error range of the SGBM algorithm measurement is between 2.97% and 6.92%, and the overall error range of the multi-scale fusion algorithm for measuring the grain truck size is 1.35% to 5.57%. As the distance between the camera and the grain truck increases, the errors of both methods increase, but at this time the error of the multi-scale fusion method is still lower than that of the SGBM algorithm. Considering the comprehensive results, the error is lower when using the optimized stereo matching algorithm to measure the grain truck data, meeting the requirements of the present disclosure.
[0070] The specific implementation manners are given above, but the present invention is not limited to the described implementation manners. The basic idea of the present invention lies in the above basic scheme. For those of ordinary skill in the art, according to the teachings of the present invention, it does not require creative labor to design various deformed models, formulas, and parameters. Changes, modifications, substitutions, and variations made to the implementation manners without departing from the principles and spirit of the present invention still fall within the protection scope of the present invention.
Claims
1. An intelligent grain sampling method based on multi-scale fusion verification and improved SGBM, characterized in that It includes the following steps: (1) Collect the left and right images of the bulk grain transport vehicle through a binocular camera; (2) Use the YOLOv8-seg model to perform image segmentation on the collected images and extract the grain area in the carriage image; (3) Divide the identified grain surface area to form multiple candidate sampling areas, and select multiple sampling areas from the candidate areas according to a preset random rule; (4) Preprocess the extracted grain area image to optimize the preparation of image data; (5) Use the improved SGBM algorithm to recover depth information from the left and right images and generate an optimized disparity map; (6) Calculate the three-dimensional spatial positions of the sampling points in each random sampling area according to the optimized disparity map, convert the position information into control instructions, and drive the sampling machine to achieve fully automatic random sampling.
2. The intelligent grain sampling method based on improved SGBM with multi-scale fusion verification according to claim 1, wherein: In step (2), the YOLOv8-seg model is used to perform instance segmentation on the corrected left-eye image, extract the grain area in the carriage image, and combine the corner detection algorithm to extract the boundary feature points of the target area.
3. The improved grain intelligent sampling method based on multi-scale fusion verification of SGBM according to claim 1, characterized in that: In step (3), according to the national standard "Inspection and Sampling Methods for Grain and Oil Seeds", the divided grain area is divided to randomly generate multiple alternative sampling areas and sampling points.
4. The intelligent grain sampling method based on multi-scale fusion verification and improved SGBM according to claim 1, wherein: In step (4), the extracted grain area image is preprocessed, specifically including image distortion correction, image registration, and noise removal.
5. The intelligent grain sampling method based on multi-scale fusion verification and improved SGBM according to claim 1, characterized in that: Step (5) includes the following sub-steps: S1: Perform image correction and noise suppression on the left and right images; S2: Use the image block comparison method to calculate the matching cost of each pixel point; S3: Use the dynamic programming method to aggregate the matching cost, and determine the disparity for each pixel point according to the minimum cost principle to generate a preliminary disparity map; S4: Post-process the preliminary disparity map, including uniqueness detection, sub-pixel interpolation, left-right consistency detection, and connected region detection, to generate a preliminary optimized disparity map; S5: Perform multi-scale fusion verification on the preliminary optimized disparity map, combine multi-scale strategies, disparity map fusion, and filtering collaborative verification to obtain the final optimized disparity map.
6. The improved intelligent grain sampling method based on multi-scale fusion verification of SGBM according to claim 5, characterized in that: The multi-scale fusion verification in step S5 includes the following steps: S51: Scale the original image multiple times respectively to construct a multi-scale image pyramid, and calculate the initial disparity map of the image at multiple resolutions respectively; S52: Restore the disparity maps at different scales to the original resolution by interpolation method, and perform weighted average on the disparity values at all scales at each pixel position to generate a unified mean disparity map; S53: Optimize the mean disparity map by introducing weighted least squares (WLS) filtering and guided filtering respectively to generate two groups of smoothed disparity maps; S54: Perform multi-scale collaborative verification operations on the two groups of smoothed disparity maps obtained in S53. Taking the WLS filtering disparity map as the main reference and the guided filtering disparity map as the auxiliary verification source, perform consistency detection based on pixel-level difference analysis.
7. The intelligent grain sampling method based on multi-scale fusion verification and improved SGBM according to claim 6, characterized in that: In step S54, the WLS filtering result and the guided filtering result are collaboratively verified, which specifically includes: for each pixel point, the disparity value after WLS filtering and the disparity value after guided filtering are respectively obtained, and the difference between the two is calculated. If the difference is less than the set threshold and the disparity values of both are greater than 0, then the average value of the two is taken as the final disparity value; otherwise, the WLS result is retained as the final disparity value to generate an optimized final disparity map, as shown in the following formula: Among them, D o represents the final optimized disparity map, D wls (x, y) is the disparity value after WLS filtering, D g (x, y) is the disparity value after guided filtering, and T is the set threshold.
8. The intelligent grain sampling method based on multi-scale fusion verification and improved SGBM according to claim 1, characterized in that: The final optimized disparity map is used to calculate the three-dimensional spatial positions of the sampling points in each randomly sampled area, and the three-dimensional coordinate information is converted into instructions to drive the sampling machine to perform automatic sampling operations.