A rice and wheat lodging three-dimensional attribute detection and distribution map construction method for a combine harvester, an electronic device, and a computer readable storage medium
By combining binocular vision and deep learning technologies with the pose information of combine harvesters, real-time and accurate detection and three-dimensional attribute extraction of lodged areas of rice and wheat were achieved, solving the problem of inaccurate detection of lodged rice and wheat in existing technologies and improving the operating efficiency and quality of combine harvesters.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU UNIV
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-29
AI Technical Summary
Existing detection methods cannot achieve accurate detection, three-dimensional attribute extraction, and multi-frame information fusion of lodged areas in rice and wheat under complex field conditions, which increases the difficulty and reduces the efficiency of combine harvester operations.
A binocular vision system combined with deep learning technology was used to identify lodged areas of rice and wheat and extract 3D point cloud data through a pre-trained neural network. Multi-frame data fusion was performed by combining the pose information of the combine harvester to construct a 3D attribute distribution map of lodged rice and wheat.
It enables real-time and accurate detection and three-dimensional attribute extraction of lodged areas of rice and wheat, reduces false detection and missed detection rates, provides stable operational data support, and improves the operational efficiency and quality of combine harvesters.
Smart Images

Figure CN122115394A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural intelligent detection technology, specifically to a method for detecting and constructing three-dimensional attributes and distribution maps of rice and wheat lodging for combine harvesters. Background Technology
[0002] Lodging of rice and wheat not only disrupts the normal growth structure of the plants, reduces photosynthetic efficiency, and leads to yield reduction, but also significantly increases the difficulty of combine harvester operation, causing poor header feeding, decreased threshing efficiency, and consequently, increased grain loss and reduced harvesting quality. Therefore, accurately determining the location and condition of lodged areas before or during combine harvester operation is crucial for adjusting operating parameters and improving harvesting efficiency.
[0003] Existing detection methods have the following limitations: manual surveys are inefficient and highly subjective; satellite remote sensing has insufficient resolution, and UAV remote sensing is difficult to operate in real-time and continuously; vision-based detection methods are mostly validated in controlled environments and are not adaptable to complex field environments (such as changes in lighting, occlusion, and background interference), and generally lack the ability to extract refined three-dimensional attributes (position, angle, and direction) of lodging. In addition, existing methods mostly analyze single-frame images, without considering the pose changes of combine harvesters during continuous operation, making it difficult to achieve spatial alignment and fusion of multi-frame detection results, and thus unable to form a global lodging distribution map at the field scale.
[0004] Therefore, a solution is needed that can adapt to complex field environments, support continuous operation, and achieve accurate detection of lodging areas, 3D attribute extraction, and multi-frame information fusion. This invention aims to integrate binocular vision and deep learning technologies to achieve real-time detection of 3D attributes and construction of spatial distribution maps for rice and wheat lodging, providing a reliable basis for optimizing combine harvester operations. Summary of the Invention
[0005] The technical problem to be solved by this application is to overcome the shortcomings of the prior art and provide a method for detecting and constructing three-dimensional attributes of rice and wheat lodging for combine harvesters, an electronic device, and a computer-readable storage medium.
[0006] Firstly, a method for detecting the three-dimensional attributes and constructing the distribution map of rice and wheat lodging for combine harvesters is provided, including the following:
[0007] Images of the rice and wheat area ahead are acquired using a binocular vision system mounted on a combine harvester;
[0008] The acquired images are segmented based on a pre-trained neural network model to identify areas where people have fallen.
[0009] The identified collapsed area is used as a segmentation mask to extract the 3D point cloud data of the collapsed area;
[0010] Based on the spatial location of pixels in the two-dimensional image of the lodged area and the corresponding three-dimensional point cloud distribution features, combined with the intrinsic parameters of the binocular vision system and the current position of the combine harvester, the area and spatial location of the lodged area in a single frame image are estimated; based on the point cloud distribution corresponding to the lodged area, the boundary of the lodged area projected in the direction perpendicular to the ground is proposed.
[0011] The collapsed area is divided into several regular grid units on the vertical ground direction projection. For each grid unit, the number of pixels marked as collapsed area inside it is counted, and the ratio between the number and the total number of pixels in the grid is calculated. Grid units with a collapsed pixel ratio higher than the threshold are filtered out.
[0012] Within the effective grid cell, based on the 3D point cloud distribution characteristics, the lodging angle and lodging direction of the lodging area within each grid cell are calculated;
[0013] The first frame of the initial operation is used as the initial keyframe, and the next keyframe is selected based on the distance the combine harvester has moved relative to the previous keyframe.
[0014] A temporary world coordinate system is constructed, and further, regularized ground space units for carrying the lodging attribute information are constructed in the temporary world coordinate system. The spatial points of the lodging area after coordinate transformation in each key frame are projected to the corresponding ground space unit, and the ground space unit to which it belongs is determined according to the spatial position relationship.
[0015] Based on the boundary of the collapsed area projected in the vertical direction of the ground in the current keyframe, calculate the two-dimensional spatial boundary range of the collapsed area in the current keyframe and the overlap between it and the two-dimensional spatial boundary range of each collapsed area in the existing fusion results;
[0016] If the overlap is greater than or equal to the preset overlap threshold, the collapsed area of the current frame and the corresponding existing fusion result area are determined to be the same ground collapsed area, and the collapsed attributes of the collapsed areas in the two are fused and updated; if the overlap is less than the overlap threshold, the collapsed area of the current frame is taken as a new collapsed area, and its collapsed attributes are written into the corresponding ground spatial unit.
[0017] The fusion refers to spatially aligning two overlapping collapsed areas to form a new collapsed area boundary, updating the position and area of the boundary, and weighting the collapsed angle and collapsed direction of the collapsed areas in the grid cells within the new collapsed area boundary.
[0018] In a second aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the method described in the first aspect.
[0019] Thirdly, a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the method described in the first aspect.
[0020] Beneficial effects: Compared with existing technologies, this invention brings significant technological progress and application benefits at the three core levels of integration, continuity, and globalization. The specific beneficial effects are as follows:
[0021] This invention achieves intelligent integration of multiple observation data from the same lodging area through a multi-frame fusion mechanism based on spatial overlap. This method not only correlates data by determining bounding box overlap, but also updates attributes such as lodging angle and direction within the fused area using a weighted average, effectively smoothing out fluctuations caused by noise, segmentation errors, or angle estimation biases in single-frame detection. By cross-validating and supplementing data from multiple perspectives and time points, the confidence and stability of lodging state estimation are improved, reducing false positives and false negatives in complex field environments.
[0022] For scenarios involving continuous operation of combine harvesters, this invention designs a displacement-based keyframe selection and spatiotemporal alignment process. Instead of processing each frame independently, the system selects keyframes with sufficient spatial differences based on the vehicle's actual displacement for in-depth processing and fusion. This avoids redundant calculations of adjacent high-definition frame data, ensures the system's real-time processing efficiency, and guarantees that all lodging data can be coherently analyzed and accumulated within the same spatial reference frame in dynamic operating environments where the vehicle is constantly moving and its posture is constantly changing.
[0023] This invention constructs a global distribution map based on regularized ground spatial units, systematically organizing all fused lodging information (including location, area, angle, and direction) obtained through continuous processing into a unified field map. This allows for a clear visualization of severely lodged areas, main lodging directions, and spatial distribution patterns, thus providing direct and comprehensive data support for combine harvester path planning and adaptive adjustment of header and travel parameters. Attached Figure Description
[0024] Figure 1 Overall flowchart for three-dimensional attribute detection of lodging in rice and wheat.
[0025] Figure 2 Schematic diagram of the construction of an in-vehicle binocular vision system.
[0026] Figure 3 Schematic diagram of the grid division of the lodged area.
[0027] Figure 4 Schematic diagram of the method for detecting the location of landslides.
[0028] Figure 5 Schematic diagram of methods for detecting lodging angle and direction.
[0029] Figure 6 A spatial distribution map of landslides. Detailed Implementation
[0030] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0031] To address the challenges of real-time and accurate acquisition of three-dimensional lodging attributes in rice and wheat fields during combine harvester operations, given the complex spatial distribution, diverse lodging morphologies, and continuous changes in lodging direction and angle, coupled with the continuous movement of vehicles, this invention proposes a method for detecting and constructing three-dimensional lodging attributes and distribution maps for rice and wheat using combine harvesters. Figure 1 As shown, this method first constructs a binocular vision information acquisition system suitable for vehicle-mounted environments based on the structural characteristics of the combine harvester's header and the requirements of field operation scenarios. By reasonably setting the camera's installation position and attitude, stable acquisition of rice and wheat lodging information in the working area in front of the header is achieved. Subsequently, considering the characteristics of uneven field lighting, strong background interference, and large scale differences in lodging areas, the lodging area detection network is improved based on the Deeplabv3+ semantic segmentation framework. By enhancing the multi-scale feature expression capability and strengthening the boundary and tilt texture features of lodging areas, pixel-level accurate identification of rice and wheat lodging areas is achieved, obtaining lodging area segmentation results with complete structure and clear boundaries. On this basis, combined with the depth information and 3D point cloud data acquired by binocular vision, geometric analysis is performed on the lodging areas to extract key 3D attributes such as the spatial location, lodging area, lodging angle, and lodging direction of the lodging areas. Finally, by further combining the pose information of the combine harvester during operation, the lodging detection results of consecutive frames are spatially aligned and fused to construct a lodging spatial distribution map that reflects the overall lodging characteristics of the field. Through the above steps, continuous and stable detection of the three-dimensional attributes of lodged rice and wheat areas was achieved under vehicle-mounted operation conditions, providing reliable perception support for intelligent operation and operation parameter optimization of combine harvesters.
[0032] This embodiment uses a combine harvester platform as a basis and describes rice and wheat fields with lodging as the detection objects. The header of the combine harvester is a key component at its front end used for cutting and transporting crops. In this embodiment, the header is a conventional rigid structure, and its cutting width (i.e., the width of the crop rows that can be harvested in one operation) is one of the basic parameters for determining the coverage of the vision system. The cutting width should be within a conventional range (e.g., 2-5 meters), and the front edge of the header should have a relatively flat mechanical structure to facilitate the installation of the front vision sensor and field of view planning.
[0033] The specific steps are as follows:
[0034] In this invention, considering the concentrated distribution of lodged rice and wheat information and the frequent changes in the working environment in front of the combine harvester's header, a vehicle-mounted binocular vision information acquisition platform is constructed to continuously and stably perceive the lodged areas within the close-range working zone in front of the header. See also... Figure 2 The binocular camera is mounted on the vehicle structure in front of the combine harvester's header, using a forward-tilted mounting method to create a preset pitch angle (preferably approximately 35°) between the camera's optical axis and the ground, thereby expanding the effective field of view of the working area in front. The camera's mounting height is determined based on the combine harvester's header structure dimensions, crop height, and operational safety requirements, with an optimal mounting height of approximately 2.8m. This ensures that the camera can stably cover the main working area of approximately 3m × 4m in front of the header, even under conditions of vehicle movement and slight ground undulations. By coordinating the camera's mounting height and pitch angle, the binocular vision system maintains stable mounting while balancing the cutting width and visibility requirements in front of the work area, acquiring more complete and continuous information on lodged areas.
[0035] In this embodiment, a binocular camera with high-resolution image output and real-time depth measurement capabilities is selected as the vision sensor. Its output resolution and frame rate parameters meet the requirements for real-time performance and detection accuracy during continuous operation of the combine harvester. The effective field of view of the camera is configured according to the working width of the header and the working distance in front, so that it always covers the main distribution area of the lodged area in front of the header during the operation, thereby ensuring that the boundary, shape and local structural features of the lodged area can be clearly obtained.
[0036] By using the above-mentioned installation method and parameter configuration of the vehicle-mounted binocular camera, the vision system can stably acquire images and depth information within the working area in complex field environments, providing a reliable data foundation for subsequent lodging area detection based on semantic segmentation and lodging three-dimensional attribute calculation based on point cloud analysis.
[0037] After the hardware system is installed, the binocular cameras are precisely calibrated. Multiple sets of calibration board images in different poses are acquired, corner feature information is extracted, the intrinsic parameter matrices and distortion coefficients of the left and right cameras are calculated, and the rotation matrix and translation vector between the binocular cameras are further solved. Based on the calibration parameters, distortion correction and epipolar correction are performed on the left and right images to ensure that the same spatial point is located in the same row in both images, guaranteeing the accuracy of disparity calculation from a geometric perspective. Considering the drastic changes in lighting conditions, the presence of shadows and highly reflective areas during field operations, the calibrated images undergo brightness normalization and contrast enhancement before subsequent processing. This enhances the grayscale and texture differences between the lodged area and the background area, improving the robustness of subsequent lodging detection.
[0038] After image preprocessing, the binocular images are input into the lodging area detection model for pixel-level identification of lodged areas in rice and wheat. The lodging area detection model is built upon a semantic segmentation framework, using the Deeplabv3+ semantic segmentation network as its basic structure to perform pixel-level segmentation of lodged areas in complex field environments. Addressing the challenges of large scale variations, complex textures, blurred boundaries, and susceptibility to interference from shadows, weeds, and stubble in lodged areas, this invention specifically enhances the Deeplabv3+ lodging detection network. During model training, the loss function design is improved to ensure the network focuses more on identifying lodged area boundaries and small lodged areas. Combined with multi-scale training samples and data augmentation strategies, the model's adaptability to different operating distances, lodging degrees, and lighting conditions is improved, thereby enhancing its robustness and generalization ability in actual field environments.
[0039] In terms of network structure, a more expressive backbone feature extraction network is introduced to replace the original lightweight backbone structure, thereby improving the model's ability to extract small-scale fallen areas and complex texture features. Simultaneously, a multi-scale attention module is added during feature extraction and fusion, enabling the network to adaptively allocate weights among features of different scales, enhancing the response intensity to real fallen areas and reducing the false detection probability of shadow areas and non-fallen backgrounds. Furthermore, a boundary refinement structure is introduced in the decoding stage to strengthen the constraints on the edges of fallen areas, improving issues such as unclear contours and broken boundaries, and enhancing the spatial integrity and precision of the segmentation results. The Deeplabv3+ semantic segmentation network outputs a probability map of fallen areas, which, through thresholding, generates a binary segmentation mask of the same size as the input image to identify all pixels in the image belonging to fallen areas.
[0040] During the model training phase, the loss function design was improved to make the network pay more attention to the recognition effect of lodged area boundaries and small lodged areas during training. Combined with multi-scale training samples and data augmentation strategies, the model's adaptability to different working distances, lodging degrees, and lighting conditions was enhanced, thereby improving the model's robustness and generalization ability in actual field working environments. The lodged area detection model ultimately outputs a lodged area probability map with the same size as the input image. A binary segmentation mask for the lodged area is generated through threshold determination, used to identify the set of pixels in the image belonging to the lodged area, providing reliable two-dimensional constraints for subsequent calculation of lodged three-dimensional attributes based on binocular point clouds.
[0041] After obtaining the two-dimensional segmentation results of the collapsed area, this invention uses a programmatic method to call the three-dimensional measurement interface provided by the ZED binocular camera to convert the image information into three-dimensional spatial points, generating dense three-dimensional point cloud data corresponding one-to-one with the image. Specifically, the system loads the SVO data acquired by the camera by setting initialization parameters, configures the depth mode to high-precision mode, and unifies the spatial coordinate units and coordinate system type to ensure the consistency of subsequent three-dimensional data in spatial scale and direction. In the processing of each frame of data, the left-eye color image is first acquired as the source of color information, and then the ZED camera interface is called to directly obtain the three-dimensional measurement results corresponding to that frame. The output is three-dimensional point matrix data with the image coordinate system as a reference, and each pixel position corresponds to a three-dimensional coordinate value.
[0042] 3D point cloud data Stored in the form of , where This represents the spatial distance of a pixel along the camera's optical axis. and These represent the horizontal and vertical spatial positions of the point in the camera coordinate system, respectively. By reading the 3D measurement results pixel by pixel from the entire image, a dense point cloud with the same resolution as the original image can be constructed. To ensure the reliability of the point cloud data, this invention performs validity screening on the original 3D measurement results during the point cloud generation stage. When an invalid depth value or a non-finite value exists in the 3D coordinates corresponding to a pixel, the 3D point corresponding to that pixel is discarded, and only point cloud data with valid spatial coordinates is retained.
[0043] Considering the repetitive textures and severe occlusion of rice and wheat plants in the field environment, which may lead to low confidence levels in depth measurements in some areas, this invention further sets a depth confidence threshold during the 3D data acquisition process to filter out measurement results with insufficient depth confidence, thereby reducing the impact of mismatched points on point cloud quality. Through this method, outliers and unreliable points are eliminated in the initial stage of point cloud generation, making the retained point cloud spatially more consistent with the actual crop morphology. After obtaining a valid 3D point cloud, the color information of corresponding pixels in the left-eye color image is mapped to the 3D point cloud, so that each 3D point carries a corresponding color attribute, thus constructing a colored 3D point cloud model. This colored point cloud not only preserves the spatial geometric information of the lodged areas of rice and wheat but also integrates color and texture features, providing a richer data foundation for subsequent point cloud selection based on semantic segmentation results and 3D geometric analysis of lodged areas.
[0044] After obtaining the 2D segmentation results of the lodged area and generating the 3D point cloud, this embodiment further utilizes the lodged area segmentation mask to spatially crop the 3D point cloud, thereby achieving accurate extraction of the lodged area point cloud. Specifically, since each point in the 3D point cloud corresponds one-to-one with a pixel position in the image, the 2D segmentation mask can be directly mapped onto the 3D point cloud data. 3D points whose corresponding pixel positions are marked as lodged areas are retained, while 3D points corresponding to non-lodged areas are discarded. In this way, a subset of the point cloud containing only lodged areas can be directly filtered from the entire scene point cloud, strictly limiting the subsequent 3D geometric analysis object to the interior of the lodged rice and wheat area.
[0045] After filtering the point cloud of the collapsed area, to achieve local fine-grained analysis of the three-dimensional attributes of the collapsed area and improve the spatial consistency of the results, this invention further divides the collapsed area into several regular grid cells on the ground projection plane, and uses the grid cell as the smallest analysis unit for calculating the collapsed attributes. See [link to relevant documentation]. Figure 3 Specifically, the 3D point cloud of the collapsed area is first projected onto a ground reference plane. The coverage area of the collapsed area is determined based on the plane coordinates in the camera coordinate system or a unified world coordinate system. Then, according to a preset grid size, the area is regularly divided along the horizontal and forward directions to form uniformly sized and orderly arranged rectangular grid units. The grid size can be set according to the required accuracy and computational efficiency of the collapsed area detection, ensuring that local collapsed features are reflected while avoiding excessive computation due to an excessively small grid or smoothing of local features due to an excessively large grid.
[0046] After completing the regular mesh division, the 2D collapsed region segmentation results are mapped to mesh cells. For each mesh cell, the number of pixels marked as collapsed regions by the segmented mask within it is counted, and the ratio between this number and the total number of pixels in the mesh is calculated as the collapsed proportion index of that mesh cell. When the collapsed proportion of a mesh cell is lower than a preset threshold, it indicates that the mesh contains only a small number of scattered collapsed pixels, which may originate from segmentation edge errors, noise interference, or local false detection areas. Such meshes are judged as invalid meshes and are not included in the subsequent calculation of collapsed attributes. When the collapsed proportion meets or exceeds the threshold, the collapsed information in the mesh is considered to have good spatial continuity and representativeness, and it is marked as a valid mesh for subsequent calculation of 3D attributes such as collapsed angle, collapsed direction, and collapsed area.
[0047] Furthermore, considering that combine harvesters harvest continuously along the forward direction using rows as basic working units during actual operation, this invention, after calculating the lodging attributes of effective grid units, fuses the results of multiple effective grid units located within the same working row. Specifically, the lodging angle and lodging direction results of each effective grid unit within the same row are weighted and averaged according to their lodging area or lodging pixel ratio to obtain the overall lodging angle and lodging direction characterization results corresponding to that working row.
[0048] Next, we will detect the lodging property. The specific steps are as follows:
[0049] This invention first performs a unified quantitative calculation of the area and spatial location of the collapsed areas detected in the entire frame of the image. Unlike methods that are based on grid cell-by-cell statistics, this invention directly uses the complete collapsed area in the current frame as the analysis object, thereby avoiding the area dispersion error caused by grid division and improving the overall consistency of collapsed area and location estimation.
[0050] For the current frame image, the number of pixels identified as lying areas in the current frame image is calculated by performing pixel statistics on the fallen area mask. Since two-dimensional pixel area cannot directly reflect the actual physical scale, this invention further combines depth information acquired by binocular vision to perform physical scale transformation on the pixel area of the fallen area. For each pixel within the fallen area... The corresponding depth value is denoted as Given the camera's intrinsic parameters, the physical area of a single pixel in real space is proportional to the square of its depth. The physical area of the corresponding pixel is calculated below. .
[0051]
[0052] in, and These are the focal length parameters of the camera in the horizontal and vertical directions, respectively.
[0053] Because binocular cameras are mounted with a forward tilt, there is a fixed angle between their imaging plane and the field surface. This causes the spatial area of the same pixel on the imaging plane to differ from its actual projected area on the ground. Directly performing physical scale conversion based on pixel spatial area will introduce a systematic area estimation error along the camera's line of sight. This error will be further amplified, especially when the camera tilt angle is large or the working distance changes.
[0054] To eliminate area-scale distortion caused by tilted camera shooting, this invention introduces a ground-plane-based projection correction mechanism during the calculation of the collapsed area. Specifically, the spatial attitude relationship of the camera coordinate system relative to the ground reference plane is obtained through camera calibration and pose estimation, and a ground plane model is established, whose normal vector is denoted as n. g For any pixel within the collapsed area, the viewing direction vector of its corresponding spatial point in the camera coordinate system is calculated using binocular vision. Based on spatial geometry, the angle α between the viewing direction and the ground plane is calculated as follows:
[0055]
[0056]
[0057] Based on this, the physical area of a pixel on the imaging plane can be further corrected according to its orthographic projection onto the ground plane. According to projection geometry principles, the actual area of a pixel after ground plane projection correction can be expressed as:
[0058]
[0059] This correction process is equivalent to mapping the spatial area observed by a pixel from a tilted perspective to a ground reference plane, thereby obtaining an area metric consistent with the actual ground coverage.
[0060] Therefore, by summing up the actual areas of all pixels within the collapsed area after projection correction, the actual physical area S of the collapsed area on the ground plane in the current frame is obtained by the following formula:
[0061]
[0062] By introducing the aforementioned projection correction mechanism, the impact of the forward tilt of the camera on the area calculation is effectively compensated, making the lodging area estimation results more consistent with the real scale of the field ground, and improving the accuracy and stability of the three-dimensional attribute calculation of lodging under vehicle-mounted operation conditions.
[0063] Parameter description: Ω represents the set of pixels in the collapsed area, (u,v) represents the pixel coordinates of the collapsed area, and Z(u,v) represents the pixel depth value. and Here, represents the camera's intrinsic parameters, specifically the focal length parameters in the horizontal and vertical directions; d(u,v) is the viewing direction vector of the spatial point corresponding to pixel (u,v) in the camera coordinate system; c x ,c y Here, represents the principal point coordinates of the camera; represents the camera intrinsic parameters; represents the pixel coordinates of the intersection point of the camera's optical axis and the image plane; n g This is the unit normal vector for the ground, describing the orientation of the ground reference plane.
[0064] Based on the calculation of the collapsed area, this embodiment further provides a quantitative description of the location of the collapsed area in the actual working space. The collapsed location is used to characterize the center position of the collapsed area in the working space in the current frame, and is an important foundation for subsequent collapse distribution map construction and multi-frame fusion.
[0065] First, the centroid position of the fallen region is calculated in a two-dimensional image coordinate system. Based on the fallen region mask, the two-dimensional centroid coordinates of the fallen region in the image plane are calculated; this centroid reflects the geometric center position of the fallen region in the current frame image. The method for calculating the spatial position of the fallen region in a single frame image is as follows:
[0066]
[0067] Subsequently, combining the depth information acquired through binocular vision, the two-dimensional centroid is mapped onto three-dimensional space. Let the depth value at the centroid location be... Based on the camera imaging model, it is converted into three-dimensional spatial coordinates in the camera coordinate system. , used to describe the spatial location of the collapsed area in the camera coordinate system.
[0068]
[0069] Parameter description: Where X c ,Y c Z c The coordinates of the three-dimensional center point of the fallen area in a single frame image represent the three-dimensional spatial center position of the fallen area in the camera coordinate system; N is the total number of pixels in the fallen area. Location of the centroid of the collapsed area ( , The depth value corresponding to ().
[0070] Based on obtaining the spatial coordinates of the center point of the collapsed area, this invention extracts the boundary of the point cloud corresponding to the collapsed area to further describe the actual coverage of the collapsed area in the work space. Specifically, the boundary point cloud is extracted from the point cloud within the collapsed area, and the boundary point cloud is projected onto the ground plane or... In the plane, we analyze its distribution in the planar coordinate system.
[0071] By analyzing the boundary point cloud Extreme value statistics were performed on the coordinate components in the plane to obtain the minimum and maximum coordinate values of the collapsed area in the plane. , , and These four corner points are used to determine the axis-aligned bounding box aligned with the ground coordinate system, thereby characterizing the spatial extent of the collapsed area within the work space.
[0072] Finally, this invention simultaneously outputs the spatial coordinates of the center point of the collapsed area and the spatial boundary range formed by the four corner points, as a representation of the position and range of the collapsed area in the current frame. See [link to relevant documentation]. Figure 4 This method not only accurately locates the center of the lodged area in the field operation space, but also effectively describes the spatial expansion scale of the lodged area, providing more complete and stable spatial information for subsequent cross-frame correlation of the lodged area, distribution map construction, and operation decisions.
[0073] See Figure 5 After calculating the spatial location and area of the lodged region, this invention further performs local geometric modeling of the point cloud of the lodged region within the effective grid cell to extract key attributes reflecting the lodging posture of rice and wheat, such as the lodging angle and lodging direction. For the point cloud of the lodged region within a certain effective grid cell, the least squares method is used to perform plane fitting on the point cloud to describe the overall tilting trend of the lodged plants in the local area. The general expression of the fitted plane is shown in the formula, where the normal vector of the fitted plane is used to describe the spatial orientation of the local lodged region.
[0074]
[0075] in, Let n be the normal vector of the fitted plane. Simultaneously, the normal vector of the ground plane is obtained by fitting the ground plane, denoted as n. g .
[0076] In the process of calculating the lodging angle, this invention uses the ground plane as a unified reference benchmark. The angle between the fitting plane normal vector and the ground plane normal vector is calculated. Used to characterize the degree to which rice and wheat plants lean relative to the ground.
[0077]
[0078] in, The lodging angle is used to reflect the degree to which rice and wheat plants lean relative to the ground. The larger the angle value, the more severe the lodging; when the angle is close to zero, it means that the plants are basically upright.
[0079] In the process of calculating the direction of collapse, in order to eliminate the influence of the collapse angle on the direction determination, this embodiment projects the normal vector of the fitting plane onto the ground plane to obtain its projection vector n on the ground plane. kThe direction of this projection vector is used as the lodging direction of the corresponding lodging area. In this way, the three-dimensional tilt attitude of the lodging area can be decomposed into two independent attributes with clear physical meaning: "degree of tilt" and "direction of tilt," which facilitates subsequent analysis and visualization of the lodging direction distribution.
[0080]
[0081] By repeating the above calculation process of the collapsed area, collapsed area, collapsed angle, and collapsed direction for each effective grid cell, this invention achieves local fine-grained extraction of the three-dimensional attributes of the collapsed area and provides structured and quantifiable basic data for subsequent multi-frame fusion and the construction of collapsed distribution maps.
[0082] During continuous operation of combine harvesters, changes in camera mounting posture, ground undulations, and vehicle attitude during movement can lead to spatial biases when directly fusing multi-frame lodging detection results based on the camera coordinate system, making it difficult to ensure the consistency of lodging distribution results at a global scale. Therefore, this invention uses the field ground plane as a spatial reference to construct a temporary world coordinate system, and on this basis, maps multi-frame lodging information to a unified global coordinate system.
[0083] For any given image frame, after acquiring the point cloud, the field ground is first modeled using the point cloud data of that frame. Considering the presence of non-ground point cloud interference such as crops and weeds in the field environment, this invention employs a random sampling consensus algorithm to perform planar fitting on the point cloud, extracting the principal plane model representing the ground from it.
[0084] See Figure 2 After obtaining the ground plane normal vector, this normal vector is used as the spatial alignment reference to perform attitude correction on the camera coordinate system. Specifically, by constructing a rotation matrix, the ground plane normal vector is rotated to be consistent with the vertical direction in the world coordinate system, thereby eliminating the influence of camera pitch and roll angles on spatial representation and achieving vertical correction of the camera coordinate system. Simultaneously, by maintaining the consistency of the ground plane in the horizontal direction, horizontal attitude alignment is achieved. After attitude correction, the coordinate system is translated along the normal direction so that the fitted ground plane falls entirely within the temporary world coordinate system. On the plane, through the above rotation and translation operations, a temporary world coordinate system based on the ground is constructed, such that in this coordinate system... The axis is perpendicular to the ground. – The plane corresponds to the actual field ground.
[0085] To construct a continuous and stable spatial distribution map of rice and wheat lodging, this invention further defines a global reference coordinate system based on a temporary world coordinate system. Specifically, the temporary world coordinate system obtained at the start of the operation after the aforementioned ground plane correction is used as the global coordinate reference system throughout the entire operation. .
[0086] During continuous operation of combine harvesters, due to the high frequency of image acquisition, there is often significant spatial overlap between adjacent frames. If the lodging detection results of each frame are directly superimposed onto the distribution map, the same lodging area is easily counted repeatedly, which not only increases the computational burden but also affects the accuracy and stability of the lodging distribution results. To address this, this invention introduces a fusion strategy based on operation displacement during multi-frame fusion, filtering key frames that participate in the distribution map update.
[0087] Specifically, by combining the pose information acquired by the inertial measurement unit (IMU) built into the binocular camera, the spatial displacement between adjacent frames during continuous operation of the combine harvester is calculated in real time. Only when the displacement distance of the current frame relative to the previous frame participating in the fusion exceeds a preset spatial threshold is the lodging detection result of that frame included in the global distribution map update process; when the displacement does not reach the threshold, it is only used for the calculation of the lodging attribute of the current frame and does not participate in global accumulation. This method ensures that the distribution map update frequency matches the actual operating cycle of the combine harvester, avoiding spatial redundancy caused by high-frequency images.
[0088] After keyframe selection, this invention spatially fuses the landslide information in a global reference coordinate system. Specifically, with the ground plane as a reference, regularized ground spatial units are constructed on the X–Y plane of the global coordinate system to carry landslide attribute information. The spatial points of the landslide area in each keyframe, after coordinate transformation, are projected onto the corresponding ground spatial unit, and their unit is determined according to their spatial position.
[0089] When the same ground spatial unit is observed multiple times in different keyframes, this invention does not simply superimpose the collapsed area and collapsed attributes, but instead uses a fusion update method. If the overlap is greater than or equal to a preset overlap threshold, the collapsed area in the current frame and the corresponding existing fusion result area are determined to be the same ground collapsed area, and the collapsed attributes of the collapsed areas in both are fused and updated to obtain a more stable and reliable collapsed attribute estimation result; if the overlap is less than the overlap threshold, it can be identified as a collapsed area that has not been observed before, and the collapsed area in the current frame is taken as a new collapsed area, and its collapsed attributes are written into the corresponding ground spatial unit.
[0090] The fusion refers to spatially aligning two overlapping collapsed areas to form a new collapsed area boundary, updating the position and area of the boundary, and weighting the collapsed angle and collapsed direction of the collapsed areas in the grid cells within the new collapsed area boundary.
[0091] By employing the aforementioned displacement-triggered keyframe filtering mechanism and lodging information fusion method based on ground spatial units, this invention achieves effective deduplication and fusion of multiple lodging detection results acquired during continuous operations, while ensuring computational efficiency and real-time performance. This gradually constructs a spatial distribution map of rice and wheat lodging that reflects the overall lodging distribution characteristics of the field. (See also...) Figure 6 .
[0092] This invention proposes a method for three-dimensional attribute detection and distribution map construction of rice and wheat lodging for combine harvesters. It systematically innovates three key aspects: accurate detection of lodging areas, efficient extraction of lodging attributes, and global representation of lodging information. Compared with existing technologies, it has the following advantages:
[0093] To address the problems of traditional lodging attribute analysis relying on complete 3D reconstruction, which is computationally complex and lacks real-time performance, this invention directly constructs a high-quality point cloud within the lodging area by combining binocular depth information. It also employs a gridded local region partitioning method to spatially decompose the lodging area, thereby improving the stability and noise resistance of local point cloud analysis. Within each local grid cell, a lightweight point cloud plane fitting method is used to model the lodging point cloud, and a normal vector consistency constraint is introduced to ensure that the fitting results accurately reflect the overall lodging trend of rice and wheat plants within that area. Based on this, the lodging angle is accurately calculated by fitting the geometric relationship between the plane normal vector and the ground normal vector, and the lodging direction is obtained by the projection direction of the normal vector onto the ground plane. This method can complete lodging attribute analysis without complex 3D reconstruction, significantly reducing the computational load while maintaining computational accuracy, meeting the dual requirements of real-time performance and accuracy under continuous operation conditions of combine harvesters.
[0094] A method for constructing a global lodging attribute distribution map based on combine harvester pose is developed, achieving a unified representation of lodging information at the field scale. Addressing the problem that single-frame lodging detection results cannot accurately reflect the lodging distribution characteristics of the entire field, this invention constructs a stable world coordinate system through ground plane fitting. Utilizing a binocular vision system and combine harvester pose estimation results, the lodging areas detected in consecutive frames, along with their corresponding lodging positions, angles, and directions, are uniformly aligned to the same global coordinate frame. During cross-frame fusion, spatial statistics and consistency updates of lodging attributes are performed, effectively suppressing errors caused by single-frame noise and repeated detections. Ultimately, a global lodging spatial distribution map reflecting the lodging distribution characteristics of rice and wheat across the entire field is constructed. This method transforms lodging information from local perception to global representation, providing an intuitive and reliable spatial basis for adjusting harvester operating parameters.
[0095] In summary, this invention effectively solves the problems of unstable lodging detection, difficulty in real-time and accurate acquisition of lodging attributes, and lack of global representation of lodging information in existing technologies by improving the semantic segmentation network of lodging areas, proposing a three-dimensional attribute analysis method based on local plane fitting, and constructing a global distribution map of lodging attributes by combining the pose of combine harvesters. It has significant technological progress and practical application value.
[0096] Corresponding to the above method, this application embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above method.
[0097] The functions of each functional unit of the electronic device provided in the above embodiments of this application can be implemented through the above methods and steps. Therefore, the specific working process and beneficial effects of each unit in the electronic device provided in the embodiments of this application will not be repeated here.
[0098] Corresponding to the above method, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above method.
[0099] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the methods described in the above embodiments.
[0100] Those skilled in the art will understand that the embodiments in this application can be provided as methods, systems, or computer program products. Therefore, the embodiments in this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the embodiments in this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0102] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0103] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0104] Although preferred embodiments have been described in this application, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of this application.
[0105] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims in this application and their equivalents, then this application also intends to include these modifications and variations.
Claims
1. A method for detecting three-dimensional attributes and constructing distribution maps of lodging in rice and wheat for combine harvesters, characterized in that, Including the following: Images of the rice and wheat area ahead are acquired using a binocular vision system mounted on a combine harvester; The acquired images are segmented based on a pre-trained neural network model to identify areas where people have fallen. The identified collapsed area is used as a segmentation mask to extract the 3D point cloud data of the collapsed area; Based on the spatial location of pixels in the two-dimensional image of the lodged area and the corresponding three-dimensional point cloud distribution features, combined with the intrinsic parameters of the binocular vision system and the current position of the combine harvester, the area and spatial location of the lodged area in a single frame image are estimated; based on the point cloud distribution corresponding to the lodged area, the boundary of the lodged area projected in the direction perpendicular to the ground is proposed. The collapsed area is divided into several regular grid units on the vertical ground direction projection. For each grid unit, the number of pixels marked as collapsed area inside it is counted, and the ratio between the number and the total number of pixels in the grid is calculated. Grid units with a collapsed pixel ratio higher than the threshold are filtered out. Within the effective grid cell, based on the 3D point cloud distribution characteristics, the lodging angle and lodging direction of the lodging area within each grid cell are calculated; The first frame of the initial operation is used as the initial keyframe, and the next keyframe is selected based on the distance the combine harvester has moved relative to the previous keyframe. A temporary world coordinate system is constructed, and further, regularized ground space units for carrying the lodging attribute information are constructed in the temporary world coordinate system. The spatial points of the lodging area after coordinate transformation in each key frame are projected to the corresponding ground space unit, and the ground space unit to which it belongs is determined according to the spatial position relationship. Based on the boundary of the collapsed area projected in the vertical direction of the ground in the current keyframe, calculate the two-dimensional spatial boundary range of the collapsed area in the current keyframe and the overlap between it and the two-dimensional spatial boundary range of each collapsed area in the existing fusion results; If the overlap is greater than or equal to the preset overlap threshold, the collapsed area of the current frame and the corresponding existing fusion result area are determined to be the same ground collapsed area, and the collapsed attributes of the collapsed areas in the two are fused and updated; if the overlap is less than the overlap threshold, the collapsed area of the current frame is taken as a new collapsed area, and its collapsed attributes are written into the corresponding ground spatial unit. The fusion refers to spatially aligning two overlapping collapsed areas to form a new collapsed area boundary, updating the position and area of the boundary, and weighting the collapsed angle and collapsed direction of the collapsed areas in the grid cells within the new collapsed area boundary.
2. The method for detecting and constructing three-dimensional attributes of rice and wheat lodging for combine harvesters according to claim 1, characterized in that: The formula for calculating the area S of the collapsed region in a single frame image is as follows: Where Ω represents the set of pixels in the collapsed region, (u,v) represents the pixel coordinates of the collapsed region, and Z(u,v) represents the pixel depth value. and Here, represents the camera's intrinsic parameters, specifically the focal length parameters in the horizontal and vertical directions; d(u,v) is the viewing direction vector of the spatial point corresponding to pixel (u,v) in the camera coordinate system; c x ,c y Here, represents the principal point coordinates of the camera; represents the camera intrinsic parameters; represents the pixel coordinates of the intersection point of the camera's optical axis and the image plane; n g This is the unit normal vector for the ground, describing the orientation of the ground reference plane.
3. The method for detecting and constructing three-dimensional attributes of rice and wheat lodging for combine harvesters according to claim 2, characterized in that: The method for calculating the spatial location of the fallen area in the single-frame image is as follows: Where X c Y c Z c represents the three-dimensional center point coordinates, indicating the three-dimensional spatial center position of the collapsed area in the camera coordinate system; N is the total number of pixels in the collapsed area. Location of the centroid of the collapsed area ( , The depth value corresponding to ().
4. The method for detecting and constructing three-dimensional attributes of rice and wheat lodging for combine harvesters according to claim 3, characterized in that: The method for calculating the lodging angle includes: For the three-dimensional point cloud of the collapsed area within the effective grid cell, the least squares method is used to perform plane fitting to obtain the fitted plane equation Ax+By+Cz+D=0, and the normal vector n=(A,B,C) of the fitted plane is extracted. The unit normal vector n of the ground plane is obtained by fitting the ground plane. g ; Calculate the fitting plane normal vector n and the ground normal vector n g The included angle θ between the two is taken as the bend angle of the grid cell, and the calculation formula is as follows: Where θ represents the tilting angle of the rice and wheat plants relative to the ground.
5. The method for detecting and constructing three-dimensional attributes of rice and wheat lodging for combine harvesters according to claim 4, characterized in that: The method for calculating the direction of the collapse includes: Projecting the normal vector n of the fitted plane onto the ground plane yields its projection vector n on the ground plane. k The calculation formula is: With the projection vector n k The direction is used as the lodging direction of the corresponding lodging area to characterize the main lodging direction of rice and wheat plants on the ground plane.
6. The method for detecting and constructing three-dimensional attributes of rice and wheat lodging for combine harvesters according to claim 5, characterized in that: The pre-trained neural network model is the Deeplabv3+ semantic segmentation network. In the process of feature extraction and fusion, the Deeplabv3+ semantic segmentation network introduces a multi-scale attention module to enable the network to adaptively allocate weights to features of different scales, enhance the response to the fallen area and suppress background interference. A boundary refinement structure is introduced in the decoding stage to strengthen the constraint on the edge of the fallen area; the Deeplabv3+ semantic segmentation network outputs a probability map of the fallen area, and generates a binary segmentation mask with the same size as the input image by threshold determination to identify all pixels in the image that belong to the fallen area.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.