Ploughing monitoring method and system based on unmanned aerial vehicle

Through drone video data splicing and SEEM model segmentation technology, real-time monitoring and refined management of cultivated land are achieved, and the problems of low efficiency, high cost and poor real-time performance in the existing technology are solved, and the efficiency and accuracy of monitoring are improved.

CN120182866APending Publication Date: 2025-06-20CHENGDU JOUAV AUTOMATION TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510249309.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing arable land monitoring methods are inefficient, costly and poor real-time, and are unable to achieve efficient and accurate monitoring of arable land.

Method used

By obtaining adjacent key image frames of drone video data for stitching, combining with SEEM model for panoramic segmentation, extracting mask profiles and converting them into geographical coordinate systems, real-time monitoring and refined management of cultivated land are achieved.

Benefits of technology

Real-time monitoring of cultivated land is achieved, the efficiency and accuracy of monitoring is improved, the costs are reduced, and problems existing in cultivated land protection can be discovered and dealt with in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182866A_ABST
    Figure CN120182866A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a cultivated land monitoring method and system based on an unmanned aerial vehicle. The method comprises the following steps: obtaining adjacent key image frames of unmanned aerial vehicle video data, and splicing the adjacent key image frames to obtain a panoramic image; performing panoramic segmentation on the panoramic image and obtaining masks reflecting different types of cultivated land through an SEEM model; and extracting the contour of the mask, converting the extracted contour into coordinates in a geographic coordinate system, and mapping each coordinate point with the land category to which the coordinate point belongs. According to the application, farmland monitoring can be carried out by using the aerial image of the unmanned aerial vehicle, the farmland is monitored in real time, meanwhile, image processing is carried out by using the panoramic image splicing technology and the SEEM model, a wider view can be obtained, the limitation of a single image is eliminated, and refined management of the farmland is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and particularly relates to a method and system for cultivated land monitoring based on an unmanned aerial vehicle (UAV). Background Art

[0002] With the acceleration of the urbanization process and the development of social economy, land resources are becoming increasingly tense, and cultivated land protection is facing unprecedented challenges. Traditional means of cultivated land monitoring, such as manual inspections and satellite remote sensing, have problems such as low efficiency, high cost, and poor real-time performance. In recent years, UAV technology has been widely used in the field of agricultural monitoring due to its flexibility, convenience, low cost, and strong real-time performance. UAV technology can obtain information on farmland through aerial images, realizing efficient and precise monitoring of cultivated land.

[0003] Existing cultivated land protection methods mainly rely on manual inspections and satellite remote sensing. Manual inspection is carried out by personnel conducting on-site inspections on the ground. Although this method is intuitive, it has low efficiency, high cost, and poor real-time performance. Satellite remote sensing obtains remote sensing images of farmland through satellites. Although it can cover a large area, affected by factors such as weather and clouds, the image quality is unstable, and the cost is relatively high.

[0004] Existing cultivated land protection methods have many problems in practical applications. First, manual inspections are inefficient and costly, and real-time monitoring cannot be achieved. Second, the image quality of satellite remote sensing is unstable, affected by factors such as weather and clouds, and the cost is relatively high. In addition, existing methods cannot achieve refined management of cultivated land and are difficult to meet the needs of modern agricultural development. Therefore, how to achieve efficient and precise monitoring of cultivated land is an important issue faced in the current field of cultivated land protection. Summary of the Invention

[0005] The technical problems to be solved by the present application are how to combine UAV video data with a map to achieve comparative analysis of changes in the flight area and observation and analysis of the situation environment, and how to realize the visualization and secondary development of UAV video data to improve the utilization rate of video data.

[0006] One of the technical solutions adopted by the present application to solve its technical problems is: a method for cultivated land monitoring based on an unmanned aerial vehicle, the method comprising:

[0007] Obtaining adjacent key image frames of UAV video data and stitching them to obtain a panoramic image;

[0008] Performing panoramic segmentation on the panoramic image and obtaining a mask reflecting different categories of cultivated land through the SEEM model;

[0009] Extracting the contour of the mask, converting the extracted contour into coordinates in the geographic coordinate system, and mapping each coordinate point to the land category to which it belongs.

[0010] Further, the video data includes POS data corresponding to each frame of image.

[0011] Further, obtaining adjacent key image frames of the UAV video data and stitching them to obtain a panoramic image specifically includes:

[0012] Extracting key image frames;

[0013] Extracting the POS data of the key image frames and performing preliminary alignment on adjacent key image frames;

[0014] Performing feature extraction on the preliminarily aligned adjacent key image frames;

[0015] Calculating a homography matrix and screening to obtain an optimal homography matrix;

[0016] Using the optimal homography matrix to optimize the preliminarily aligned adjacent key image frames and stitching the optimized adjacent key image frames to obtain a panoramic image.

[0017] Further, the extracting the POS data of the key image frames and performing preliminary alignment on adjacent key image frames specifically includes:

[0018] Rectifying adjacent key image frames;

[0019] Obtaining relative displacement data of the rectified adjacent key image frames in mutually perpendicular directions;

[0020] Using the relative displacement data to perform preliminary translational alignment on the rectified adjacent key image frames to obtain preliminarily aligned adjacent key image frames.

[0021] Further, the performing feature extraction on the preliminarily aligned adjacent key image frames specifically includes:

[0022] Selecting the central pixel of the preliminarily aligned adjacent key image frames;

[0023] Obtaining the pixel points around the central pixel;

[0024] If the absolute value of the difference between the pixel value of the central pixel and the pixel values of the surrounding pixel points is greater than a preset threshold, the central pixel is marked as a feature point.

[0025] Further, the calculating a homography matrix and screening to obtain an optimal homography matrix specifically includes:

[0026] Obtaining 4 random pairs of matching points between the rectified key image frames and the preliminarily aligned adjacent key image frames and calculating a preliminary homography matrix;

[0027] Calculate the projection error between the corrected key image frame and the initially aligned key image frame adjacent to it through the homography matrix;

[0028] Set the projection error threshold thresh, and determine that the matching points within the threshold thresh are inliers;

[0029] Recalculate the homography matrix using the inliers;

[0030] Repeat the above steps and screen the recalculated homography matrix to obtain the optimal homography matrix.

[0031] Furthermore, the panoramic image is panoramically segmented and a mask reflecting cultivated land of different categories is obtained through the SEEM model, specifically including:

[0032] Perform panoramic segmentation on the panoramic image, and calculate the original size of each initial slice image after panoramic segmentation of the panoramic image;

[0033] Modify the initial slice image to the input size of the SEEM model through bilinear interpolation to obtain the modified slice image;

[0034] Input the modified slice image into the SEEM model, and output a mask containing the corresponding cultivated land category;

[0035] Restore the mask to the size of the initial slice image through the nearest neighbor interpolation algorithm to obtain the final mask.

[0036] Furthermore, the step of modifying the slice image to the input size of the SEEM model through bilinear interpolation to obtain the modified slice image specifically includes

[0037] Obtain the pixel position (x block , y block ) of each pixel point in the initial slice image, and the pixel position corresponding to the size of the modified slice image is (x block_resize , y block_resize ): Among them, scale x and scale y are the scaling factors of the pixel point in the horizontal and vertical directions;

[0038] Obtain the pixel value of the pixel point corresponding to the coordinates in the size of the modified slice image at the adjacent integer coordinates of the pixel point;

[0039] Use the pixel values of the adjacent integer coordinates to calculate the interpolated pixel value of the pixel point in the size of the modified slice image;

[0040] The modified slice image is obtained according to the pixel values after interpolation of each pixel point in the initial slice image.

[0041] Further, the extraction of the contour of the mask specifically includes:

[0042] Smoothing the contour of the final mask by Gaussian filtering;

[0043] Then, performing dilation operation on the final mask to obtain a dilated mask and performing erosion operation to obtain an eroded mask;

[0044] Then, subtracting the dilated mask from the eroded mask to obtain the contour.

[0045] Further, the conversion of the extracted contour into coordinates in the geographic coordinate system specifically includes:

[0046] Converting the pixel coordinates of the contour into geographic coordinates through the following formula:

[0047] longitude = longitude origin +(pixel x ×longitude pixel_size ) ;

[0048] latitude = latitude origin +(pixel y ×latitude pixel_size ) ;

[0049] where longitude is the longitude, latitude is the latitude, (pixel x , pixel y ) are the pixel coordinates of the contour; (longitude origin , latitude origin ) are the origin coordinates, longitude pixel_size is the scale ratio of the longitude to the pixel coordinate pixel x , and latitude pixel_size is the scale ratio of the latitude to the pixel coordinate pixel y .

[0050] The beneficial effects of the farmland monitoring method based on drones in this application are as follows: It can utilize drone aerial images for farmland monitoring, achieve real-time monitoring of farmland, and promptly discover and handle problems existing in farmland protection. Meanwhile, by adopting panoramic image stitching technology and the SEEM model for image processing, a broader field of view can be obtained, eliminating the limitations of single images and realizing refined management of farmland. The real-time data transmission method can ensure that after the ground base station receives the video data, it performs real-time decoding and processing to achieve real-time monitoring of farmland.

[0051] Another technical solution adopted by this application to solve its technical problems is: A farmland monitoring system based on drones, including: a drone video transmission module and a data processing module;

[0052] The drone video transmission module acquires video data through the drone and is used for encoding and real-time transmission of the video data;

[0053] The data processing module acquires the real-time transmitted video data and performs real-time decoding, and is used to execute the following operations:

[0054] Stitch adjacent key image frames of the decoded video data to obtain a panoramic image;

[0055] Perform panoramic segmentation on the panoramic image and obtain a mask reflecting different categories of farmland through the SEEM model;

[0056] Extract the contour of the mask, convert the extracted contour into coordinates in the geographic coordinate system, and map each coordinate point to the land category it belongs to.

[0057] The beneficial effects of the farmland monitoring system based on drones in this application are as follows: Utilize drone aerial images for farmland monitoring. The flight altitude of the drone is moderate, which can not only cover a large area of farmland but also ensure a relatively high image resolution. During the video acquisition process, the camera carried by the drone captures the farmland video at a high frame rate, and at the same time records the POS information corresponding to each frame of the image to ensure the precise geolocation of the image. This real-time data acquisition method can achieve real-time monitoring of farmland, promptly discover and handle problems existing in farmland protection. The real-time data transmission method can ensure that after the ground base station receives the video data, it performs real-time decoding and processing to achieve real-time monitoring of farmland. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The following further describes this application in conjunction with the drawings and embodiments.

[0059] Figure 1 It is a flowchart of the farmland monitoring method based on drones in this application;

[0060] Figure 2Schematic diagram of the cultivated land monitoring system based on UAV in this application; Detailed implementation manners

[0061] In order to make the technical means, creative features, achieved purposes and functions implemented in this application easy to understand, the following further elaborates this application in combination with specific implementation manners.

[0062] The technical solutions in the embodiments of this application are clearly and completely described below. Obviously, the described embodiments are only one of the embodiments of this application, rather than all embodiments. Based on the embodiments in this application, other embodiments popularized by those of ordinary skill in the art all fall within the scope of protection of this application.

[0063] As Figure 1 shown, a cultivated land monitoring method based on UAV provided by this application includes the following steps:

[0064] Obtain adjacent key image frames of the UAV and splice them to obtain a panoramic image and related data;

[0065] Perform panoramic segmentation on the panoramic image and obtain masks reflecting different categories of cultivated land through the SEEM model;

[0066] Extract the contours of the masks, convert the extracted contours into coordinates in the geographic coordinate system, and map each coordinate point to the land category to which it belongs.

[0067] Among them, in some specific embodiments, after the UAV obtains video data, the video data can be encoded and transmitted in real time through the UAV video transmission module; the data processing module obtains the real-time transmitted video data and performs real-time decoding to obtain adjacent key image frames. Further, the H.264 video coding standard can be used to compress and encode the video data to reduce the data volume while maintaining the image quality. The encoded video data can be transmitted in real time through WebRTC and saved locally in the form of recording. The WebRTC protocol is a real-time communication protocol that supports peer-to-peer audio and video transmission without an intermediate server.

[0068] After obtaining the real-time transmitted video data, perform real-time decoding on it and extract the key image frames therein. Since the video captured by the UAV may be jittery or blocked, in order to obtain a complete cultivated land image, it is necessary to splice the continuous key image frames in the video. Using panoramic image stitching technology, seamless stitching between key image frames is achieved by matching the feature points between adjacent key image frames. The stitched panoramic image and related data not only contain a wider field of view, but also eliminate the limitations of a single image.

[0069] The stitched panoramic image and related data are fed into the SEEM (Segment Everything Everywhere All at Once) model for panoramic segmentation. The SEEM model can handle various input prompts (such as points, masks, text, etc.) and form prompts in the same visual-semantic space, thus achieving powerful combinability. The masks output by the model reflect different categories of cultivated land, such as farmland, buildings, roads, etc.

[0070] Extract the contours of the masks from the segmentation results, simplify them, and use Gaussian filtering to make the contour edges smoother. Then use the polygon similarity calculation method to reduce the number of nodes to reduce the computational complexity during subsequent geographical coordinate conversion. Convert the extracted contours into coordinate points in the geographical coordinate system, and then map each geographical coordinate point to the land category identified by the deep learning model.

[0071] Combine Figure 1 and Figure 2 As shown in

[0072] In some specific embodiments, the video data includes POS data corresponding to each frame of the image. Among them, the POS data can include GPS position information, longitude and latitude, altitude, pitch angle, roll angle, heading angle, yaw angle, etc. The POS data can be embedded in the video data, and the video data is compressed into compressed data by means of the H264 compression technology, and the compressed data is transmitted back to the ground end through the WebRTC method. After receiving the compressed data, the ground end decodes and restores the compressed data to video data by means of the H264 decoding technology.

[0073] In some specific embodiments, the process of obtaining adjacent key image frames of the drone video data and stitching them into a panoramic image specifically includes:

[0074] Extract key image frames. Among them, the video data can be frame-extracted at a rate of one frame every 5 frames to obtain key image frames for panoramic image stitching;

[0075] Extract the POS data of the key image frames and perform preliminary alignment on adjacent key image frames;

[0076] The ORB (Oriented Fast and Rotated Brief) feature extraction algorithm can be used to extract features from adjacent key image frames that are initially aligned;

[0077] Calculate the homography matrix, and use the RANSAC (Random Sample Consensus) algorithm to filter and obtain the optimal homography matrix;

[0078] Use the optimal homography matrix to optimize the adjacent key image frames that are initially aligned, and splice the optimized adjacent key image frames to obtain a panoramic image.

[0079] In some specific embodiments, extracting the POS data of the key image frames and initially aligning the adjacent key image frames specifically includes:

[0080] Rectify the adjacent key image frames. Among them, in some further embodiments, the following formula can be used to rectify the adjacent key image frames I1 and I2: R = R roll #R pitch ·R yaw , where I1 and I2 are the original key image frames, and are the rectified key image frames, R roll is the roll angle transformation, R pitch is the pitch angle transformation, R yaw is the yaw angle transformation. In addition, θ yaw is the yaw angle, θ pitch is the pitch angle, θ roll is the roll angle;

[0081] Obtain the relative displacement data of the rectified adjacent key image frames in mutually perpendicular directions. Among them, in some further embodiments, the relative displacement data ΔX and ΔY between the rectified adjacent key image frames in two mutually perpendicular directions can be obtained, where the key image frames I1 and I 2+1 , are respectively at geographical positions (X1, Y1) and (X2, Y2). The relative displacements ΔX and ΔY between the images can be obtained by the formulas ΔX = X2 - X1 and ΔY = Y2 - Y1;

[0082] Use the relative displacement data to perform preliminary translational alignment on the rectified adjacent key image frames to obtain the initially aligned adjacent key image frames. Among them, in some further embodiments, the following formula can be used to perform preliminary translational alignment on the adjacent key image frames: Among them, is the aligned key image frame, x W and xW The key image frames after alignment The geographical coordinates of each pixel point, where x and y are The geographical coordinates of each pixel point of W and y W The geographical coordinates of each pixel point of the key image frames after alignment.

[0083] In some specific embodiments, the feature extraction of the preliminarily aligned adjacent key image frames specifically includes:

[0084] Select the central pixel p of the preliminarily aligned adjacent key image frames;

[0085] Obtain the pixel points around the central pixel p;

[0086] If the absolute value of the difference between the pixel value I(p) of the central pixel and the pixel values I(p i ) of the surrounding pixel points is greater than a preset threshold t, that is, |I(p i ) - I(p)| > t, then the central pixel p is marked as a feature point.

[0087] Furthermore, in order to handle rotational invariance, the direction of each feature point needs to be calculated. For the central pixel p marked as a feature point, the formula for its direction is where (x i , y i ) are the coordinates of the surrounding pixel point p i relative to p, and I(p i ) is its pixel value. The descriptor is generated using BRIEF (Binary Robust Independent Features), and a binary descriptor is generated by comparing the intensities of a pair of pixel points in the neighborhood of the feature point, where p a and p b are a pair of pixels randomly selected in the neighborhood;

[0088] The similarity of the feature descriptors in two images is calculated using the Hamming distance. For the descriptors d1 and d2 extracted from images and respectively, the Hamming distance is: where represents the binary exclusive OR operation, and n is the length of the binary descriptor;

[0089] In some specific embodiments, the calculation of the homography matrix and the screening to obtain the optimal homography matrix specifically includes:

[0090] Obtain the corrected key image frames and the preliminarily aligned key image frames adjacent to it Randomly select 4 pairs of matching points and calculate the initial homography matrix H rand ;

[0091] Among them, the homography matrix can be obtained specifically through the following method:

[0092] For the matching points of two images and There is an expression for the homography matrix as:

[0093]

[0094] The matching points satisfy: Estimate H using the least squares method through multiple matching points. The calculation process is as follows. There is a pair of matching image points (x A , y A ) and (x B , y B ), and the following constraint relationships can be established;

[0095]

[0096] Among them, λ is the scale factor because the image coordinates are homogeneous coordinates. By expanding the matrix multiplication, the following three equalities can be obtained;

[0097] Equation 1: λx B = h 11 x A + h 12 y A + h 13 ;

[0098] Equation 2: λy B = h 21 x A + h 22 y A + h 23 ;

[0099] Equation 3: λ = h 31 x A + h 32 y A + h 33 ;

[0100] To eliminate the scale factor, substitute Equation 3 into the above Equation 1 and Equation 2 to obtain the following two equalities;

[0101] Equation 4: x B (h 31 x A + h 32 y A + h 33 ) = h 11 xA +h 12 y A +h 13 ;

[0102] Equation 5: y B (h 31 x A +h 32 y A +h 33 ) = h 21 x A +h 22 y A +h 23 ;

[0103] Expanding the above Equation 4 and Equation 5, the following two equations can be obtained;

[0104] x A h 11 +y A h 12 +h 13 -x B x A h 31 -x B y A h 32 -x B h 33 = 0;

[0105] x A h 21 +y A h 22 +h 23 -y B x A h 31 -y B y A h 32 -y B h 33 = 0;

[0106] Since the homography matrix H has 8 degrees of freedom, at least 4 pairs of matching points are required to solve it, and a system of linear equations can be constructed in the form: A·h = 0;

[0107] where A is the coefficient matrix and h is the column vector containing the elements of the homography matrix:

[0108] h = (h 11 h 12 h 13 h 21 h 22 h 23 h 31 h 32h 33 ) T ;

[0109] By minimizing the two - norm of A·h = 0, the least - squares method or SVD (Singular Value Decomposition) is used to solve it. Specifically, A = U∑V T , where U is an orthogonal matrix containing the left singular vectors of A; ∑ represents a diagonal matrix whose diagonal elements are the singular values of matrix A; V T is an orthogonal matrix containing the right singular vectors of A, and the singular value corresponding to the last column of V T is the smallest. Therefore, h is the last column of V. Finally, the obtained h is a homogeneous solution, and the resulting homography matrix usually needs to be normalized to obtain the homography matrix;

[0110] Among them, for the preliminary homography matrix H rand , the RANSAC algorithm can be used to improve the robustness;

[0111] The specific process is to randomly select 4 pairs of matching points and calculate the preliminary homography matrix H through the above steps of the homography matrix rand ;

[0112] Calculate the projection error between the corrected key image frame rand and the adjacent preliminarily aligned key image frame using the homography matrix H . In some specific embodiments, all pixel points on one key image frame can be projected to another key image frame through the homography matrix H rand and the projection error is calculated;

[0113] Set the projection error threshold thresh, and determine that the matching points within the threshold thresh are inliers;

[0114] Recalculate the homography matrix using the inliers;

[0115] Repeat the above steps and screen the recalculated homography matrix to obtain the optimal homography matrix H final .

[0116] Furthermore, through the formula obtain the deformed key image frame

[0117] Set the length and width of the size of the canvas C for image stitching to be W canvas and H canvas , and stack with onto the canvas;

[0118] On With Overlaying on the canvas, the overlapping area of two key image frames appears, and weighted averaging can be used for image fusion. Let α be the fusion coefficient, and the fusion formula is where (x canvas , y canvas ) represents the pixel coordinates of the overlapping area. The fusion weight is dynamically adjusted according to the characteristics of the overlapping area to achieve a more natural transition and obtain the final canvas;

[0119] Repeat the above steps to obtain the final panoramic image, and the calculation process is where I stitched is the final panoramic image, n is the number of key frames extracted, and C n represents the canvas for each key frame stitched.

[0120] In some specific embodiments, performing panoramic segmentation on the panoramic image and obtaining a mask reflecting different categories of cultivated land through the SEEM model specifically includes:

[0121] Performing panoramic segmentation on the panoramic image based on the pixels of the image acquisition, and calculating the original size of each initial slice image I block after panoramic segmentation of the panoramic image;

[0122] Modifying the initial slice image I block to the input size of the SEEM model by bilinear interpolation to obtain the modified slice image;

[0123] Inputting the modified slice image into the SEEM model and outputting a mask mask containing the corresponding cultivated land category;

[0124] Restoring the mask mask to the size of the initial slice image I block through the nearest neighbor interpolation algorithm to obtain the final mask mask′.

[0125] In some specific embodiments, since the panoramic image is too large, a preprocessing process is required to put the panoramic image into the neural network model SEEM. The preprocessing process first calculates the width w block and height h block of each initial slice image I block based on the pixels of the image acquisition. At the same time, an overlapping process (overlap) is performed to ensure that each pixel is covered by a more accurate step size for each movement. The calculation formula is as follows:

[0126]

[0127] Calculate the horizontal step size x step and the vertical step size y step, the calculation formula is:

[0128] x step = w block ×(1 - overlap);

[0129] y step = h block ×(1 - overlap);

[0130] The initial slice image I is modified to the input size of the SEEM model through bilinear interpolation, specifically including: block Modify to obtain the modified slice image, specifically including:

[0131] Obtain the pixel positions (x block , y block ) of each pixel point in the initial slice image I through the following formula. The corresponding pixel positions in the size of the modified slice image are (x block , y block_resize ) as follows: block_resize ) Among them, scale x and scale y are the scaling factors of the pixel point in the horizontal and vertical directions;

[0132] Obtain the pixel values corresponding to the coordinates of the pixel points with adjacent integer coordinates of the pixel point in the size of the modified slice image. Among them, the four adjacent pixel value integer coordinates are the pixel values of the corresponding coordinates after compressing the image, Q 11 = I(x int , y int ), Q 12 = I(x int , y int + 1), Q 21 = I(x int + 1, y int ), Q 22 = I(x int + 1, y int + 1);

[0133] Use the pixel values of the adjacent integer coordinates to calculate the interpolated pixel value of the final pixel position (x block_resize , y block_resize ) of the pixel point in the size of the modified slice image. Q(x block_resize , y block_resize ) = (1 - x float )(1 - y float )Q 11 + x float (1 - y float )Q 21+(1 - x float )y float Q 12 +x float y float Q 22 ;

[0134] Use the interpolated pixel values of four neighboring pixel values to calculate the final pixel position (x block_resize , y block_resize ). Q(x block_resize , y block_resize ) = (1 - x float )(1 - y float )Q 11 +x float (1 - y float )Q 21 +(1 - x float )y float Q 12 +x float y float Q 22

[0135] Obtain the modified slice image I block based on the interpolated pixel values of each pixel point in the initial slice image I block_resize .

[0136] Send I block_resize into the SEEM network. Its formula is Re(mask, class) = SEEM(I block_resize ), where the output result contains the mask mask and the corresponding cultivated land class class, and SEEM(·) represents the SEEM model;

[0137] Enlarge the mask mask to the same size as I block through the nearest neighbor interpolation algorithm;

[0138] Calculation process where represents rounding down, x mask and y mask are the coordinates of the mask mask, x' mask and y' mask are the coordinates of the enlarged final mask mask', and S x and S y represent the enlargement factors. W' and H' are the length and width of the final mask mask', and W and H represent the length and width of the mask mask;

[0139] Repeat the above calculation process for all pixels of the mask mask to obtain the final mask mask'.

[0140] In some specific embodiments, extracting the contour of the mask specifically includes:

[0141] Smoothing the contour of the final mask mask′ through Gaussian filtering;

[0142] Then, performing a morphological dilation operation on the final mask mask′ respectively to obtain a dilated mask mask′ dilatation and an erosion operation to obtain an eroded mask mask′ corrode ;

[0143] Then, subtracting the eroded mask mask′ dilatation obtained by the erosion operation from the dilated mask mask′ corrode to obtain a contour mask′ silhouettes .

[0144] In some specific embodiments, converting the extracted contour into coordinates in a geographic coordinate system specifically includes:

[0145] Converting the pixel coordinates of the contour mask′ silhouettes into geographic coordinates through the following formula:

[0146] longitude = longitude origin +(pixel x × longitude pixel_size );

[0147] latitude = latitude origin +(pixel y × latitude pixel_size );

[0148] where longitude is the longitude, latitude is the latitude, (pixel x , pixel y ) are the pixel coordinates of the contour; (longitude origin , latitude origin ) are the origin coordinates, longitude pixel_size is the scale ratio of the longitude to the pixel coordinate pixel x , and latitude pixel_size is the scale ratio of the latitude to the pixel coordinate pixel y .

[0149] The converted (longitude, latitude) and its corresponding cultivated land category of the pixel coordinates, and the above-mentioned cultivated land category class and the historical cultivated land category classhistory Make a comparison to determine whether they are consistent. If they are not, an alarm prompt will be given. Among them, the historical cultivated land category class history is pre-stored.

[0150] In some further embodiments, the masked contours after being mapped with the land category to which they belong can be classified to obtain the cultivated land classification result at the current moment. Further, the cultivated land classification result at the current moment can be compared and analyzed with historical data to evaluate the changes in the use of cultivated land, including changes in cultivated land area, changes in land use types, etc. Through the above steps, the method realizes the efficient monitoring and protection of cultivated land, provides timely and accurate data support for relevant departments, helps to formulate reasonable cultivated land protection policies, and promotes sustainable development.

[0151] This application also provides an unmanned aerial vehicle (UAV)-based cultivated land monitoring system, including: a UAV video transmission module and a data processing module;

[0152] The UAV video transmission module acquires video data through the UAV and is used for real-time transmission after encoding the video data;

[0153] The data processing module acquires the real-time transmitted video data and performs real-time decoding, and is used to perform the following operations:

[0154] Stitch adjacent key image frames of the decoded video data to obtain a panoramic image;

[0155] Perform panoramic segmentation on the panoramic image and obtain a mask reflecting different categories of cultivated land through the SEEM model;

[0156] Extract the contour of the mask, convert the extracted contour into coordinates in the geographic coordinate system, and map each coordinate point to the land category to which it belongs.

[0157] In addition, the UAV-based cultivated land monitoring system in this application can execute the UAV-based cultivated land monitoring method as described above. In the UAV-based cultivated land monitoring system and the UAV-based cultivated land monitoring method provided in this application, UAV aerial images are used for cultivated land monitoring. The UAV flight height is appropriate, which can not only cover a large area of farmland but also ensure a high image resolution. During the video acquisition process, the camera carried by the UAV shoots the cultivated land video at a high frame rate, and at the same time records the POS information corresponding to each frame of the image to ensure the accurate geolocation of the image. This real-time data acquisition method can realize the real-time monitoring of cultivated land and timely discover and handle problems existing in cultivated land protection.

[0158] Compared with traditional farmland monitoring methods, such as manual inspections and satellite remote sensing, there are problems of low efficiency and high cost. This method uses drone technology for farmland monitoring. The purchase and maintenance costs of drones are relatively low, and a large amount of manpower input is not required, greatly reducing the cost of farmland monitoring.

[0159] Moreover, this application uses panoramic image stitching technology and the SEEM model for image processing, which can obtain a broader field of view, eliminate the limitations of single images, and simultaneously achieve refined management of farmland. This high-precision image processing method can improve the accuracy of farmland monitoring and provide more reliable data support for relevant departments.

[0160] This application can further compare and analyze the current farmland classification results with historical data to evaluate the changes in the use of farmland, including changes in farmland area, land use types, etc. This comparison and analysis function can better understand the dynamic changes of farmland and provide a basis for formulating reasonable farmland protection policies.

[0161] Finally, this application uses the RTMP protocol for real-time data transmission. The RTMP protocol has the characteristic of low latency and is suitable for real-time data transmission. This real-time data transmission method can ensure that after the ground base station receives the video data, it decodes and processes it in real time to achieve real-time monitoring of the farmland.

[0162] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media provided in this application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0163] It should be noted that, in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements not only includes those elements but also other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.

[0164] The foregoing has shown and described the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited by the above embodiments, and the above-described embodiments and descriptions in the specification are only for explaining the principles of the present application. Without departing from the spirit and scope of the present application, the present application will have various changes and improvements, and all such changes and improvements fall within the scope of the claims of the present application. The scope of protection claimed by the present application is defined by the appended claims and their equivalents.

Claims

1. The method for monitoring cultivated land based on drones is characterized by: The method comprises: Acquire adjacent key image frames of drone video data and stitch them together to obtain a panoramic image; The panoramic image is segmented and the masks reflecting different types of cultivated land are obtained through the SEEM model; The contour of the mask is extracted, the extracted contour is converted into coordinates in a geographic coordinate system, and each of the coordinate points is mapped to the land category to which it belongs.

2. The method for monitoring cultivated land based on an unmanned aerial vehicle according to claim 1, characterized in that: The video data includes POS data corresponding to each image frame.

3. The method for monitoring cultivated land based on an unmanned aerial vehicle according to claim 1, characterized in that: The step of obtaining adjacent key image frames of the drone video data and stitching them together to obtain a panoramic image specifically includes: Extract key image frames; Extract POS data of key image frames and perform preliminary alignment of adjacent key image frames; Perform feature extraction on the adjacent key image frames that are initially aligned; Calculate the homography matrix and select the optimal homography matrix; The optimal homography matrix is ​​used to optimize the initially aligned adjacent key image frames, and the optimized adjacent key image frames are spliced ​​to obtain a panoramic image.

4. The method for monitoring cultivated land based on an unmanned aerial vehicle according to claim 3, characterized in that: The extracting of POS data of key image frames and preliminary alignment of adjacent key image frames specifically includes: Correcting adjacent key image frames; Obtaining relative displacement data of adjacent key image frames after correction in mutually perpendicular directions; The relative displacement data is used to perform preliminary translation alignment on the corrected adjacent key image frames to obtain preliminary aligned adjacent key image frames.

5. The method for monitoring cultivated land based on an unmanned aerial vehicle according to claim 4, characterized in that: The feature extraction of the initially aligned adjacent key image frames specifically includes: Selecting central pixels of adjacent key image frames for preliminary alignment; Get the pixel points around the central pixel; If the absolute value of the difference between the pixel value of the central pixel and the pixel values ​​of the surrounding pixels is greater than a preset threshold t, the central pixel is marked as a feature point.

6. The method for monitoring cultivated land based on an unmanned aerial vehicle according to claim 5, characterized in that: The calculating of the homography matrix and screening to obtain the optimal homography matrix specifically includes: Obtain four random matching points between the rectified key image frame and its adjacent preliminary aligned key image frame, and calculate the preliminary homography matrix; Calculate the projection error between the rectified key image frame and the adjacent preliminary aligned key image frame through the homography matrix; Set a projection error threshold thresh, and determine that matching points within the projection error range of the threshold thresh are inliers; Recompute the homography matrix using the interior points; Repeat the above steps and screen the recalculated homography matrix to obtain the optimal homography matrix.

7. The method for monitoring cultivated land based on drone according to claim 1, characterized in that: The panoramic image is segmented and masks reflecting different types of cultivated land are obtained through the SEEM model, specifically including: Performing panoramic segmentation on the panoramic image, and calculating the original size of each initial slice image after the panoramic segmentation of the panoramic image; The initial slice image is modified to the input size of the SEEM model by bilinear interpolation to obtain a modified slice image; The modified slice image is input into the SEEM model and a mask containing the corresponding cultivated land category is output; The mask is restored to the size of the initial slice image through the nearest neighbor interpolation algorithm to obtain the final mask.

8. The method for monitoring cultivated land based on an unmanned aerial vehicle according to claim 7, characterized in that: The step of modifying the initial slice image to the input size of the SEEM model by the bilinear interpolation method to obtain the modified slice image specifically includes: Get the pixel position (x) of each pixel in the initial slice image block ,y block ), the corresponding pixel position of the modified slice image size is (x block_resize ,y block_resize ): Among them, scale x With scale y is the scaling factor of the pixel in the horizontal and vertical directions; Obtaining pixel values ​​of coordinates corresponding to pixel points of integer coordinates adjacent to the pixel point in the modified slice image size; Calculate the interpolated pixel value of the pixel point in the modified slice image size using the pixel values ​​of the adjacent integer coordinates; The modified slice image is obtained according to the interpolated pixel value of each pixel point in the initial slice image.

9. The method for monitoring cultivated land based on an unmanned aerial vehicle according to claim 7, characterized in that: The extracting the contour of the mask specifically includes: The contour of the final mask is smoothed by Gaussian filtering; Then the final mask is subjected to dilation operation to obtain a dilation mask and erosion operation to obtain an erosion mask; Then subtract the dilation mask from the erosion mask to get the contour.

10. The method for monitoring cultivated land based on an unmanned aerial vehicle according to claim 8, characterized in that: The step of converting the extracted contour into coordinates in a geographic coordinate system specifically includes: The pixel coordinates of the contour are converted to geographic coordinates using the following formula: longitude=longitude origin +(pixel x ×longitude pixel_size ); latitude=latitude origin +(pixel y ×latitude pixel_size ); Where longitude is longitude, latitude is latitude, (pixel x ,pixel y ) are the pixel coordinates of the contour; (longitudeorigin, latitudeorigin) are the origin coordinates, longitude pixel_size Longitude and pixel coordinates x The scale ratio, latitude pixel_size is the latitude and pixel coordinate pixel y The scale ratio size.

11. The farmland monitoring system based on drone is characterized by: include: UAV image transmission module and data processing module; The drone image transmission module obtains video data through the drone, and is used to encode the video data and transmit it in real time; The data processing module obtains the real-time transmitted video data and performs real-time decoding, and is used to perform the following operations: splicing adjacent key image frames of the decoded video data to obtain a panoramic image; The panoramic image is segmented and the masks reflecting different types of cultivated land are obtained through the SEEM model; The contour of the mask is extracted, the extracted contour is converted into coordinates in a geographic coordinate system, and each of the coordinate points is mapped to the land category to which it belongs.

Citation Information

Cited By

  • Crop disaster monitoring, early warning and evaluation platform for main grain production area

    CN120374299A

  • Intelligent monitoring and real-time early warning system based on unmanned aerial vehicle

    CN121214267A