Intelligent harvesting method, agricultural harvesting robot chip system and medium

By applying an intelligent harvesting method based on artificial intelligence in agricultural harvesting robots, using neural network models to identify crop areas and perform precise harvesting, the problems of low efficiency and low automation of traditional harvesting methods are solved, and efficient and accurate intelligent harvesting is achieved.

CN119942328APending Publication Date: 2025-05-06NANTONG INST OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510013709.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The traditional manual and mechanized harvesting methods cannot adapt to the changing farmland environment, resulting in low harvesting efficiency and manual control, and the inability to achieve fully automated intelligent harvesting.

Method used

Using an intelligent harvesting method based on artificial intelligence models, we take crop images, use pre-trained neural network models to perform image processing, identify crop areas, and perform precise harvesting.

Benefits of technology

It improves the efficiency and accuracy of crop harvesting, reduces the harvesting of weeds and other debris, and realizes fully or semi-automatic intelligent harvesting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942328A_ABST
    Figure CN119942328A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent harvesting method, an agricultural harvesting robot chip system and a medium, is applied to the field of intelligent agriculture, and can improve the crop harvesting efficiency and accuracy. The method specifically comprises the steps that a first image is obtained through shooting, the first image is processed based on a first neural network model to obtain a second image, and the second image is used for enhancing crop-related features in the second image and weakening non-crop-related features in the second image so as to improve the subsequent crop row recognition effect. Further, based on the reflectivity of each pixel point in the second image in the preset wave band, a first operation is executed on the second image to obtain a third image, and the first operation comprises removing the pixel points with the reflectivity exceeding a preset threshold value in the preset wave band in the second image. Processing the third image based on a second neural network model to obtain a first recognition result, wherein the first recognition result is used for indicating an area where the crops are located in the first image; and finally harvesting the crops based on the first identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of smart agriculture, and in particular relates to an intelligent harvesting method, an agricultural harvesting robot chip system and a medium. Background Art

[0002] With the development of agricultural modernization, traditional manual harvesting methods can no longer meet the growing demand for agricultural products. Although mechanized harvesting solutions have been proposed, mechanized harvesting tools still cannot adapt to the changing farmland environment, so they still need to be manually controlled or driven for harvesting, which consumes manpower and has low harvesting efficiency. In view of this, how to achieve fully automated intelligent harvesting of crops is a problem that needs to be solved at present. Summary of the invention

[0003] The embodiments of the present application provide an intelligent harvesting method, an agricultural harvesting robot chip system and a medium, which can improve the efficiency and accuracy of crop harvesting.

[0004] In a first aspect, a smart harvesting method is provided, the method comprising:

[0005] Capturing and obtaining a first image, wherein image content of the first image includes crops;

[0006] The first image is processed based on the first neural network model to obtain the second image, wherein the decoder parameters of the first neural network model are trained for the purpose of reducing the first loss function value, and the first loss function value is used to characterize the feature difference between the input image and the output image; the encoder parameters of the first neural network model are trained for the purpose of reducing the second loss function value, the second loss function value is a composite calculation result of the first value and the second value, and the weight values ​​corresponding to the first value and the second value are determined based on the model training target of the first neural network model, the first value is used to characterize the feature similarity between the input image and the output image, and the second value is used to characterize the difference between the crop recognition result and the target recognition result for the input image;

[0007] Performing a first operation on the second image based on the reflectivity of each pixel in the second image in a preset band to obtain a third image, wherein the first operation includes removing pixels in the second image whose reflectivity exceeds a preset threshold in the preset band, and the reflectivity of the color corresponding to the crop in the preset band is less than the preset threshold;

[0008] Processing the third image based on the second neural network model to obtain a first recognition result, where the first recognition result is used to indicate an area where the crops in the first image are located;

[0009] The crops are harvested based on the first recognition result.

[0010] Optionally, the third image is processed based on the second neural network model to obtain a first recognition result:

[0011] Based on the second neural network model, each pixel in the third image is mapped to an n-dimensional vector in the feature space using a clustering loss function, and then the first recognition result is further determined according to the mapping result.

[0012] Optionally, before capturing the first image, the method further includes:

[0013] capturing a fourth image, wherein the fourth image also includes the crops;

[0014] The mechanical orientation of the agricultural harvesting robot is adjusted according to the fourth image.

[0015] Optionally, adjusting the mechanical direction of the agricultural harvesting robot according to the fourth image comprises:

[0016] Processing a fifth image using a third neural network model to obtain a second recognition result, the fifth image being the fourth image or an image obtained by processing the fourth image, the second recognition result comprising baselines of n crop rows, the third neural network model being a lightweight neural network model, and n being an integer greater than or equal to 1;

[0017] Use the baseline of the n crop rows as a guide to adjust the machine direction.

[0018] Optionally, before using the third neural network model to process the fifth image to obtain the second recognition result, the method further includes:

[0019] The fourth image is cropped to obtain a fifth image, where the fifth image is an image of a preset size and includes at least one independent crop row, where the crop row consists of parts of crops.

[0020] Optionally, the value of n is determined based on the type of crop and / or the size of the field in which the crop is located.

[0021] In a second aspect, a terminal device is provided, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the terminal device implements the steps of the method described in any one of the above-mentioned first aspects.

[0022] In a third aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.

[0023] In a fourth aspect, a computer program product is provided. When the computer program product is executed on a terminal device, the terminal device executes any one of the methods described in the first aspect.

[0024] In a fifth aspect, a chip system is provided, which includes a processor coupled to a memory, and the processor executes a computer program stored in the memory to implement the method described in any one of the first aspects above.

[0025] The chip system may be a single chip or a chip module composed of multiple chips.

[0026] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 An exemplary flow chart of the smart harvesting method for crops in this application is provided;

[0028] Figures 2 to 4 Here are some examples of crop row distribution;

[0029] Figure 5 is an example of the relative position of the camera and the crop row;

[0030] Figure 6 The baseline diagram corresponding to the crop row;

[0031] Figure 7 It is a schematic diagram of the relationship between the edge line and the baseline;

[0032] Figure 8 It is a schematic diagram of the relationship between the machine direction, crop row direction and horizontal line direction;

[0033] Fig. 9 A schematic diagram of a model training process. DETAILED DESCRIPTION

[0034] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0035] With the development of agricultural modernization, traditional manual harvesting methods can no longer meet the growing demand for agricultural products. Although mechanized harvesting solutions have been proposed, mechanized harvesting tools still cannot adapt to the changing farmland environment, so they still need to be manually controlled or driven for harvesting, which consumes manpower and has low harvesting efficiency.

[0036] In order to improve the harvesting efficiency of agricultural crops, this application relates to an intelligent harvesting solution for crops. However, due to the complex environment in the farmland, in addition to crops, there are usually various weeds. If a large number of weeds are involved while harvesting crops, it will cause great trouble for the later crop processing. Therefore, how to harvest crops accurately is one of the issues that need to be considered when implementing intelligent harvesting solutions for crops.

[0037] In view of this, the present application proposes an intelligent harvesting method for crops, in which crop identification is assisted by an artificial intelligence model, and crops are harvested accurately. First, before performing the harvesting operation, an image containing crops to be harvested can be taken, and then the captured image can be processed by a pre-trained first neural network model to accurately identify the area where the crops are located, and finally the crops are accurately harvested based on the recognition results. Through this solution, the efficiency of crop harvesting can be improved, the harvesting accuracy can be improved, and the harvesting of other debris (such as weeds) other than crops can be reduced.

[0038] Combine the following Figure 1 The method 100 in the embodiment of the present application exemplifies the intelligent harvesting method for crops provided in the embodiment of the present application. The method 100 can be executed by an agricultural harvesting robot, or by an internal module or component of the agricultural harvesting robot, which is not limited in the present application, and the agricultural harvesting robot execution method 100 is used as an example for explanation below. The agricultural harvesting robot described in the embodiment of the present application is a fully automatic or semi-automatic robot capable of intelligently harvesting crops. The agricultural harvesting robot can be a special-purpose robot, that is, it is specifically used for harvesting crops, or the agricultural harvesting robot can also be a multifunctional robot, that is, the agricultural harvesting robot can also have other functions in addition to the harvesting function, which is not limited in the present application.

[0039] The present application does not limit the types of crops that the agricultural harvesting robot is used to harvest. The crops can be various grains, such as wheat, rice, etc., or various vegetables, such as leeks, etc. The present application does not limit this.

[0040] S110: Capture and obtain a first image, where image content of the first image includes crops.

[0041] Exemplarily, the agricultural harvesting robot in the present application includes a visual recognition system, which is composed of one or more cameras, for example. These cameras can be RGB cameras, ToF cameras, hyperspectral cameras, or a combination of different types of cameras, which is not limited in the present application.

[0042] The agricultural harvesting robot can control one or more cameras in the visual recognition system through a control system to shoot the farmland to obtain a first image, and the image content of the first image includes crops, or in other words, the first image is an image obtained by shooting the crops.

[0043] It should be understood that the crops in the first image are usually only part of the crops in the farmland.

[0044] It should also be understood that the first image in the present application may be one or more images.

[0045] It should also be understood that the first image may be an image directly captured by one or more cameras, or an image obtained after pre-processing (such as noise reduction, cropping, rotation, smoothing, etc.) the image captured by one or more cameras, or an image obtained after fusion processing of multiple images captured by one or more cameras. This application does not limit this.

[0046] Optionally, before executing step S110, the agricultural harvesting robot may also pre-adjust the mechanical direction. The so-called adjustment of the mechanical direction refers to adjusting the overall direction of the agricultural harvesting robot, or refers to adjusting the direction of the harvesting device of the agricultural harvesting robot. The purpose of pre-adjusting the mechanical direction is to make the mechanical direction parallel to the crop rows as much as possible. For example, the purpose of adjusting the mechanical direction of the agricultural harvesting robot is to make the angular deviation between the mechanical direction and the crop rows to be harvested within a preset range. The crop rows to be harvested here refer to a continuous area composed of crops to be harvested. That is to say, in a possible example adapted by the present application, the crops are distributed in the form of crop rows. Figure 2 One possible example is given.

[0047] Figure 2 This is an example of crop row distribution. In this example, four crop rows are shown as examples (as shown in the shaded part in the figure), and these four crop rows are parallel. The area outside the crop rows is usually a road or weeds in the farmland. It should be understood that Figure 2 The crop row distribution and the number of crop rows shown are only examples, and the crop rows in the figure are an ideal state. In actual scenarios, the crop rows are usually not completely parallel, and the sizes of different crop rows may vary. More importantly, the edges of the crop rows are usually bumpy, not the completely straight edges shown in the figure, because when sowing, the seeds are generally randomly and discretely distributed, and the growth of crops is irregular. There will be many weeds in the farmland, and the weeds will also invade the edges of the crop rows.

[0048] In this application, the machine direction is adjusted for at least two purposes: first, a new image needs to be retaken for crop row recognition. Specifically, if the machine direction deviation is large, part of the crop row area to be harvested may not be detected. For example, the retaken image is Figure 3 As shown in the figure, part of the crop row is included in the image (the left-hand shaded part in the figure). Due to the large deviation in the mechanical direction, the crop row is cut obliquely, which may cause part of the crop row area to be outside the image, such as Figure 4 As shown in the shaded part on the middle right. Moreover, the fact that this part of the crop row is outside the image is not due to the limitation of the visual range of the visual recognition system. Therefore, although it is possible to alleviate this problem by adjusting the angle of the visual recognition system and taking multiple shots, it will also make the overall process very complicated (for example, it will involve the problem of multiple image synthesis and the order problem when performing crop row recognition), and it will also cause a waste of resources. In summary, if crop row detection is performed according to such an image, some crop rows will not be detected, resulting in the inability to harvest the crops in these crop rows, or the system complexity will be increased, resulting in a waste of resources. Secondly, pre-adjusting the mechanical direction can facilitate the subsequent agricultural harvesting robot to move in a direction close to parallel to the crop and perform the harvesting operation, which can reduce the number of subsequent adjustments to the mechanical direction and improve the efficiency of crop harvesting.

[0049] The following is an exemplary description of a possible implementation method for adjusting the mechanical direction of an agricultural harvesting robot.

[0050] The agricultural harvesting robot collects a fourth image through a camera, and the fourth image may be an RGB image, a ToF image, or another type of image. The agricultural harvesting robot then adjusts the mechanical direction of the agricultural harvesting robot based on the fourth image.

[0051] In a possible implementation, the agricultural harvesting robot determines the baseline of n crop rows based on the fourth image, and adjusts the mechanical direction according to the baseline. In the present application, the baseline of the crop row is a straight line used to indicate the direction of the crop row, which may refer to a line fitted based on the crop row area, or an edge line of the agricultural crop row, or a median line of the agricultural crop row (i.e., a connecting line between the midpoint of the top area of ​​the crop row and the midpoint of the bottom area), etc., which is not limited in the present application.

[0052] First, let's explain why the machine direction should be adjusted based on the baseline of the crop row: it should be understood that Figure 2The crop row image shown is an ideal state and is the crop row shape observed from a bird's-eye view. However, in reality, the visual recognition system of the agricultural harvesting robot shoots the crop rows at an oblique angle in the outer area of ​​the farmland. Therefore, the crop rows captured by the camera in the visual recognition system are generally in an oblique state. Figure 5 An example is given: because the image is taken at a certain tilt angle in front of the crop row, the crop row on the left side of the camera will tend to tilt from right to left in the image, while the crop row on the right side of the camera will tend to tilt from left to right in the image. In addition, the farther away from the camera, the greater the tilt angle of the crop row. In other words, due to the rule of "near is big and far is small", the captured image is quite different from the bird's-eye view. Even if the crop rows are originally parallel, they will not be arranged in parallel in the image. Of course, Figure 5 This is just an example. The actual shooting effect will usually not be bilaterally symmetrical.

[0053] Agricultural harvesting robots can determine a baseline for each crop row and then adjust the direction of the machine based on the baseline. An ideal baseline state is as follows: Figure 6 As shown, Figure 6 The examples in Figure 5 That is, the solution described in this application hopes to use a straight line to represent the direction of each crop row. This straight line can be a line fitted based on the crop row, or an edge line of the agricultural crop row, or a median line of the agricultural crop row, etc. Among them, when the baseline is a line fitted based on the crop row, a possible implementation method is to determine n horizontal center points to Figure 7 For example, n horizontal lines can be drawn between the left and right edge lines of a crop row ( Figure 7Taking the n horizontal lines as an example), these n horizontal lines can be arranged equidistantly, and the first horizontal line and the last horizontal line can be the upper edge line and the lower edge line of the crop row respectively. Since in actual scenes, the edge lines of crop rows are usually not straight lines, the midpoints of these n horizontal lines are generally not on a straight line. Therefore, these n center points can be linearly fitted to obtain a straight line as the baseline of the crop row. In this way, the baseline can better characterize the direction of the crop row, so as to more accurately adjust the mechanical direction. It can be understood that n is a preconfigured value and a positive integer greater than or equal to 2. In addition, in a further optional implementation, after determining the n center points, more center points can be obtained by interpolation, and then linear fitting is performed based on all the center points, so that the accuracy of the baseline obtained is higher. In the case where the baseline is the edge line of the crop row, the baseline can be the left edge line of the crop row, or the right edge line, or a straight line fitted from the left edge line or the edge line, and this application does not limit this; in the case where the baseline is the median line of the crop row, the straight line obtained by connecting the center of the upper edge and the center of the lower edge of the crop row can be used as the baseline.

[0054] However, to determine the above-mentioned baseline, it is necessary to identify the exact position of the crop row based on the fourth image, or to identify the left and right edge lines of the crop row. In one possible implementation, the third neural network model can be used to process the fifth image to obtain a second recognition result, and the second recognition result includes the baselines of multiple crop rows in the fifth image. Among them, the fifth image is the fourth image itself, or the fifth image is an image obtained by processing the fourth image. The processing here specifically refers to one or more of histogram equalization, filtering and denoising, image enhancement, and normalization processing, which is not limited in this application. Through these processing operations, the accuracy of subsequent image recognition can be improved.

[0055] In one example, the third neural network includes the structures of the first neural network model and the second neural network model involved in the subsequent scheme. The first neural network model in the present application is an image processing model for enhancing the image, and the second neural network model is a crop row recognition model, which can be used to identify the area where the crop rows are located. The superposition of the first neural network model and the second neural network model can be used to accurately identify the area where the crop rows are located. For details, please refer to the description of steps S120 and S140, which will not be repeated here. In other words, the third neural network model can accurately identify the area where the crop rows are located based on the fifth image, and then determine the baselines of n crop rows based on the area where the crop rows are located. This implementation method can identify the crop rows more accurately and determine a baseline with high accuracy based on this, thereby helping the agricultural harvesting robot to adjust the mechanical direction to a more appropriate orientation.

[0056] However, the above high-precision recognition scheme is bound to increase power consumption and latency. In view of this, the third neural network model in the present application may also include a structure of a lightweight neural network model for crop row detection. Lightweight neural network models are usually designed to run efficiently on devices with limited computing resources. Such models usually have the characteristics of a small number of parameters, small amount of calculation, and small model size, and can achieve fast inference and low-latency processing.

[0057] As a possible implementation method, the lightweight neural network model involved in the present application is a lightweight model obtained after compression processing for crop row identification, wherein the basic model for compression processing can be the above-mentioned high-precision neural network model including the first neural network model and the second neural network model structure, or it can be any other neural network model in the relevant technology that can be used for crop row identification, and the present application does not limit this.

[0058] A possible model compression method provided by the present application is exemplified below.

[0059] In one possible implementation, the purpose of compressing the model is achieved by eliminating the redundant structure of the model while ensuring that the accuracy of the model is within a preset range. In this implementation, it is first necessary to accurately identify the redundant structure of the model, otherwise the model after compression may have a problem of too low accuracy. The redundant structure of the model can be represented by the redundant channels in the convolutional layer of the model. In one example, a dependency coefficient can be calibrated to characterize the degree of dependence of the model on a certain channel, that is, the dependency coefficient corresponding to each channel in the model can be calculated, and then the channels whose dependency coefficients are less than the preset threshold are removed, thereby achieving the purpose of removing the redundant structure of the model. It is understandable that the output of a model is jointly determined by the input and model parameters, and the input is determined by the training set. Therefore, the input can be controlled to remain unchanged and the model output can be observed as the model parameters change. Therefore, in order to measure the degree of dependence of a certain channel on the model, the model parameters corresponding to the channel can be observed during the sparse training process. For example, during the coefficient training process, if the parameters corresponding to a certain channel gradually become smaller, it means that the channel has a weaker influence on the output of the model, so the dependence coefficient is also smaller; correspondingly, if the parameters corresponding to a certain channel gradually increase or remain unchanged, it means that the channel has a stronger influence on the output of the model, so the dependence coefficient is also larger. Among them, the calculation method of the dependence coefficient is not specifically limited in this application. After determining the dependence coefficient corresponding to each channel, the channels whose dependence coefficients are less than the preset threshold are subtracted to achieve the purpose of compressing the model.

[0060] It is understandable that the above scheme is explained by taking the redundant channels of the convolution layer as the redundant structure of the model as an example, but the present application is not limited to this. For example, the redundant layer in the model can also be used as the redundant structure of the model. In this case, a dependency coefficient can also be calibrated to characterize the degree of dependence of the model on a certain layer. The principle of the scheme is similar to the above scheme and will not be repeated here.

[0061] It can also be understood that the solution of the present application can remove only redundant channels or redundant layers, or can remove both redundant channels and redundant layers at the same time to compress the model to the greatest extent, and the present application does not limit this.

[0062] Through the above model compression method, a streamlined model without irrelevant structures and unnecessary parameters can be obtained. Although the model accuracy may decrease, the size of the model can be greatly reduced and the processing speed of the model can be improved.

[0063] It is understandable that the lightweight neural network model has the characteristics of fast processing speed and low resource consumption, but at the same time, the processing accuracy of the lightweight neural network model will be relatively low. However, the above characteristics are in line with the requirements of the third neural network model of this application, because the purpose of determining the crop row baseline in this application is to adjust the mechanical direction, and the adjustment of the mechanical direction does not require too high precision. Too high precision will cause a waste of resources. Even if the mechanical direction is adjusted by the baseline determined by the low-precision crop row recognition result, it can usually meet the needs, while also reducing resource consumption and latency, and has a high cost-effectiveness.

[0064] Optionally, in a possible implementation, after the above-mentioned lightweight model is obtained through compression processing, the model data can be further fine-tuned using knowledge distillation technology to restore the accuracy of the model. The specific implementation method is not limited in this application.

[0065] In another example, the lightweight neural network model involved in the embodiments of the present application can also be a customized model dedicated to determining the crop row baseline, that is, the model can also detect the crop row baseline directly instead of detecting the crop row area. For example, the third neural network model can automatically calculate the vegetation index of each pixel in the fifth image, and then automatically select the target area as the basis to determine the baseline. In the target area, the vegetation index of the pixels exceeding a preset proportion meets the preset requirements. In layman's terms, this solution is: the area with darker crop color is assumed to be the crop row area, and then the crop row baseline is determined based on this area. The accuracy of the baseline determined by this solution is lower, but the consumption is also minimal, which can basically meet the actual requirements.

[0066] Optionally, the fifth image may also be an image obtained by cropping the fourth image according to a preset size, and the fifth image obtained after cropping includes at least one independent crop row. In the solution of the present application, cropping may be performed without detecting the crop row, because the purpose of cropping is to improve the processing speed of the third neural network model (theoretically, the less the image data content of the fifth image, the faster the processing speed). In a possible implementation, the size of the preset size may be preset, and the preset size may be determined based on the type of crop and / or the size of the farmland where the crop is located. For example, based on the type of crop and / or the size of the farmland where the crop is located, it can be inferred that the fifth image usually includes at least x crop rows. In this case, if 1 / (x-1) of the fifth image is taken as the preset size, any interception can generally intercept at least one crop row. Of course, this is just an example. If taking 1 / (x-1) as the preset size still does not meet the requirements, other values ​​can be taken as the preset size according to actual conditions. Or, in another possible implementation, it can also be determined based on the type of crop and / or the size of the farmland where the crop is located.

[0067] The following is an illustrative description of possible implementations of an agricultural harvesting robot adjusting the mechanical direction according to the crop row baseline. As an example, adjusting the mechanical direction may include adjusting the position and orientation of the agricultural harvesting robot. For example, when a baseline is determined, the agricultural harvesting robot can be moved to the starting point at the bottom of the baseline, and then the orientation of the agricultural harvesting robot is adjusted to within a preset range of the error between the orientation of the agricultural harvesting robot and the direction of the baseline. As another example, adjusting the mechanical direction only includes adjusting the orientation of the agricultural robot, and the present application does not limit the position adjustment method of the agricultural harvesting robot. For example, when a baseline is determined, since the baseline is in an inclined state for the agricultural harvesting robot at the current position, the agricultural harvesting robot adjusts the mechanical direction based on the direction and angle of the baseline. For example, if the angle between the baseline and the horizontal line is 60°, the agricultural harvesting robot can adjust its orientation to an angle of 30° with the baseline to adjust the mechanical direction to a direction close to the vertical horizontal line, such as Figure 8 It should be understood that in actual application, there is a high probability that there will be a certain error in the adjusted mechanical direction, but this will not affect the execution of subsequent solutions.

[0068] S120. Process the first image based on the first neural network model to obtain a second image.

[0069] Exemplarily, after the first image is captured, the first image is not directly used for crop row detection, but the first image is first processed using the first neural network model to obtain the second image. The first neural network model is used to enhance the first image so that the features related to the crop rows in the first image are enhanced, while the features unrelated to the crop rows (such as weed features, soil features, light features, etc.) are weakened. This can greatly improve the subsequent crop row recognition effect, because these irrelevant features can cause interference. For example, weeds and soil related to crops can make it difficult to identify the edges of crop rows, and light scattered on crops at different times can also produce some shadows or highlights, making the crop recognition results inaccurate. The first neural network model involved in the present application can weaken the impact of these features and improve the accuracy of subsequent crop row recognition.

[0070] That is, the second image is an image obtained by enhancing the first image. Compared with the first image, the features related to crops in the second image are enhanced, for example, the color of the features related to crops is deepened, or the brightness is increased, or the contrast is enhanced; in addition, the features related to non-crops in the second image are weakened, for example, the color of the features related to non-crops is faded, or the transparency is increased.

[0071] The first neural network model in this application is a crop row feature enhancement model creatively proposed in combination with the scene of crop row recognition. Fig. 9 An example of a training method for the first neural network model is given.

[0072] The model structure of the first neural network model includes an encoder and a decoder. The encoder is used to obtain an input image and extract a feature vector from the input image. The feature vector is used to describe the characteristics of different contents in the image. The decoder is used to process the processed data and output an image. Therefore, the input of the encoder can be regarded as the input of the first neural network model, and the output of the decoder can be regarded as the output of the first neural network model. In other words, the input of the encoder is the original image to be processed, and the output of the decoder is the processed image.

[0073] The parameter training process of the first neural network model provided in the present application is essentially a training process for encoder parameters and decoder parameters. The following is an exemplary explanation of the training method of encoder parameters and decoder parameters: the decoder parameters are trained for the purpose of reducing the first loss function value, and the encoder parameters are trained for the purpose of reducing the second loss function value. Among them, the first loss function value is used to characterize the feature difference between the input image and the output image. Therefore, the smaller the first loss function value, the greater the difference between the input image and the output image. In other words, in the training process of the decoder parameters, it is hoped that the image difference is as large as possible. But we don’t want the image difference to be too large, and the image changes without any rules. Therefore, when we train the encoder parameters, we calculate a second loss function value. We hope that the second loss function value can make the image change in the direction we expect, that is, to enhance the crop-related features and weaken the non-crop-related features. In order to achieve this purpose, the present application proposes a new loss function value calculation method, that is, the composite calculation result of the first value and the second value is used as the second loss function value, the first value is used to characterize the feature similarity between the input image and the output image, and the second value is used to characterize the difference between the crop recognition result and the target recognition result for the input image. We train the encoder parameters with the purpose of reducing the second loss function value, so that the output image and the input image are as similar as possible, and the crop recognition process is as accurate as possible. That is to say, in the decoder training process, we hope that the input image and the output image are quite different; in the encoder training process, we hope that the input image and the output image are less different. This seems to be contradictory, but due to the existence of the second value, when the image changes in the direction of greater difference, the main change is the non-crop features, while the crop features are kept unchanged or changed as little as possible. Because the second value is used to characterize the difference between the crop recognition result and the target recognition result, and the smaller the second value (corresponding to the smaller second loss function value), the higher the recognition accuracy, where the crop recognition result here refers to the actual recognition result after the input image in the training data set is extracted through feature extraction and the crop row recognition is performed using the second neural network model. The second neural network model can be any existing crop recognition model, and the target recognition result refers to the real crop row recognition result corresponding to the input image, which can be obtained by manual calibration. As a result, the training process of the encoder and decoder forms an adversarial process, and the adversarial training process causes some features in the input image to undergo significant changes, while some features change slightly or do not change. Moreover, the crop recognition accuracy of the image after the change is significantly improved. Therefore, the result of the adversarial process will cause the features related to crops in the input image to be retained or even enhanced, while the features unrelated to crops will be weakened or even eliminated.After such a processing process, the second image obtained after being processed by the first neural network model will become an image that is more conducive to crop recognition, thereby improving the crop row recognition result.

[0074] In addition, when performing a composite calculation, a weight value will be configured for the first numerical value and the second numerical value, and the weight value is used to characterize the importance of the first numerical value and the second numerical value in the composite calculation, and this weight value can be determined based on the model training target of the first neural network model, and this model training target is actually associated with the current application scenario. For example, in one scenario, the final model obtained by training tends to be biased towards the accuracy of the final crop recognition result when applied, and does not care about the processing of the intermediate image. The weight of the second numerical value can be appropriately increased, while the weight of the first numerical value can be reduced. For another example, in another scenario, the final model obtained by training tends to be biased towards image enhancement results when applied, then the weight of the first numerical value can be appropriately increased, while the weight of the second numerical value can be reduced. In the application scenario of the present application, the first example is usually preferred, so the weight of the first numerical value can be set to be less than the weight of the second numerical value.

[0075] It is understandable that the first image may include only one image or multiple images (such as m images, where m is an integer greater than 1), which is not limited in this application. If the first image includes multiple images, step S120 may be performed on the multiple images in sequence, and then the multiple second images obtained may be fused to improve the effect of the final enhanced image. Alternatively, the image with the best effect (such as the image with the most uniform light and the highest clarity) may be selected from the multiple images to perform step S120, and the other images may be discarded. Alternatively, before executing step S120, the multiple images may be fused, and step S120 may be performed on the fused image. It is understandable that performing fusion processing of multiple images may refer to fusing image features to achieve the purpose of feature equalization and improve the effect of subsequent image processing.

[0076] S130: Perform a first operation on the second image based on the reflectivity of each pixel in the second image in a preset band to obtain a third image.

[0077] Exemplarily, after the first image is processed by the first neural network model to obtain the second image, the first operation can be further performed on the second image to further filter out non-crop related features. This is because for the second image, although the non-crop features are weakened, it cannot guarantee that a crop recognition result with sufficiently high accuracy can be achieved in the end, especially for the situation where the image includes some weeds and the weeds are intertwined with crops, and the weed features are similar to the crop features.

[0078] The first operation described in the embodiment of the present application is mainly used to further filter non-crop features. The first operation may specifically refer to clearing the pixels in the second image whose reflectivity exceeds a preset threshold value in the preset band. The first operation is explained below: First, in addition to crops, the images captured by the crop harvesting robot usually include soil, weeds, etc. These features are non-crop features and will most likely be weakened after the processing of step S120. The weakening of these features usually means that in the second image, the clarity of these non-crop features becomes lower or the color becomes lighter. On the contrary, the crop features may be enhanced, at least similar to the original feature strength. Based on this, different features can be distinguished according to the reflectivity of different pixels in different bands. The reason for this is that different colors, especially colors of different intensities, have different absorption effects in different bands (such as near-infrared band, blue light band, green light band, etc.). We can determine in advance through experiments the reflection and absorption effects of the mature colors of the crops involved in this application in different bands. Usually, the color of a single variety of crops will be within a certain range. For example, the reflectivity of the color corresponding to the crops in this application in the preset band is less than the preset threshold. Therefore, if the reflectivity of a pixel in the preset band is greater than or equal to the preset value, it means that the pixel does not belong to the crop and can be cleared. It should be understood that step S130 must be executed after step S120 to be effective, because step S120 can weaken the non-crop features in the image. Without this key step, some pixels are likely to be mistakenly cleared in step S130, because before the non-crop features are weakened, some original colors may be close to the crop colors, or have similar absorption effects in the preset band, especially some green crops, which are usually difficult to distinguish from other green weeds. Only after the feature weakening process of step S120 can the colors of crops and non-crops be distinguished in category and intensity, so that the absorption effects in the preset band can be significantly different, and the non-crop features can be effectively filtered through step S140, thereby improving the recognition effect of the subsequent step S140.

[0079] S140: Process the third image based on the second neural network model to obtain a first recognition result.

[0080] For example, after a series of processing steps S120 and S130, a third image is finally obtained, in which most non-crop features have been filtered, while crop features are basically retained or even enhanced. In this case, the recognition result can be improved by using the second neural network model to identify the area where the crops are located (or the crop row area) based on the third image. For example, the third image is processed using the second neural network model to obtain a first recognition result, and the first recognition result is used to indicate the area where the crops are located in the first image.

[0081] The present application does not limit the specific implementation method of the second neural network model to obtain the first recognition result. In one possible example, the second neural network model first maps each pixel in the third image to an n-dimensional vector in the feature space through a clustering loss function, so that the feature vectors of pixels belonging to the same crop row instance are closer, while keeping the feature vectors of pixels of different crop row instances farther and shorter, thereby distinguishing different crop rows. Finally, the first recognition result is further determined based on the mapping result. This scheme helps to distinguish different crop rows, and the specific method is not limited here.

[0082] S150: Harvesting the crops based on the first recognition result.

[0083] Exemplarily, after determining the first recognition result, the crops are harvested based on the first recognition result. For example, the agricultural harvesting robot controls the harvesting device to harvest the crops based on the first recognition result. Since the accuracy of the first recognition result is greatly improved, harvesting crops based on the first recognition result can reduce the omission of crops and reduce the harvested weeds and the like, thereby improving the efficiency and effect of crop harvesting.

[0084] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0085] An embodiment of the present application provides a computer program product. When the computer program product runs on a device, the device can implement the steps in the above-mentioned method embodiments when the device is executed.

[0086] The embodiment of the present application provides a chip, the chip is used to execute instructions, when the chip is running, the technical solution in the above embodiment is executed. The implementation principle and technical effect are similar, and will not be repeated here.

[0087] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0088] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0089] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0090] It should be understood that the "embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the various embodiments in the entire specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0091] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

[0092] In addition, it should be noted that the various numerical numbers involved in this application (such as the terms "first", "second", "third", "fourth" and other various terminology labels (if any) in the specification and claims and the above-mentioned drawings) are only distinguished for the convenience of description and are not used to limit the scope of this application. The size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic.

[0093] The terms "including" and "having" and any variations thereof mean "including but not limited to" unless specifically emphasized otherwise. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed but may include other steps or units not explicitly listed or inherent to such process, method, product or apparatus.

Claims

1. An intelligent harvesting method, characterized in that: The method is applied to an agricultural harvesting robot, and the method comprises: Capturing a first image, wherein the image content of the first image includes crops; The first image is processed based on a first neural network model to obtain a second image, wherein the decoder parameters of the first neural network model are obtained by training for the purpose of reducing a first loss function value, and the first loss function value is used to characterize the feature difference between an input image and an output image of the first neural network model; the encoder parameters of the first neural network model are obtained by training for the purpose of reducing a second loss function value, and the second loss function value is a composite calculation result of a first value and a second value, and the weight values ​​corresponding to the first value and the second value are determined based on a model training target of the first neural network model, the first value is used to characterize the feature similarity between the input image and the output image, and the second value is used to characterize the difference between a crop recognition result and a target recognition result for the input image; Performing a first operation on the second image based on the reflectivity of each pixel in the second image in a preset band to obtain a third image, wherein the first operation includes removing pixels in the second image whose reflectivity exceeds a preset threshold in the preset band, and the reflectivity of the color corresponding to the crop in the preset band is less than the preset threshold; Processing the third image based on the second neural network model to obtain a first recognition result, where the first recognition result is used to indicate an area where the crop in the first image is located; The crops are harvested based on the first recognition result.

2. The method according to claim 1, characterized in that: The third image is processed based on the second neural network model to obtain a first recognition result: Based on the second neural network model, each pixel in the third image is mapped to an n-dimensional vector in a feature space using a clustering loss function, and then the first recognition result is further determined based on the mapping result.

3. The method according to claim 1 or 2, characterized in that: Before capturing the first image, the method further includes: capturing a fourth image, wherein the fourth image also includes crops; The mechanical direction of the agricultural harvesting robot is adjusted according to the fourth image, and the angular deviation between the adjusted mechanical direction and the crop row to be harvested is within a preset range.

4. The method according to claim 3, characterized in that The step of adjusting the mechanical direction of the agricultural harvesting robot according to the fourth image comprises: Processing a fifth image using a third neural network model to obtain a second recognition result, the fifth image being the fourth image or an image obtained by processing the fourth image, the second recognition result comprising baselines of n crop rows, the third neural network model being a lightweight neural network model, and n being an integer greater than or equal to 1; The machine direction is adjusted using the baseline of the n crop rows as a guide line.

5. The method according to claim 4, characterized in that Before the fifth image is processed by the third neural network model to obtain the second recognition result, the method further includes: The fourth image is cropped to obtain the fifth image, where the fifth image is an image of a preset size and includes at least one independent crop row, where the crop row is composed of parts of the crops.

6. The method according to claim 4 or 5, characterized in that: The value of n is determined based on the type of the crop and / or the size of the farmland where the crop is located.

7. An agricultural harvesting robot, characterized in that: The agricultural harvesting robot comprises: A shooting module, used for shooting and obtaining a first image, wherein the first image includes crops; A model processing module, configured to process the first image based on a first neural network model to obtain a second image, wherein the decoder parameters of the first neural network model are obtained by training for the purpose of reducing a first loss function value, and the first loss function value is used to characterize the feature difference between the first image and the second image; the encoder parameters of the first neural network model are obtained by training for the purpose of reducing a second loss function value, and the second loss function value is a composite calculation result of a first value and a second value, and the weight values ​​corresponding to the first value and the second value are determined based on a model training target of the first neural network model, the first value is used to characterize the feature similarity between the first image and the second image, and the second value is used to characterize the difference between a crop recognition result and a target recognition result; a processing module, configured to perform a first operation on the second image based on the reflectivity of each pixel in the second image in a preset band to obtain a third image, wherein the first operation includes removing pixels in the second image whose reflectivity exceeds a preset threshold in the preset band, and the reflectivity of the color corresponding to the crop in the preset band is less than the preset threshold; a recognition module, configured to process the third image based on a second neural network model to obtain a first recognition result, wherein the first recognition result is used to indicate an area in the first image where the crop is located; A harvesting module is used to control a harvesting device to harvest the crops based on the first recognition result.

8. An agricultural harvesting robot comprising one or more processors and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program codes, wherein the computer program codes include computer instructions, and the one or more processors call the computer instructions to enable the agricultural harvesting robot to perform the method as described in any one of claims 1 to 7.

9. A chip system, characterized in that: The chip system is applied to an agricultural harvesting robot, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the agricultural harvesting robot executes the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises instructions, which, when executed on an electronic device, cause the electronic device to perform the method as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Crop row detection method and device based on deep learning image segmentation

    CN113128576A

  • Farmland contour detection method and device based on generative adversarial network

    CN114078213A

  • Image-based big data analysis method

    CN117333409A

  • Agricultural robot control method, device and equipment and storage medium

    CN119225376A