Image recognition system based on ambrosia trifida on grassland

By designing a multi-modular image recognition system in the grassland environment, using technical means such as image preprocessing, deep feature extraction, object separation and multi-scale enhancement, the problem of three-leaf ragweed recognition in complex grassland backgrounds was solved, and efficient and accurate recognition effect was achieved.

CN120071144APending Publication Date: 2025-05-30SHIHEZI UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510147129.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the complex plant background of grassland, how to accurately identify the tri-leaf ragweed and solve the image blur problem caused by plant occlusion and overlap has become a key challenge in the design of image recognition systems.

Method used

An image recognition system based on the three-leaved ragweed on the grassland is adopted, which includes an image preprocessing module, a multi-channel depth feature extraction module, an object separation and overlap elimination module, a multi-scale and context enhancement module, and a post-processing and post-correction module. These modules gradually improve image quality and feature recognition capabilities through technical means such as adaptive image enhancement, multi-channel input, depth image assisted separation, semantic separation and constraint optimization, multi-scale convolution and context perception, and solve the occlusion and overlapping problems.

Benefits of technology

It realizes the automatic and efficient identification of ragweeds in complex grassland environments, improves the efficiency of grassland ecological monitoring, enhances the identification ability of complex backgrounds and occlusion situations, and ensures the accurate identification and ecological management of ragweeds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071144A_ABST
    Figure CN120071144A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition system based on ambrosia trifida on a grassland. The image recognition system comprises adaptive image enhancement, de-noising and fuzzy processing and background segmentation; extracting multi-dimensional and multi-scale detail features of a plant by using a convolutional neural network and multi-channel input and combining convolutional layer fusion and feature map comprehensive learning; by combining a depth image, a segmentation reconstruction network and a semantic separation and constraint optimization technology, precise separation and independent identification of ambrosia trifida are realized; through a multi-scale convolutional network and a context sensing network, the recognition effect on a target in an occluded or complex background is improved by using context information at the same time; the output of a preorder module is subjected to refined optimization, and the target positioning precision and the classification accuracy are improved through non-maximum suppression, form correction based on geometric constraints and classification optimization. According to the method, challenges caused by plant shielding, overlapping and background complexity in the grassland environment are solved, and a more comprehensive and accurate monitoring result can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of image recognition, and in particular to an image recognition system based on ragweed on grasslands. Background Art

[0002] Three-lobed ragweed (Ambrosia trifida L.) is a common grassland weed that is widely distributed in temperate regions around the world, especially in my country's grasslands, farmlands, and roadsides. As a plant with extremely strong growth ability, three-lobed ragweed is regarded as a harmful plant in many areas because it can compete with crops for water and nutrients and pose a threat to biodiversity. In addition to its potential impact on agricultural production, three-lobed ragweed can easily spread to new areas through its gravity diffusion and the spread of cattle and sheep, causing serious disturbances to the ecological environment. Due to its rapid growth and strong adaptability, the monitoring and control of three-lobed ragweed has become an important issue that needs to be urgently addressed in grassland and agricultural ecological protection.

[0003] With the development of remote sensing technology, computer vision and image recognition technology, plant monitoring using image recognition has become an important tool in modern ecological research and agricultural management. Traditional manual monitoring methods are often inefficient and easily affected by human factors, while automated monitoring based on image recognition can not only improve work efficiency, but also more accurately identify and locate plant populations, especially in large-scale grassland and farmland monitoring, which can effectively solve the problem of difficulty in manual identification.

[0004] However, in practical applications, the complexity of grassland environment brings challenges to image recognition. The grassland has undulating terrain, a wide variety of plants, and changeable growth environment conditions. As a typical plant on the grassland, three-lobed ragweed often grows together with other herbaceous plants to form dense plant communities. In this environment, the plant targets in the image are often blurred due to the overlap of grass, mutual occlusion between plants, and interference from light and background, which affects the accuracy of the image recognition system. In particular, when part of the three-lobed ragweed is blocked by other plants or overlaps with other plants, it is difficult to clearly identify the independent form of the three-lobed ragweed in the image. This situation is more prominent in the process of large-scale image acquisition and processing, which greatly increases the difficulty of the recognition task. Therefore, how to accurately identify three-lobed ragweed in the complex plant background of the grassland and solve the image blur problem caused by plant occlusion and overlap has become a key challenge in the design of image recognition systems. Summary of the invention

[0005] In order to solve the above problems, the present invention provides an image recognition system based on Ambrosia trifida on grassland.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] An image recognition system based on Ambrosia trifida on the grassland, comprising:

[0008] An image preprocessing module: By means of adaptive image enhancement, denoising and blurring processing, and background segmentation, the image quality is improved, the characteristics of the target plant are highlighted, and high-quality input is provided for the subsequent modules;

[0009] A multi-channel depth feature extraction module: Using a convolutional neural network and multi-channel input, combined with convolutional layer fusion and feature map comprehensive learning, multi-dimensional and multi-scale detailed features of the plant are extracted, and the adaptability of the model to complex environments is enhanced;

[0010] An object separation and overlap elimination module: Combining depth images, segmentation and reconstruction networks, and semantic separation and constraint optimization technologies, the problems of plant occlusion and overlap are solved, and the precise separation and independent identification of Ambrosia trifida are realized;

[0011] A multi-scale and context enhancement module: Through a multi-scale convolutional network and a context-aware network, the recognition ability of plants at different scales is enhanced, and at the same time, the identification effect of the target in occlusion or complex backgrounds is improved by using context information;

[0012] A post-processing and post-correction module: Refined optimization is performed on the output of the previous modules. Through non-maximum suppression, morphological correction based on geometric constraints, and classification optimization, the target positioning accuracy and classification accuracy are further improved.

[0013] Furthermore: The image preprocessing module includes:

[0014] The image is enhanced by adaptive histogram equalization, and the brightness and details of the image are improved by adjusting the contrast of the local area, avoiding distortion caused by global enhancement;

[0015] Gaussian filtering and median filtering are used to remove the noise in the image, and the blurred area is repaired by wavelet transform to further improve the clarity of the image;

[0016] A segmentation algorithm based on color and texture is adopted, the RGB image is converted to the HSV color space, and combined with local binary pattern texture features, effective separation of the grassland background and the plant target is realized.

[0017] Furthermore: The multi-channel depth feature extraction module includes:

[0018] The convolutional neural network is used to extract image features layer by layer. Among them, the low-level convolutional layer extracts edge and texture information, the high-level convolutional layer identifies shape and structure features, and through the convolutional layer fusion strategy, feature maps at different levels are combined to enhance the richness of feature expression;

[0019] Introduce a multi-channel input mechanism, taking multiple data sources such as RGB images, depth maps, and thermal imaging maps as inputs. The RGB image provides color and shape information, the depth map is used to distinguish occluded and overlapping plants, and the thermal imaging map enhances the recognizability of plants through temperature differences;

[0020] Combine the Feature Pyramid Network and Spatial Pyramid Pooling technology. The Feature Pyramid Network captures multi-scale information by fusing low-level and high-level feature maps, improving the recognition ability for plants of different sizes. The Spatial Pyramid Pooling further enhances the comprehensive recognition of plant details through multi-scale pooling operations.

[0021] Furthermore: The object separation and overlap elimination module includes:

[0022] Based on the distance information provided by the depth image and combined with a generative adversarial network, precisely repair and reconstruct the occluded parts of the plants to restore their contours;

[0023] The segmentation and reconstruction network uses an encoder-decoder architecture to refine the image segmentation boundary through deconvolution operations, achieving precise separation of overlapping plant regions;

[0024] The semantic separation model combines the graph cut algorithm to optimize the overlapping regions, and precisely separates different plant targets and eliminates interference by minimizing the energy function.

[0025] Furthermore: The object separation and overlap elimination module includes:

[0026] Based on the distance information provided by the depth image and combined with a generative adversarial network, precisely repair and reconstruct the occluded parts of the plants to restore their contours;

[0027] The segmentation and reconstruction network uses an encoder-decoder architecture to refine the image segmentation boundary through deconvolution operations, achieving precise separation of overlapping plant regions;

[0028] The semantic separation model combines the graph cut algorithm to optimize the overlapping regions, and precisely separates different plant targets and eliminates interference by minimizing the energy function.

[0029] Furthermore: The post-processing and post-correction module includes:

[0030] Use non-maximum suppression to remove redundant or overlapping detection results. By calculating the intersection over union between candidate regions, retain the target region with the highest confidence, and remove overlapping detections caused by dense plant growth, thereby precisely determining the position and boundary of Ambrosia trifida.

[0031] Morphological correction based on geometric constraints analyzes the morphological features of the target area, combines morphological operations to repair misjudgments caused by occlusion or complex backgrounds, and further optimizes the target boundary.

[0032] Classification optimization uses a Softmax classifier to perform refined classification on the target area, calculates the probability of each target belonging to different plant categories, and sets a confidence threshold to exclude misclassification results, thereby improving the ability to distinguish Ambrosia trifida from other plants.

[0033] Compared with the prior art, the technological progress achieved by the present invention is as follows:

[0034] First, the system can automatically and efficiently monitor and identify Ambrosia trifida, greatly improving the efficiency of grassland ecological monitoring. Compared with traditional manual detection methods, image recognition technology can quickly process a large amount of image data in a vast grassland area, automatically identify the growth area of Ambrosia trifida, avoiding the cumbersome and error-prone manual annotation, and improving the accuracy and real-time of data collection. Especially when the grassland is vast and the number of monitoring objects is large, manual monitoring cannot meet the rapid and accurate requirements, while the image recognition-based system can effectively solve this problem.

[0035] Second, the image preprocessing module and multi-channel depth feature extraction module of the system can effectively improve the image quality and feature recognition ability, and solve common problems such as illumination changes and noise interference in the grassland environment. Through adaptive image enhancement, denoising processing, and multi-channel input (such as RGB images, depth maps, and thermal images), the system can improve the recognition ability of Ambrosia trifida in a changing grassland environment. Whether under complex climate conditions such as low light, high light, and haze, or in different vegetation backgrounds, the detailed features of Ambrosia trifida (such as leaf shape, number of lobes, etc.) can still be clearly extracted.

[0036] Furthermore, the system has significant advantages in dealing with plant occlusion and overlap problems. Plants on the grassland grow densely, and Ambrosia trifida often overlaps and occludes with other herbaceous plants, which directly affects the accuracy of image recognition. However, the system uses an object separation and overlap elimination module, combined with depth images and segmentation and reconstruction techniques, to accurately separate overlapping plant targets. Through the combination of depth image-assisted separation and the segmentation-reconstruction network, the system can not only detect obvious Ambrosia trifida targets but also process plants that are partially occluded or overlapped, ensuring that Ambrosia trifida can be accurately recognized even in complex backgrounds.

[0037] In addition, the system combines multi-scale convolution with a context enhancement module. By extracting plant features at different scales and enhancing context awareness, it further addresses the problem of plant variations at different scales and angles. Especially for Ambrosia trifida in different growth stages or when encountering different environmental factors, it can ensure its accurate positioning and recognition in the image.

[0038] In summary, the advantages of this application are not only reflected in improving the recognition efficiency and accuracy, but also in addressing the challenges brought by plant occlusion, overlap, and background complexity in the grassland environment through multiple technical means. It can provide more comprehensive and accurate monitoring results. Especially in an environment like the grassland with dense growth and a wide variety of plant species, the innovation and high robustness of the system provide reliable technical support for the accurate recognition and ecological management of Ambrosia trifida. Brief Description of the Drawings

[0039] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.

[0040] In the drawings:

[0041] Figure 1 It is the system structure diagram of the present invention. Detailed Description of the Embodiments

[0042] The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the drawings.

[0043] As Figure 1 shown, the present invention discloses an image recognition system for Ambrosia trifida on the grassland, including:

[0044] The implementation of the image preprocessing module includes:

[0045] The goal of the image preprocessing module is to perform multiple processes on the original image, improve the image quality, enhance the discernibility of plant features, and at the same time reduce the influence of factors such as illumination changes, noise, and background interference on subsequent image analysis and deep feature extraction. By processing multiple aspects of the image, it ensures that the image input to the recognition system can maximize the features of the target object (Ambrosia trifida), ensuring that the subsequent modules can accurately perform target detection and classification.

[0046] 1. Adaptive Image Enhancement

[0047] The adaptive histogram equalization method is used to enhance the local brightness and contrast of images, thereby reducing the impact of uneven illumination or strong light environments on plant features. This method can not only enhance the details of the image but also avoid the over-enhancement of contrast caused by global enhancement during the traditional histogram equalization process.

[0048] The implementation method is as follows:

[0049] First, convert the original image to the grayscale image space, and enhance the local image area by calculating the histogram of each local area (small window).

[0050] When using CLAHE, first set a "window size" (for example, 8x8 pixels), and perform local histogram equalization within this window.

[0051] To avoid over-enhancement causing image noise, set the contrast limit parameter, usually set to 2.0, to ensure that while the contrast of the enhanced image is improved, over-saturation is avoided.

[0052] Specifically:

[0053] Calculate the local area histogram, and define the pixel value P(x, y) within the local image window:

[0054]

[0055] The formula for local histogram equalization is:

[0056]

[0057] Among them, L is the number of gray levels of the image, usually 256.

[0058] In this way, the local contrast enhancement of the image is improved. Especially in low-light or strong-light conditions, the local details of the image can be clearly displayed, avoiding the excessive distortion caused by global enhancement.

[0059] 2. Denoising and Blurring Processing

[0060] Images often contain noise. Especially during the acquisition process, noise may be introduced due to environmental interference, sensor characteristics, etc., which will affect the recognition accuracy. Using Gaussian filtering and median filtering can effectively remove noise, and using wavelet transform to repair blurred areas can further improve the clarity of the image.

[0061] 3. Background Segmentation

[0062] The grassland environment is complex, and the background often has similar colors or textures to the plant targets, making it difficult to distinguish the plant targets from the background in the image. Therefore, background segmentation is crucial, especially for accurate plant detection. In this embodiment, a segmentation algorithm based on color and texture is adopted, which can effectively separate the grassland background from the plant targets through the combination of color space conversion and texture features.

[0063] The implementation method is as follows:

[0064] Color space conversion: Convert the RGB image to the HSV color space. In the HSV space, the color information and brightness information are separated, making it easier to distinguish plants from the background. For example, by setting appropriate thresholds to separate the green area, the grassland background and the Ambrosia trifida area can be distinguished.

[0065] Texture feature extraction: Use texture descriptors such as local binary pattern to perform image segmentation at the texture level. LBP is a method of binarizing the neighborhood of each pixel and calculating its texture features. Its mathematical expression is:

[0066]

[0067] where g p is the gray value of the neighborhood pixel, g c is the gray value of the central pixel, s(x) is the sign function, P is the number of neighborhood pixels, and the generated binary value represents the texture information by converting it to decimal.

[0068] Through the combination of color space conversion (RGB to HSV) and texture feature extraction, the background segmentation is not only based on color information but also considers texture features, enabling more accurate separation of the grassland background and plant targets and avoiding misjudgments that may occur in single-color segmentation.

[0069] The processing flow of the image preprocessing module is as follows:

[0070] 1. Original image input: Obtain the original image data captured by sensors or drones.

[0071] 2. Adaptive image enhancement: First, perform adaptive histogram equalization to improve the contrast and brightness of the image and enhance the local details of the plants.

[0072] 3. Denoising and blur restoration: Denoise the enhanced image, including Gaussian filtering and median filtering, and use wavelet transform to restore the blurred areas.

[0073] 4. Background segmentation: Separate the grassland background and plant targets through color space conversion (RGB to HSV) and texture feature extraction, reducing the interference of the background on the recognition process.

[0074] 5. Output preprocessed image: The image after the above processing will be provided as input to the subsequent deep feature extraction module for target detection and classification.

[0075] The image preprocessing module effectively improves the image quality through adaptive image enhancement, denoising and blurring, and background segmentation, ensuring that the characteristics of Ambrosia trilobata can be accurately extracted in complex grassland environments. These technical means complement each other, can minimize the negative impact of environmental changes on image recognition, provide high-quality input data, and thus lay a solid foundation for subsequent deep feature extraction and object recognition modules.

[0076] The implementation of the multi-channel deep feature extraction module includes:

[0077] The core goal of the multi-channel deep feature extraction module is to extract the detailed features of Ambrosia trilobata and other plants from the image through a deep convolutional neural network, especially in key areas such as leaf shape, edge, texture, etc. This module is not limited to a single visual information input, but through the design of multi-channel input, it combines multiple data sources (such as RGB images, depth maps, thermal images, etc.) to enhance the model's adaptability to environmental factors such as complex backgrounds, occlusions, and lighting changes, and improves the comprehensive recognition and extraction of plant details.

[0078] 1. Convolutional neural network + convolutional layer fusion

[0079] In the process of image feature extraction, convolutional neural network, as a deep learning model that can efficiently extract local and global features of images, undertakes the task of extracting features at all levels from images. The multi-layer convolutional structure of CNN can extract features from edges, textures to complex shapes and structures at multiple levels. After layers of abstraction and processing, plant features with high recognition are finally obtained.

[0080] The implementation is:

[0081] Convolutional layer feature extraction: CNN is usually composed of multiple convolutional layers, each of which can extract a specific type of feature. Low-level convolutional layers are generally used to extract basic information such as edges, lines, and colors; while high-level convolutional layers can recognize more complex shapes, structures, and even advanced features such as categories.

[0082] For example, the first convolutional layer can recognize simple shapes in the image such as edges and corners through convolution operations, while in the deeper convolutional layers, it can recognize the outline of leaves, the structure of flowers, and even the specific patterns of three-leaf ragweed.

[0083] Convolutional layer fusion: To maximize the extraction of image features at different levels, a convolutional layer fusion strategy can be added to the network structure to fuse convolutional feature maps at different levels and enhance the feature representation ability. There are usually two ways of fusion: concatenation fusion and weighted fusion. Concatenation fusion directly connects low-level feature maps and high-level feature maps to form a larger and richer feature representation; weighted fusion performs weighted summation on different feature maps according to their weights, thereby comprehensively considering the contributions of features at different levels.

[0084] The output of the convolutional layer is usually represented by the following formula:

[0085]

[0086] Among them, I(x,j) is the input image, K(i,j) is the convolutional kernel (filter), and Feature Map(x,y) is the feature map obtained by the convolutional operation. In a deep network, the feature maps of multiple convolutional layers are stacked layer by layer or weighted fused to obtain richer feature information.

[0087] 2. Multi-channel input mechanism

[0088] A single image data source (such as an RGB image) is vulnerable to interference from factors such as lighting, weather, and background weeds when facing complex environmental changes, resulting in unsatisfactory feature extraction effects. Therefore, the multi-channel input mechanism is one of the innovative designs of this module, which can combine feature information from multiple sources in different channels to improve the robustness and accuracy of the model.

[0089] The implementation method is as follows:

[0090] RGB image: As a traditional image input method, the RGB image provides color information and basic shape information in the image. Although the RGB image captures visual information relatively comprehensively, it is often sensitive to environmental changes such as lighting changes and shadows.

[0091] Depth map: The depth map provides the object distance information in the image, which helps to effectively distinguish different plants in the case of occlusion or overlap. In the depth map, the relative positions and shape differences between Ambrosia trifida and other plants are clearly represented, which can help the network effectively identify and separate target objects in the image.

[0092] Thermal imaging map: By capturing the heat difference between plants and the surrounding environment, the thermal imaging map can provide additional feature information under certain environmental conditions (such as when the temperature changes greatly). For example, at dawn or dusk, the difference in thermal radiation between plants and the surrounding ground can significantly improve the recognizability of plants.

[0093] By taking these image sources (RGB images, depth maps, thermal images) as inputs for different channels, the network can perform feature learning in multiple dimensions, not only enhancing the ability to recognize plant targets but also improving the adaptability to factors such as lighting and environmental changes.

[0094] Mathematical representation of multi-channel input:

[0095] Assume the multi-channel input image is \(X = \{X 1 , X 2 , \cdots, X n \}\), where each \(X i \) represents a different channel (such as RGB, depth, thermal imaging, etc.). Then the convolution operation on each channel is represented as:

[0096]

[0097] where \(X i (x, y)\) is the input image in the \(i\)-th channel, \(W k \) is the convolution kernel, and \(Output(X i )\) is the feature map obtained through convolution. Finally, the feature maps of all channels are combined or weighted and fused to generate the final multi-channel output:

[0098]

[0099] where \(\alpha i \) is the weight coefficient for each channel.

[0100] 3. Feature map fusion and comprehensive learning

[0101] Based on multi-channel input and extraction by the convolutional layer, how to fuse feature maps of different channels is the key to achieving multi-dimensional feature learning. Through a deep feature fusion network, feature maps from different sources can be comprehensively processed, thereby maximizing the accuracy and richness of feature expression.

[0102] The implementation methods are as follows:

[0103] Deep feature fusion: By combining low-level feature maps and high-level feature maps through a feature map fusion network (such as FPN, Feature Pyramid Network), a feature representation containing multi-scale information is formed. At this time, low-level feature maps are helpful for capturing details, and high-level feature maps provide higher-level semantic information. The combination of the two can improve the comprehensive recognition of plant details.

[0104] Spatial pyramid pooling: The spatial pyramid pooling network can capture feature information at various scales in the image through pooling operations at different scales, thereby enhancing the ability to recognize plants at different scales.

[0105] After combining multi-channel inputs, the system can fuse multiple types of information such as traditional RGB image features, depth information, and temperature differences. This enables recognition to not only rely on visual information but also fully consider the influence of external environmental factors. Thus, even in a complex grassland background or when plants are partially occluded, the system can effectively distinguish Ambrosia trifida from other plants, thereby improving the recognition accuracy.

[0106] The multi-channel depth feature extraction module uses convolutional neural networks and convolutional layer fusion techniques to extract detailed features at different levels in the image layer by layer. Through a multi-channel input mechanism, it combines multiple data sources such as RGB images, depth maps, and thermal images to provide a more comprehensive and robust feature representation. This multi-dimensional and multi-scale feature extraction method can not only handle problems such as illumination changes and plant occlusion in complex environments but also enhance the system's ability to accurately recognize Ambrosia trifida, providing richer and more accurate feature inputs for subsequent object detection and classification.

[0107] The implementation of the object separation and overlap elimination module includes:

[0108] The core objective of the object separation and overlap elimination module is to effectively solve the problem of target recognition in the grassland environment caused by plant mutual occlusion and overlap, ensuring that Ambrosia trifida can still be accurately recognized and marked even when there is strong occlusion and overlap between plants. To achieve this goal, this module will combine multiple technical means, including occlusion separation technology based on depth images, segmentation and reconstruction networks, and semantic separation and constraint optimization algorithms. Guided by fine image processing and depth information, it completes the separation of overlapping plants and the elimination of interference.

[0109] 1. Occlusion separation based on depth images

[0110] In the grassland environment, plants often occlude each other due to dense growth, resulting in overlapping regions in the image and making the target difficult to identify. By introducing depth images and combining generative adversarial network (GAN) technology, this embodiment can accurately repair and reconstruct the occluded parts based on depth information, thereby achieving the three-dimensional separation of plants.

[0111] The implementation method is as follows:

[0112] Depth image-assisted separation: Depth images can provide distance information between various objects in the scene, helping to identify the spatial relationship between occluded plants and other plant targets. By fusing depth information in a deep learning network, the occluded area can be distinguished from the foreground objects.

[0113] Generative Adversarial Network (GAN): The generative adversarial network is used to perform inpainting on the overlapping regions. The generator generates inpainted images similar to the real images, while the discriminator evaluates the authenticity of the generated images. The two play against each other and gradually optimize the generated inpainted images.

[0114] Objective function of GAN: The goal is to make the generated images restore the occluded plant morphology as much as possible without generating excessive artifacts in the surrounding environment. The following loss function can be used:

[0115]

[0116] where G(z) is the image output by the generator, D(x) is the judgment of the real image output by the discriminator, p data (x) and p z (z) are the data distribution and the noise distribution respectively.

[0117] Guided by the generative adversarial network, the depth image can help restore the part of the plant information that disappears due to occlusion. Especially in the case of severe plant overlap, it can effectively separate the occluded objects.

[0118] 2. Segmentation and Reconstruction Network

[0119] After separating the objects, the next step is to accurately divide the overlapping plants into regions. For this purpose, this embodiment uses a segmentation and reconstruction network, a deep learning model. The segmentation and reconstruction network can effectively separate the overlapping plant regions by refining the segmentation boundaries in the image. Especially in the case of complex backgrounds and diverse plant morphologies, it can improve the segmentation accuracy.

[0120] Implementation method:

[0121] Segmentation network architecture: The segmentation and reconstruction network model consists of an encoder and a decoder. The encoder extracts the feature information in the image, and the decoder gradually restores the spatial resolution of the image. During the decoding process, the segmentation and reconstruction network can use deconvolution operations to restore the high-dimensional feature maps to the spatial resolution of the image and accurately mark the boundaries of the plants.

[0122] Segmentation loss function: To optimize the segmentation accuracy, this embodiment uses the cross-entropy loss function to measure the difference between the segmentation result and the true label:

[0123]

[0124] where y c is the true label of class c, p c is the class probability predicted by the network, and C is the total number of classes.

[0125] Region refinement: Reconstruct and refine the segmented image to further improve the accuracy of the boundary. For example, using the superpixel segmentation method, the image is divided into several small regions, and then fine reconstruction is performed within these regions to avoid over-fusion and mis-segmentation.

[0126] 3. Semantic Separation and Constraint Optimization

[0127] In multi-object plant recognition, especially in overlapping plant regions, in addition to image segmentation, it is crucial to accurately classify and separate different plants. The semantic separation model and optimization algorithm can further enhance the processing ability of overlapping regions, and through the combination of category recognition and spatial constraints, accurately separate objects and eliminate interference.

[0128] The implementation method is as follows:

[0129] Semantic separation model: In the object recognition of overlapping regions, first, the plant categories are initially recognized through the semantic separation model to determine the category labels of each object. This model can extract high-level semantic information in the image through a deep neural network and identify the differences between Ambrosia trifida, other herbaceous plants, and the background.

[0130] Constraint optimization (GraphCut algorithm): When separating objects, the GraphCut algorithm is used to optimize the plant targets. The GraphCut algorithm minimizes the energy function, so that the plants in the overlapping region are correctly divided into different targets. The energy function of the GraphCut algorithm is expressed as:

[0131]

[0132] where ω ij is the similarity between pixel pairs (i, j), λ i is the constraint coefficient of pixel i, and f(I i ) is the category label function of pixel i.

[0133] Optimization process: The GraphCut algorithm adjusts the category labels and region boundaries of pixels to minimize the total energy function, thereby accurately separating overlapping plant targets, eliminating unnecessary interference, and ensuring that each plant target is independently identified.

[0134] 4. Overall Process

[0135] The processing flow of the object separation and overlap elimination module can be summarized as the following steps:

[0136] 1. Input image and depth map: The original image and depth image are used as input data. After preprocessing by the depth network, effective depth information is extracted.

[0137] 2. Occlusion separation of depth images: Utilize depth image information and generative adversarial networks (GANs) to repair and reconstruct occluded plant information, and restore the outline of the object as much as possible.

[0138] 3. Image segmentation and reconstruction: Use a segmentation and reconstruction network to finely segment the image, mark the boundaries of the plant area, and further refine the boundaries through reconstruction techniques.

[0139] 4. Semantic separation and constraint optimization: Combine a semantic separation model to classify plant targets, and optimize overlapping regions through graph cut algorithms. Finally, separate plant targets independently and eliminate interference.

[0140] 5. Output separation results: After the above processing, output a clearly separated plant image, in which Ambrosia trifida and other plant targets are accurately identified.

[0141] The object separation and overlap elimination module can effectively handle plant occlusion and overlap problems in complex grassland environments by combining depth image-assisted separation, segmentation and reconstruction networks, and semantic separation and constraint optimization technologies, accurately separating different plant targets, especially Ambrosia trifida, thus achieving more accurate target detection and recognition. Through the collaborative work of multiple deep learning algorithms, this module significantly improves the ability to separate and recognize plant targets in dense grassland environments.

[0142] The implementation of the multi-scale and context enhancement module includes:

[0143] The main goal of the multi-scale and context enhancement module is to solve the recognition difficulties of plants in images due to scale changes, angle changes, and partial occlusions. Especially for plants with slender shapes like Ambrosia trifida, by combining multi-scale feature extraction and context information enhancement, even when the plant size is small, in a complex background, or partially occluded by other plants, the system can still accurately recognize its features. To achieve this goal, this module uses technical means such as multi-scale convolutional networks (such as FPN), region aggregation strategies, and context-aware networks, working together, through multi-dimensional feature learning and context enhancement, to maximize the recognition effect.

[0144] 1. Multi-scale convolution and region aggregation

[0145] In grassland environments, Ambrosia trifida and other plants may present different scales in images due to differences in distance, angle, or growth patterns. To address this challenge, this embodiment introduces multi-scale convolution technology, uses a feature pyramid network (FPN) to extract features of plants at various scales in the image, and combines a region aggregation strategy to ensure that the system can extract sufficient information from multiple scales regardless of the size of the plant target.

[0146] The implementation method is as follows:

[0147] Feature Pyramid Network (FPN): FPN extracts features at different levels by performing convolutional operations on images at different scales. In the FPN architecture, the low-level convolutional layers are responsible for extracting detailed information in the image (such as edges, textures, etc.), while the high-level convolutional layers can extract more abstract semantic information (such as shapes, categories, etc.). By fusing these features at different levels, FPN can capture multi-scale information from details to the global in the image, thereby helping the system identify ragweed and other plants of different sizes.

[0148] Region aggregation: Region aggregation is a process of effectively fusing features at different scales. For plant targets of different sizes, especially when the target is partially occluded, far from the camera, or covered by background weeds, region aggregation can perform weighted aggregation on features from multiple scales, ensuring that sufficient context information can still be extracted even when some details are lost. This aggregation not only enhances the recognition ability of small targets (such as distant ragweed), but also avoids misjudgment caused by over-focusing on local features.

[0149] In FPN, the fusion of multi-scale feature maps is expressed as:

[0150]

[0151] Among them, represents the i-th feature map generated by the j-th convolution, and α ij is the weight coefficient of each feature map, is the fused feature map, and N is the number of feature layers for fusion.

[0152] 2. Context-Aware Network

[0153] For the situation where plants in the grassland overlap or are partially occluded, a single local feature is often insufficient to accurately identify ragweed. Therefore, in this embodiment, a context-aware network is introduced to enhance the model's perception of the environment around the plant target, especially its recognition ability in complex backgrounds. The context-aware mechanism can weightedly learn the information in the surrounding area of the target, thereby helping the system accurately identify ragweed even when the target is partially occluded by other plants.

[0154] The implementation method is as follows:

[0155] Context-aware mechanism: The context-aware network learns the relationships between various regions in the image through the self-attention mechanism, especially the spatial relationships between plants and the background, and between plants. This mechanism can dynamically adjust the focus of attention for each region, assign appropriate weights to the features of the surrounding environment, and thus enhance the ability to identify targets in complex backgrounds. Specifically, the network calculates the dependence of each pixel on other pixels to generate an attention weight map, and then applies the weight map to the image features to strengthen the attention to the regions around the target.

[0156] Self-attention mechanism: The self-attention mechanism can capture long-range dependence relationships, identify the mutual influence between the target and its surrounding background, so as to better separate overlapping regions and retain the detailed information of the target. For Ambrosia trifida, especially when it is partially blocked by other plants, the model can adjust the recognition strategy through context awareness and maintain the attention to the target.

[0157] The calculation of the self-attention mechanism can be expressed by the following formula:

[0158]

[0159] where Q and K are the query and key matrices respectively, A is the attention weight matrix obtained through self-attention calculation, and d k is the dimension of the key vector. By performing a softmax operation on the weight matrix, weighted attention to different regions is obtained.

[0160] Context information weighting: By calculating the attention map of the context, the model can weight the features of the regions around the target, improving the recognition ability of Ambrosia trifida that is occluded or far away:

[0161] F countextual = A·F

[0162] where A is the context attention map, F is the original image feature, and F countextual is the feature enhanced by context awareness.

[0163] This module realizes the accurate recognition of Ambrosia trifida by combining multi-scale convolution and context-aware network, especially in the case of large morphological changes of plants, partial occlusion or being far away from the camera. The multi-scale convolution network (such as FPN) extracts features from different scales to ensure the recognition of plants of different sizes, while the context-aware mechanism effectively improves the processing ability of occluded regions by enhancing the understanding of the surrounding environment of plants. Through this multi-dimensional feature learning and context information weighting, the system can better identify Ambrosia trifida in the grassland and overcome the interference of complex backgrounds and occlusions on target recognition.

[0164] The multi-scale and context enhancement module combines a multi-scale convolutional network (such as FPN) with a context-aware network, which not only solves the problem of plant target recognition at different scales and angles, but also effectively enhances the model's recognition ability in complex backgrounds and occlusion situations. By introducing a context-aware mechanism and a region aggregation strategy, the module can accurately extract the features of Ambrosia trifida in various environmental changes, especially in the case of overlap, partial occlusion, or being far and small, and maintain a high recognition accuracy, thus greatly improving the recognition ability of the system.

[0165] The implementation of the post-processing and post-correction module includes:

[0166] In the entire image recognition process, the goal of the post-processing and post-correction module is to further refine and optimize the results obtained by the previous modules (image preprocessing, feature extraction, object separation, multi-scale enhancement, etc.) to ensure more accurate final target recognition and localization. This module mainly focuses on improving the localization accuracy and classification accuracy of the target, eliminating overlapping and redundant recognition results, and ensuring the accuracy and robustness of Ambrosia trifida in complex environments.

[0167] 1. Module input

[0168] The input of this module is the output of the previous several modules:

[0169] The image preprocessing module provides a high-quality image after denoising and enhancement;

[0170] The multi-channel depth feature extraction module extracts the depth features of the image;

[0171] The object separation and overlap elimination module generates a relatively accurate target area (not necessarily in the form of a candidate box), that is, the separated plant area;

[0172] The multi-scale and context enhancement module further strengthens the recognition ability of Ambrosia trifida at different scales and in complex backgrounds.

[0173] Therefore, the input of the post-processing and post-correction module is these feature maps and target area information.

[0174] 2. Non-maximum suppression

[0175] It is used to remove redundant or overlapping detection results. In the scenario of this embodiment, due to the dense growth of plants on the grassland, there may be some overlapping or repeated areas in the separated target areas, especially when Ambrosia trifida is close to other plants. Non-maximum suppression effectively removes redundant overlapping areas by calculating the overlap degree (IoU) between candidate regions and retains the regions most likely to be Ambrosia trifida.

[0176] The implementation process includes:

[0177] 1. The target regions obtained from the multi-channel depth feature extraction and object separation module are sorted according to the confidence level (the predicted probability by the classifier).

[0178] 2. Calculate the intersection over union (IoU) between each pair of candidate regions:

[0179]

[0180] where, b i and b j are candidate bounding boxes, and IoU is the degree of overlap between them.

[0181] 3. For redundant bounding boxes with IoU values exceeding a certain threshold, retain the bounding box with a higher confidence level and remove the region with a larger overlap.

[0182] The core purpose of non-maximum suppression is to reduce redundant results in object detection. In this way, the position and boundary of Ambrosia trifida can be accurately determined.

[0183] 3. Morphological correction based on geometric constraints

[0184] In the image, the morphological features of Ambrosia trifida (such as the number of lobes of the leaves and the shape of the flowers) have certain rules. Through morphological correction based on geometric constraints, this embodiment can further optimize the accuracy of the target and repair the problem of inaccurate morphological recognition caused by occlusion or complex background.

[0185] The implementation method is as follows:

[0186] 1. Number of lobes and morphological correction: The leaves of Ambrosia trifida usually have a definite number of lobes. During the recognition process, through morphological analysis of the target region and combining with known plant morphological features, the shape deviation is corrected. For example, morphological dilation and erosion operations can be used to finely adjust the boundary of the target.

[0187] Morphological erosion: For the target region, perform erosion operation to remove the parts that do not conform to the morphology of Ambrosia trifida.

[0188] Morphological dilation: For the target region, perform dilation operation to fill in the parts lost due to occlusion.

[0189] 2. Geometric constraint optimization: According to the geometric features of the target region, such as the shape of the leaves and the angle of the lobes, apply geometric constraints for position optimization. For example, if the number of lobes of a certain candidate target does not match the expectation, or the morphology of the petals does not conform to the characteristics of Ambrosia trifida, this embodiment can perform morphological correction through geometric constraints to ensure that the final result conforms to the actual characteristics of the plant.

[0190] 3. Calibrated Target: After geometric optimization, misjudgments caused by complex backgrounds or partial occlusions are corrected, and the boundaries of the target are more precise.

[0191] 4. Classification Optimization and Refinement

[0192] Although relatively accurate target regions have been obtained through feature extraction and object separation in the previous modules, in the actual recognition process, due to the wide variety of plant species, misclassification may still occur. Through classification optimization, this embodiment can further improve the ability to distinguish Ambrosia trifida from other plant categories and avoid misjudgment.

[0193] The implementation process includes:

[0194] 1. Softmax Classification Optimization: For each target region, use the softmax classifier to further refine the classification. By calculating the probability values of each target belonging to various plant categories, ensure that each region can be accurately classified as Ambrosia trifida.

[0195]

[0196] Among them, P(y i = c k | x i ) is the probability that the target x i belongs to the category c k , w k is the weight of the category c k , x i is the feature vector of the target region, and C is the number of all categories.

[0197] 2. Threshold Optimization: To avoid misjudgment, set an appropriate classification confidence threshold. When the probability value of the classification result is lower than a certain threshold, the target is excluded to reduce the risk of misclassification.

[0198] Based on the deep features output by the first four modules and the separated target regions, the post-processing module first finely calibrates the target regions and performs localization optimization to ensure that the boundaries of Ambrosia trifida are more accurate. Combining the geometric morphological characteristics of Ambrosia trifida, apply morphological operations to correct misjudgments caused by occlusion or background complexity. Further improve the accuracy of classification through the softmax classifier to avoid confusion between Ambrosia trifida and other plants.

[0199] This embodiment combines geometric constraints with morphological analysis, enabling plant targets in complex backgrounds to be accurately recognized and located, greatly improving the recognition accuracy.

[0200] The post - processing and post - calibration module removes redundant regions through non - maximum suppression, performs morphological calibration on the target region by combining geometric constraints and morphological analysis, and finally ensures the accurate distinction between Ambrosia trifida and other plants through classification optimization. This module plays a role in fine - tuning and optimizing the results in the entire image recognition system, improving the recognition accuracy and positioning accuracy.

[0201] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. An image recognition system based on Ambrosia trilobata on grassland, characterized in that: include: Image preprocessing module: Improve image quality, highlight target plant features, and provide high-quality input for subsequent modules through adaptive image enhancement, denoising, blurring, and background segmentation; Multi-channel deep feature extraction module: using convolutional neural networks and multi-channel input, combined with convolutional layer fusion and feature map comprehensive learning, to extract multi-dimensional and multi-scale detailed features of plants, and enhance the model's ability to adapt to complex environments; Object separation and overlap elimination module: Combines deep images, segmentation and reconstruction networks, and semantic separation and constraint optimization technology to solve plant occlusion and overlap problems, and achieve accurate separation and independent identification of trilobate ragweed; Multi-scale and context enhancement module: Through multi-scale convolutional networks and context-aware networks, the recognition ability of plants of different scales is enhanced, and the recognition effect of targets in occluded or complex backgrounds is improved by using context information; Post-processing and post-correction module: refines and optimizes the output of the previous module, and further improves the target positioning accuracy and classification accuracy through non-maximum suppression, geometric constraint-based morphological correction and classification optimization.

2. The image recognition system based on Ambrosia trilobata on grassland according to claim 1, characterized in that: The image preprocessing module comprises: The image is enhanced through adaptive histogram equalization. The brightness and detail of the image are improved by adjusting the contrast of the local area, avoiding the distortion caused by global enhancement. Use Gaussian filtering and median filtering to remove noise from the image, and use wavelet transform to repair the blurred area to further improve the clarity of the image; A color and texture-based segmentation algorithm is used to convert the RGB image into the HSV color space, and the local binary pattern texture features are combined to achieve effective separation of grassland background and plant targets.

3. The image recognition system based on Ambrosia trilobata on grassland according to claim 2, characterized in that: The multi-channel depth feature extraction module includes: Convolutional neural networks are used to extract image features layer by layer. The low-level convolutional layers extract edge and texture information, and the high-level convolutional layers recognize shape and structural features. Through the convolutional layer fusion strategy, feature maps at different levels are combined to improve the richness of feature expression. A multi-channel input mechanism is introduced, which takes multiple data sources such as RGB images, depth maps and thermal images as input. RGB images provide color and shape information, depth maps are used to distinguish occluded and overlapping plants, and thermal images enhance the recognizability of plants through temperature differences. Combining feature pyramid network and spatial pyramid pooling technology, the feature pyramid network captures multi-scale information by fusing low-level and high-level feature maps, improving the recognition ability of plants of different sizes. Spatial pyramid pooling further enhances the comprehensive recognition of plant details through multi-scale pooling operations.

4. The image recognition system based on Ambrosia trilobata on grassland according to claim 3, characterized in that: The object separation and overlap elimination module includes: Through the distance information provided by the depth image and the generative adversarial network, the occluded plant parts are accurately repaired and reconstructed to restore their contours. The segmentation and reconstruction network uses an encoder-decoder architecture to refine the image segmentation boundaries through deconvolution operations to achieve accurate separation of overlapping plant areas; The semantic separation model is combined with the graph cut algorithm to optimize the overlapping area and accurately separate different plant targets and eliminate interference by minimizing the energy function.

5. The image recognition system based on Ambrosia trilobata on grassland according to claim 4, characterized in that: The multi-scale and context enhancement module includes: In terms of multi-scale feature extraction, the module uses a feature pyramid network to combine high-level semantic information and low-level detail information through bottom-to-top and top-to-bottom paths to generate a feature map containing multi-scale information; A context-aware network is introduced to enhance the perception of the target's surrounding environment through the self-attention mechanism. The self-attention mechanism dynamically adjusts the attention weights of different areas and captures long-distance dependencies, so that three-leaved ragweed can still be accurately identified when the target is partially occluded or in a complex background.

6. The image recognition system based on Ambrosia trilobata on grassland according to claim 5, characterized in that: The post-processing and post-correction module includes: Non-maximum suppression is used to remove redundant or overlapping detection results. By calculating the intersection-over-union ratio between candidate regions, the target region with the highest confidence is retained, and overlapping detections caused by dense plant growth are removed, thereby accurately determining the location and boundary of three-leaved ragweed. Morphological correction based on geometric constraints analyzes the morphological features of the target area and combines morphological operations to repair misjudgments caused by occlusion or complex background, further optimizing the target boundary. Classification optimization uses the Softmax classifier to refine the classification of the target area, calculate the probability that each target belongs to a different plant category, and set the confidence threshold to exclude misclassification results, thereby improving the ability to distinguish three-leaved ragweed from other plants.

Citation Information

Cited By

  • Irrigated area water-saving irrigation control system fused with Internet of Things technology

    CN120959132A