Intelligent monitoring method and system for micro-insects on fruits and vegetables based on microscopic magnification imaging

CN122821587APending Publication Date: 2026-09-25PLANT PROTECTION RES INST OF GUANGDONG ACADEMY OF AGRI SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610926193.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

人工巡查和显微镜复查高度依赖操作人员的经验积累和持续专注度,劳动强度大、效率低下,难以支撑大面积、高频次的虫情普查需求,且主观判断差异导致数据的一致性和可比性难以保证

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821587A_ABST
    Figure CN122821587A_ABST
Patent Text Reader

Abstract

The application discloses a kind of fruit and vegetable micro-insect intelligent monitoring method and system based on microcosmic amplification imaging, comprising: adjustable micro-distance amplification acquisition module is close to leaf or stick insect board, focus stack acquisition is carried out, fusion generates visible light texture image and autofluorescence image;Micro target detection network is constructed, through channel attention fusion bimodal feature, enhanced feature pyramid is formed by separable convolution and spatial offset, and dense classification regression of anchor frame is carried out to extract insect candidate area;For candidate area, depth topography is generated by white light focus stack fitting, and texture and fluorescence image are stacked and sent into multi-task instance segmentation network, and dense adhesion insect segmentation is realized;Based on segmentation instance, each insect area block is extracted and input into three branch mixed classification network to identify insect species and distinguish between live and dead;Population dynamics modeling and prediction are carried out by collecting multi-modal monitoring data, and insect situation early warning is generated by introducing bayesian inference, to enhance the monitoring robustness under complex imaging conditions in field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural pest and disease monitoring technology, and in particular to an intelligent monitoring method and system for tiny insects in fruits and vegetables based on microscopic magnified imaging. Background Technology

[0002] In large-scale production in greenhouses and open orchards, millimeter-sized insects such as thrips, whiteflies, cotton aphids, two-spotted spider mites, and fungus gnats are the most damaging pests. Adults typically range from 0.8 to 4 millimeters in length, with nymphs even smaller, often hiding on the undersides of leaves, in bud crevices, or inside calyxes, making early damage symptoms extremely subtle. Currently, commonly used monitoring methods include manual visual inspection, periodically replacing yellow or blue sticky traps and then manually counting and identifying insects using a stereomicroscope, and automated image monitoring devices based on ordinary visible light cameras. Manual inspection and microscopic examination heavily rely on the operator's experience and sustained focus, resulting in high labor intensity and low efficiency, making it difficult to support large-scale, high-frequency pest surveys. Furthermore, subjective judgment differences make it difficult to guarantee data consistency and comparability. Although monitoring devices based on ordinary visible light cameras have achieved automated image acquisition, their optical systems are limited by fixed focal length and conventional magnification. When the target insect is less than a few millimeters long, the imaging area occupies only a very small number of pixels on the sensor. Key distinguishing features such as wing veins, antennae, body surface punctures, and color spots are completely unrecognizable in the image. Different species of tiny insects are highly similar in shape, making automatic classification based on visual features almost impossible.

[0003] Furthermore, insects on sticky traps often exhibit high-density aggregation, with individuals sticking together and even stacking on top of each other. Combined with unavoidable factors in the field environment such as fog condensation, leaf reflection, irrigation dust, and pesticide residue spots, these factors severely interfere with image quality, further compressing the effective visual signal of these tiny insects. Even with the introduction of high-magnification microscope lenses, their inherently shallow depth of field, coupled with the millimeter-level undulations naturally present on fruit and vegetable surfaces and the unevenness of the sticky trap layer, means that only a localized area in a single frame image can be clearly focused. A large amount of insect information is obscured by the blurred foreground or background. Simultaneously, dead insects, impurities, and lesions are highly similar to live insects in static microscopic images, making them difficult to distinguish effectively with current techniques. These multiple challenges combine to create the core pain points of current fruit and vegetable micro-pest monitoring: inability to see clearly, distinguish, or accurately count pests. There is an urgent need for an intelligent monitoring method that can stably acquire fully focused microscopic images under real field conditions, accurately separate densely clustered individuals, and reliably identify live pests. Summary of the Invention

[0004] This invention overcomes the shortcomings of the prior art and provides a method and system for intelligent monitoring of tiny insects in fruits and vegetables based on microscopic magnification imaging. Its important purpose is to improve the identification accuracy of live pests at the millimeter level and the accuracy of separating and counting densely clustered individuals, thereby enhancing the robustness of monitoring under complex imaging conditions in the field.

[0005] To achieve the above objectives, the first aspect of this invention provides an intelligent monitoring method for tiny insects on fruits and vegetables based on microscopic magnification imaging, comprising: The adjustable macro magnification acquisition module is placed close to the surface of fruit and vegetable leaves or sticky insect boards. White light and ultraviolet light are irradiated by time-division switching, and the focus is adjusted in equal steps within a preset depth range along the optical axis to acquire focus stacks. The focus stacks are fused by non-downsampling contour wave transformation to generate visible light texture images and autofluorescence images. A small target detection network is constructed. The visible light texture image and autofluorescence image are respectively input into the small target detection network. After the dual-modal features are fused by channel attention, separable convolution and spatial offset are performed to form an enhanced feature pyramid. Anchor box dense classification regression is performed in each layer of the pyramid to extract candidate regions of small insect bodies. For the candidate region of tiny insects, the pixel-by-pixel focus curve in the white light focus stack is fitted by Gaussian interpolation to generate a depth topography map, which is stacked with the visible light texture image and the autofluorescence image and fed into the multi-task instance segmentation network to obtain the foreground mask, contour distance map and orientation vector field and construct a joint edge cost map to perform densely adhered insect instance segmentation. Based on insect body segmentation examples, the corresponding region blocks of each independent insect body are extracted from visible light texture images, autofluorescence images and depth topography images. The regions are then input into a pre-constructed three-branch hybrid classification network to extract multimodal features. The extracted multimodal features are then used to identify insect species and determine their liveness or death status through a multi-classification head. Multimodal monitoring data of the target area within a preset period is collected. Based on the multimodal monitoring data, spatiotemporal modeling of population dynamics and prediction of reproductive cycle are performed. A Bayesian inference algorithm is introduced to infer the insect situation at future times based on the population dynamics evolution and the prediction results of reproductive cycle. Finally, insect situation inference information is generated and early warning is pushed out.

[0006] In this solution, the adjustable macro magnification acquisition module is placed close to the surface of fruit and vegetable leaves or sticky insect boards. White light and ultraviolet light are switched in a time-division manner, and focus is adjusted in equal steps along the optical axis within a preset depth range to acquire focus stack images. The focus stack is then fused using a non-subsampled contour wave transform to generate visible light texture images and autofluorescence images. Specifically, this includes: The adjustable macro magnification acquisition module, which integrates a variable focus liquid lens, a ring multi-band light source and an ultraviolet excitation filter, is placed close to the back of fruit and vegetable leaves or the surface of an insect sticky board. The microcontroller selects between white LED and ultraviolet LED in a time-division manner and switches the corresponding bandpass filter in front of the lens simultaneously. Under white light illumination, a stepped drive current signal is applied to the variable focus liquid lens to continuously change the diopter within a preset depth range in fixed steps. Each change in diopter triggers the CMOS image sensor to expose one frame, thereby obtaining a white light focus stack image sequence. Under ultraviolet illumination, the same driving current ladder sequence and exposure time sequence were used to repeatedly scan and obtain the autofluorescence of the insect's epidermal chitin and internal metabolites under ultraviolet excitation, and obtain the ultraviolet fluorescence focus stack image sequence registered with the spatial position of the white light focus stack. Multi-scale fusion processing based on non-downsampled contour wave transform was performed on the white light focus stacked image sequence and the ultraviolet fluorescence focus stacked image sequence respectively. The non-downsampled pyramid filter bank was used to decompose each frame of the image in the stack into a low-pass sub-band and several bandpass direction sub-bands, and the spatial size before downsampling was preserved at each decomposition level. Within the bandpass sub-band, the local phase consistency of pixels is used as a measure of focus sharpness. The corresponding sub-band coefficients of the frame with the largest phase consistency amplitude are selected for fusion. For the low-pass sub-band, the weighted average of the corresponding coefficients of all frames in the stack is taken for fusion. The fused sub-bands are then reconstructed in reverse according to the decomposition level to generate visible light texture images and autofluorescence images.

[0007] In this scheme, the construction of the micro-target detection network involves inputting the visible light texture image and the autofluorescence image into the micro-target detection network, fusing dual-modal features through channel attention, performing separable convolution and spatial offset to form an enhanced feature pyramid, and performing anchor-bound dense classification regression at each layer of the pyramid to extract candidate regions for micro-insect bodies. Specifically, this includes: A small target detection network with a texture-fluorescence dual-stream shallow feature extraction structure is constructed. The texture stream performs continuous convolution operations on visible light texture images to extract morphological details of wing vein orientation, body surface punctures, and color spot boundaries. The fluorescence stream performs convolution operations of the same depth on autofluorescence images to extract bright spot contours and fluorescence intensity distribution features. The shallow feature maps output from the texture stream and fluorescence stream are concatenated along the channel dimension and fed into the channel attention module. The concatenated feature map is compressed into a channel description vector through global average pooling. The weighting coefficients of each channel are generated through a fully connected layer and activation function. The weighting coefficients of each channel are multiplied back into the original feature map for adaptive weighted fusion of dual-modal features. The feature map after channel attention fusion is fed into the dynamic receptive field enhancement module. Multiple sets of depth separable convolutional layers with different dilation rates are arranged in parallel within the module. A learnable spatial offset is introduced during the convolution sampling process so that the receptive field sampling points are adaptively concentrated on the discriminative local structure of the insect body. The outputs of each set are spliced ​​and then compressed and convolved to form a single-layer enhanced feature map. Using the single-layer enhanced feature map as the base layer of the feature pyramid, a multi-scale feature pyramid is constructed layer by layer through top-down upsampling paths and lateral connections. Each layer has different spatial resolution and semantic abstraction to adapt to the scale changes of tiny insects of different body lengths. At each level of the feature pyramid, anchor boxes of various scales and aspect ratios are densely pre-set around each spatial location. Convolutional predictions of classification and regression branches are performed in parallel for each anchor box. The classification branch outputs the confidence score containing the tiny insect, and the regression branch outputs the bounding box offset. The prediction results of all levels of the pyramid are aggregated, and a preliminary filter is performed using a preset relaxed confidence threshold. Non-maximum suppression is then applied to the retained anchor boxes to remove duplicate boxes on the same worm, ultimately generating candidate regions for tiny worms.

[0008] In this scheme, for the candidate region of tiny insect bodies, the pixel-by-pixel focus curve in the white light focus stack is fitted with Gaussian interpolation to generate a depth topography map. This map is then stacked with the visible light texture image and the autofluorescence image and fed into a multi-task instance segmentation network to obtain the foreground mask, contour distance map, and orientation vector field, and a joint edge cost map is constructed for the segmentation of densely clustered insect body instances. Specifically, this includes: For each candidate region of a tiny insect, the improved Laplacian sum of each frame is calculated pixel by pixel in the white light focus stack image sequence with a preset window to generate a focus curve as the focus metric value. The focus curve is then fitted with a Gaussian function, and the depth index corresponding to the peak of the fitted curve is used as the height value of the insect surface corresponding to that pixel. Arrange the height values ​​of all pixels according to their spatial positions to form a depth topography map that is pixel-level aligned with the candidate region. The depth topography map shows height gradient changes at the junction depressions of the adhered worm bodies and at the interlayer transitions between stacked layers. The visible light texture image block, autofluorescence image block and the depth topography map corresponding to the same candidate region are stacked in the channel dimension to form a multi-channel input tensor, which is then imported into a multi-task instance segmentation network to obtain the foreground mask, contour distance map and orientation vector field. The multi-task instance segmentation network adopts a shared encoder and multi-branch decoder structure. The shared encoder extracts abstract features that fuse texture, fluorescence and depth morphology information from the multi-channel input. The foreground mask branch, contour distance map branch and orientation vector field branch are arranged in parallel. The three branches share the same encoder features and are jointly optimized during training. The contour distance map, orientation vector field and depth topography map are fused pixel by pixel to construct a joint edge cost map. The contour distance map provides the individual boundary prior on the two-dimensional plane of the worm, the orientation vector field provides the pixel orientation change signal at the adhesion boundary, and the depth topography map provides the interlayer stacking and surface depression information in the three-dimensional space of the worm. The effective processing area is constrained by the foreground mask. Local minimum points of the contour distance map within the foreground range are detected on the joint edge cost map as seed points for each potential individual. A marker watershed algorithm is introduced to simulate the water immersion process from various seed points. When the immersion fronts of different seed points meet at the ridge line of the cost map, a segmentation boundary is formed. Finally, the densely adhered and stacked worm bodies are cut into independent single instances.

[0009] In this scheme, the insect body segmentation instance extracts the corresponding region blocks of each independent insect body from visible light texture image, autofluorescence image, and depth topography image. These regions are then input into a pre-constructed three-branch hybrid classification network to extract multimodal features. The extracted multimodal features are then used to identify the insect species and determine its liveness / death status through a multi-classification head. Specifically, this includes: Based on the outline boundary of each insect segmentation instance, visible light texture region blocks, autofluorescence region blocks, and depth morphology region blocks with corresponding positions and sizes are cropped from the visible light texture image, autofluorescence image, and depth morphology image, respectively. Each region block contains only the pixel information of the independent insect itself. The visible light texture region block, autofluorescence region block, and depth morphology region block are respectively fed into the corresponding branches of the pre-constructed three-branch hybrid classification network. The three-branch hybrid classification network consists of a texture feature extraction branch, a fluorescence feature extraction branch, and a depth morphology feature extraction branch. The three branches are structurally independent of each other. The texture feature extraction branch extracts microscopic morphological features, including the direction of wing vein bifurcation, the arrangement of body surface punctures, and the outline of color spots, from the visible light texture region block by stacking multiple layers of small-sized convolution kernels. It also gradually compresses the spatial dimension through pooling operations to integrate the global texture representation of various parts of the insect body. The fluorescence feature extraction branch uses the same convolution depth as the texture branch but with independent weight parameters. It extracts the intensity distribution pattern of autofluorescence signal, the spatial concentration of fluorescent bright spots, and the decay gradient features of fluorescence from the center of the insect body to the edge from the autofluorescence region block layer by layer, forming a feature expression that reflects the physiological state of the insect body. The deep morphology feature extraction branch extracts three-dimensional geometric features from the deep morphology region block, including the curvature distribution of the worm's back, the height difference between segments, and the relative inclination angle between the worm and the attachment surface, to describe the worm's morphology features. Finally, the feature vectors extracted by the three branches are spliced ​​and fused in the fully connected layer to form a multimodal feature vector. On top of the fully connected layer, a fine-grained insect species classification head and a binary classification head for live / dead states are set up in parallel. The insect species classification head outputs the probability of the target insect segmentation instance belonging to each insect species, and the insect species with the highest probability is taken as the identification result. The live / dead state classification head outputs the binary classification probability value of live insects and dead insects.

[0010] In this scheme, the collection of multimodal monitoring data of the target area within a preset period, the performance of population dynamics spatiotemporal modeling and reproductive cycle prediction based on the multimodal monitoring data, the introduction of a Bayesian inference algorithm to infer insect infestation at future times based on the population dynamics evolution and reproductive cycle prediction results, and the final generation of insect infestation inference information for early warning push, specifically includes: All single-insect data collected multiple times within a preset monitoring period are aggregated according to the collection timestamp and spatial location label to form a multimodal monitoring dataset. Each record contains the insect species label, live / dead status label, spatial coordinates, and the cumulative environmental temperature value corresponding to the collection of a single insect in a single collection. The multimodal monitoring dataset is aggregated according to a preset statistical granularity, and the number of live insects, the number of dead insects, and the total insect population density of each insect species are counted in each time period. The average temperature and effective accumulated temperature are obtained from the environmental sensors deployed in the monitoring area in the corresponding time period to form the insect population density time series and the accumulated temperature time series. A hybrid prediction framework combining a fusion mechanism model and a data-driven model is constructed. The first pathway calculates the accumulation rate of temperature based on the developmental starting temperature and effective accumulated temperature constant of each target insect species, combined with daily accumulated temperature data. The rate and model output population trend prediction values ​​based on biological mechanisms are then used. The second pathway inputs the insect population density time series into a long short-term memory network with an attention mechanism. Through a gating mechanism, it selectively memorizes and forgets the long-term trends, periodic fluctuations, and short-term mutations in the historical insect population density series. The attention mechanism assigns weights to different historical moments and outputs a data-driven predicted value of insect population density. A Bayesian inference algorithm is introduced to integrate the population trend prediction and insect population density prediction under a unified probability framework. The posterior probability distribution of the fused insect population density is calculated, and the posterior mean is output as the fused prediction value and confidence interval. A three-level early warning judgment logic is constructed based on the posterior probability distribution. When the posterior mean of the fused predicted density is lower than the proliferation threshold, it is judged as low density level; when it exceeds the proliferation threshold and the lower limit of the confidence interval breaks through the rising judgment line, it is judged as proliferation period; when it exceeds the outbreak threshold and the posterior probability is greater than the confidence level requirement, it is judged as outbreak period. Insect situation inference information is generated and early warning is pushed.

[0011] A second aspect of the present invention provides an intelligent monitoring system for microscopic insects in fruits and vegetables based on microscopic magnification imaging. The system includes: a memory, a processor, and a communication interface. The memory contains a program for an intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging. When the program for the intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging is executed by the processor, it implements the steps of the intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging as described in any of the above claims. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments or examples of the present invention, the drawings used in the embodiments or examples will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained according to these drawings without creative effort.

[0013] Figure 1 The first method flowchart of an intelligent monitoring method for tiny insects in fruits and vegetables based on microscopic magnification imaging is provided in an embodiment of the present invention; Figure 2 The flowchart of the second method of an intelligent monitoring method for tiny insects in fruits and vegetables based on microscopic magnification imaging is provided in an embodiment of the present invention. Figure 3 A block diagram of an intelligent monitoring system for tiny insects in fruits and vegetables based on microscopic magnification imaging, provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0014] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0015] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0016] Figure 1 The first method flowchart of an intelligent monitoring method for tiny insects in fruits and vegetables based on microscopic magnification imaging is provided in an embodiment of the present invention; like Figure 1 As shown, the present invention provides a first method flowchart for an intelligent monitoring method for tiny insects in fruits and vegetables based on microscopic magnification imaging, comprising: S102, the adjustable macro magnification acquisition module is placed close to the surface of fruit and vegetable leaves or sticky insect board, and white light and ultraviolet light are irradiated by time-division switching, and the focus is adjusted in equal steps along the optical axis within a preset depth range to acquire focus stack. The focus stack is fused by non-subsampled contour wave transformation to generate visible light texture image and autofluorescence image. S104, construct a small target detection network. Input the visible light texture image and autofluorescence image into the small target detection network respectively. After fusing the dual-modal features through channel attention, perform separable convolution and spatial offset to form an enhanced feature pyramid. Perform anchor box dense classification regression in each layer of the pyramid to extract candidate regions for small insect bodies. S106, For the candidate region of tiny insects, the pixel-by-pixel focus curve in the white light focus stack is fitted by Gaussian interpolation to generate a depth topography map, which is stacked with the visible light texture image and the autofluorescence image and fed into the multi-task instance segmentation network to obtain the foreground mask, contour distance map and orientation vector field and construct a joint edge cost map to perform densely adhered insect instance segmentation. S108: Based on insect body segmentation instances, extract the corresponding region blocks of each independent insect body from visible light texture images, autofluorescence images and depth topography images, input them into a pre-constructed three-branch hybrid classification network to extract multimodal features, and use the extracted multimodal features to identify insect species and determine their liveness or death status through a multi-classification head. S110: Collect multimodal monitoring data of the target area within a preset period, perform spatiotemporal modeling of population dynamics and prediction of reproductive cycle based on the multimodal monitoring data, introduce Bayesian inference algorithm to infer the insect situation at future times based on the population dynamic evolution trend and the prediction results of reproductive cycle, and finally generate insect situation inference information for early warning push.

[0017] Furthermore, in a preferred embodiment of the present invention, the adjustable macro magnification acquisition module is placed close to the surface of fruit and vegetable leaves or sticky insect boards, and white light and ultraviolet light are irradiated by time-division switching, and focus is adjusted in equal steps along the optical axis within a preset depth range to perform focus stack acquisition. The focus stack is then fused using non-downsampling contour wave transform to generate a visible light texture image and an autofluorescence image. Specifically, this includes: The adjustable macro magnification acquisition module, which integrates a variable focus liquid lens, a ring multi-band light source and an ultraviolet excitation filter, is placed close to the back of fruit and vegetable leaves or the surface of an insect sticky board. The microcontroller selects between white LED and ultraviolet LED in a time-division manner and switches the corresponding bandpass filter in front of the lens simultaneously. Under white light illumination, a stepped drive current signal is applied to the variable focus liquid lens to continuously change the diopter within a preset depth range in fixed steps. Each change in diopter triggers the CMOS image sensor to expose one frame, thereby obtaining a white light focus stack image sequence. Under ultraviolet illumination, the same driving current ladder sequence and exposure time sequence were used to repeatedly scan and obtain the autofluorescence of the insect's epidermal chitin and internal metabolites under ultraviolet excitation, and obtain the ultraviolet fluorescence focus stack image sequence registered with the spatial position of the white light focus stack. Multi-scale fusion processing based on non-downsampled contour wave transform was performed on the white light focus stacked image sequence and the ultraviolet fluorescence focus stacked image sequence respectively. The non-downsampled pyramid filter bank was used to decompose each frame of the image in the stack into a low-pass sub-band and several bandpass direction sub-bands, and the spatial size before downsampling was preserved at each decomposition level. Within the bandpass sub-band, the local phase consistency of pixels is used as a measure of focus sharpness. The corresponding sub-band coefficients of the frame with the largest phase consistency amplitude are selected for fusion. For the low-pass sub-band, the weighted average of the corresponding coefficients of all frames in the stack is taken for fusion. The fused sub-bands are then reconstructed in reverse according to the decomposition level to generate visible light texture images and autofluorescence images.

[0018] It should be noted that in actual monitoring environments in fruit and vegetable fields, tiny insects are typically less than a few millimeters long and often hide in the velvet layer on the underside of leaves or in the recesses of the adhesive layer on sticky insect traps. Ordinary fixed-focus macro lenses are limited by their inherent depth of field, and a single exposure can only obtain a clear image within a very shallow focal plane. Most of the insect's structure is blurred due to defocusing, and key identifying features such as wing veins and antennae are severely lost. To solve this problem, this solution uses a variable-focus liquid lens as the focusing actuator. Its refractive power can be continuously changed in milliseconds via the driving current, enabling rapid focus scanning along the optical axis without mechanical displacement. Combined with the stepped drive signal generated by the microcontroller, it acquires images layer by layer within a preset depth range with a fixed step size, obtaining a complete focus stack covering the leaf surface velvet layer to the shallow mesophyll layer or from the adhesive layer on the sticky insect trap to the top of the insect. This ensures that insect structures of any height are precisely in the optimal focus plane in a single frame of the stack.

[0019] In terms of illumination strategy, white light and ultraviolet light sources were switched time-divisionally while the bandpass filter in front of the lens was switched synchronously. Two image sequences—one with white light focus stack and one with ultraviolet fluorescence focus stack—were acquired sequentially at the same spatial location. Under white light illumination, the surface texture and pigmentation of the insect were clearly visible. Under ultraviolet excitation, the chitin in the insect's epidermis and flavin metabolites produced autofluorescence with peak wavelengths of approximately 450 to 520 nanometers. The spectral distribution of chlorophyll fluorescence in leaves differed significantly from this, resulting in a significant brightness contrast between the insect area and the plant background in the autofluorescence image. Both stacks shared the same driving current step sequence and exposure timing during acquisition, ensuring that the subsequently generated fully focused texture image and fully focused fluorescence image naturally maintained pixel-level spatial registration, laying the foundation for dual-modal feature fusion. In the image fusion stage, a multi-scale fusion algorithm based on non-subsampled contourlet transform was used to process each focus stack separately. This transformation decomposes each frame of the image into a low-pass sub-band and several bandpass directional sub-bands through a non-downsampled pyramidal filter bank. No downsampling is performed at any stage of decomposition, ensuring that each sub-band maintains the same spatial size as the original image and possesses translation invariance, thus avoiding artifacts at depth-of-field abrupt changes in the fusion result. Within the bandpass directional sub-bands, local phase consistency is used as a measure of focus sharpness. This measure is insensitive to changes in local image contrast and brightness, accurately determining the focus state even in areas where contrast is reduced due to hair occlusion or gel wetting at the edges of the insect body, demonstrating higher robustness compared to traditional gradient operators. The low-pass sub-band is fused using the weighted average of the corresponding coefficients from all frames in the stack to maintain the continuity of the overall brightness distribution. After fusion, the image is reconstructed in reverse according to the decomposition levels, ultimately generating a globally clear visible light texture image and autofluorescence image of the insect structure across the entire scanning depth.

[0020] Furthermore, in a preferred embodiment of the present invention, the construction of the micro-target detection network involves inputting the visible light texture image and the autofluorescence image into the micro-target detection network, fusing dual-modal features through channel attention, performing separable convolution and spatial offset to form an enhanced feature pyramid, and performing anchor-bound dense classification regression at each layer of the pyramid to extract candidate regions for micro-insect bodies. Specifically, this includes: A small target detection network with a texture-fluorescence dual-stream shallow feature extraction structure is constructed. The texture stream performs continuous convolution operations on visible light texture images to extract morphological details of wing vein orientation, body surface punctures, and color spot boundaries. The fluorescence stream performs convolution operations of the same depth on autofluorescence images to extract bright spot contours and fluorescence intensity distribution features. The shallow feature maps output from the texture stream and fluorescence stream are concatenated along the channel dimension and fed into the channel attention module. The concatenated feature map is compressed into a channel description vector through global average pooling. The weighting coefficients of each channel are generated through a fully connected layer and activation function. The weighting coefficients of each channel are multiplied back into the original feature map for adaptive weighted fusion of dual-modal features. The feature map after channel attention fusion is fed into the dynamic receptive field enhancement module. Multiple sets of depth separable convolutional layers with different dilation rates are arranged in parallel within the module. A learnable spatial offset is introduced during the convolution sampling process so that the receptive field sampling points are adaptively concentrated on the discriminative local structure of the insect body. The outputs of each set are spliced ​​and then compressed and convolved to form a single-layer enhanced feature map. Using the single-layer enhanced feature map as the base layer of the feature pyramid, a multi-scale feature pyramid is constructed layer by layer through top-down upsampling paths and lateral connections. Each layer has different spatial resolution and semantic abstraction to adapt to the scale changes of tiny insects of different body lengths. At each level of the feature pyramid, anchor boxes of various scales and aspect ratios are densely pre-set around each spatial location. Convolutional predictions of classification and regression branches are performed in parallel for each anchor box. The classification branch outputs the confidence score containing the tiny insect, and the regression branch outputs the bounding box offset. The prediction results of all levels of the pyramid are aggregated, and a preliminary filter is performed using a preset relaxed confidence threshold. Non-maximum suppression is then applied to the retained anchor boxes to remove duplicate boxes on the same worm, ultimately generating candidate regions for tiny worms.

[0021] It should be noted that in the task of detecting tiny insects, when the target body length is less than millimeters, it occupies only a very small number of pixels in the image. During the process of extracting high-level semantic features through multiple downsampling in conventional target detection networks, the spatial information of tiny insects is easily compressed and obscured, causing the network to misclassify them as background noise and miss detection. To solve this problem, this solution constructs a texture-fluorescence dual-stream shallow feature extraction structure. The texture stream is responsible for capturing high-frequency morphological details such as the direction of wing vein bifurcation, the density of punctures on the body surface, and the contours of color spots from the full-focus visible light texture image. The fluorescence stream extracts the contours of bright spots on the insect body and the spatial distribution features of fluorescence intensity from the full-focus autofluorescence image. The two streams complete feature extraction and retain the feature map in the shallow stage before the spatial resolution is significantly reduced, avoiding the irreversible loss of small target signals caused by deep downsampling.

[0022] Specifically, after concatenating the dual-stream features along the channel dimension, they are fed into the channel attention module for adaptive fusion. This module compresses the concatenated feature map into channel description vectors through global average pooling. These vectors are then processed by two fully connected layers and an activation function to learn the contribution weights of each feature channel in distinguishing the insect from the background. Finally, each channel is multiplied back into the original feature map, enabling the network to dynamically enhance effective modal channels and suppress ineffective channels caused by local reflections or fluorescence quenching, thus achieving intelligent integration of dual-modal information. The feature map after channel attention fusion enters the dynamic receptive field enhancement module. This module contains multiple sets of depthwise separable convolutional layers with different dilation rates arranged in parallel. A learnable spatial offset is introduced during the convolution sampling process, allowing the sampling point positions of the receptive field to adaptively shift according to the slender shape and curved posture of the insect. This concentrates the sampling points on discriminative local structures such as antennae, legs, and the end of the abdomen, rather than fixed rectangular neighborhoods, thereby effectively suppressing the introduction of background noise on small targets. The outputs of each group are concatenated and compressed into a single-layer enhanced feature map. This map is then used as a base to construct a multi-scale feature pyramid through top-down upsampling and lateral connections, enabling the network to respond simultaneously to adults with smaller body lengths. Anchor-box dense classification regression is performed at each layer of the pyramid, and candidate regions are preserved with a lenient threshold to ensure that adhered individuals are not lost. By synergistically employing dual-stream shallow feature preservation, adaptive channel attention fusion, and deformable sampling of dynamic receptive fields, the challenge of signal loss in deep feature maps for small insects is effectively overcome. Even when the insect body is partially obscured by leaf hairs or impurities, it can still respond based on local discriminative components.

[0023] Furthermore, in a preferred embodiment of the present invention, for the candidate region of the tiny insect, the depth topography map is generated by Gaussian interpolation fitting of the pixel-by-pixel focus curve in the white light focus stack, and stacked with the visible light texture image and the autofluorescence image and fed into a multi-task instance segmentation network to obtain the foreground mask, contour distance map and orientation vector field, and construct a joint edge cost map for densely adhered insect instance segmentation, specifically including: For each candidate region of a tiny insect, the improved Laplacian sum of each frame is calculated pixel by pixel in the white light focus stack image sequence with a preset window to generate a focus curve as the focus metric value. The focus curve is then fitted with a Gaussian function, and the depth index corresponding to the peak of the fitted curve is used as the height value of the insect surface corresponding to that pixel. Arrange the height values ​​of all pixels according to their spatial positions to form a depth topography map that is pixel-level aligned with the candidate region. The depth topography map shows height gradient changes at the junction depressions of the adhered worm bodies and at the interlayer transitions between stacked layers. The visible light texture image block, autofluorescence image block and the depth topography map corresponding to the same candidate region are stacked in the channel dimension to form a multi-channel input tensor, which is then imported into a multi-task instance segmentation network to obtain the foreground mask, contour distance map and orientation vector field. The multi-task instance segmentation network adopts a shared encoder and multi-branch decoder structure. The shared encoder extracts abstract features that fuse texture, fluorescence and depth morphology information from the multi-channel input. The foreground mask branch, contour distance map branch and orientation vector field branch are arranged in parallel. The three branches share the same encoder features and are jointly optimized during training. The contour distance map, orientation vector field and depth topography map are fused pixel by pixel to construct a joint edge cost map. The contour distance map provides the individual boundary prior on the two-dimensional plane of the worm, the orientation vector field provides the pixel orientation change signal at the adhesion boundary, and the depth topography map provides the interlayer stacking and surface depression information in the three-dimensional space of the worm. The effective processing area is constrained by the foreground mask. Local minimum points of the contour distance map within the foreground range are detected on the joint edge cost map as seed points for each potential individual. A marker watershed algorithm is introduced to simulate the water immersion process from various seed points. When the immersion fronts of different seed points meet at the ridge line of the cost map, a segmentation boundary is formed. Finally, the densely adhered and stacked worm bodies are cut into independent single instances.

[0024] It should be noted that in high-density monitoring scenarios using sticky insect boards, tiny insects such as thrips and whiteflies often cluster together, their bodies pressed tightly together and overlapping. Conventional target detection boxes can only select the entire cluster and cannot penetrate to the individual insects, resulting in multiple insects being merged and counted, severely affecting the accuracy of density statistics. When relying solely on 2D textures for instance segmentation, the texture difference between two parallel, closely packed insects at their boundary is very weak, making it difficult for the segmentation network to perceive the boundary. Furthermore, in cases of stacking, the 2D image itself does not carry depth information, making it impossible to determine the overlap relationship based solely on appearance. This solution utilizes the 3D spatial cues contained in the acquired white light focus stack to address this bottleneck.

[0025] Specifically, for each candidate region, each frame in the white light focus stack is traversed pixel by pixel. The improved Laplacian sum is calculated using a preset small window as the focus measure. This measure is stable for high-frequency texture responses and exhibits good single-peak characteristics. Then, Gaussian interpolation is performed on the focus curve to locate the depth index corresponding to the peak value. This index is used as the relative height of the insect's surface at that pixel. The height values ​​of all pixels are arranged according to their spatial positions to generate a depth topography map that is pixel-level aligned with the candidate region. Significant height gradient changes are observed at adhesion junctions and interlayer transitions in this map. Subsequently, the visible light texture blocks, autofluorescent blocks, and depth topography maps corresponding to the candidate regions are stacked in the channel dimension to form a multi-channel input tensor, which is then fed into the multi-task instance segmentation network. This network synchronously extracts joint features of texture, fluorescence, and depth through a shared encoder. At the decoding end, a foreground mask branch, a contour distance map branch, and an orientation vector field branch are deployed in parallel. The contour distance map branch predicts the normalized distance from each pixel to the nearest boundary within the foreground region, naturally forming a valley-shaped response at adhesion points. The orientation vector field branch predicts the unit vector pointing from each pixel towards the worm's head, with abrupt discontinuities in the lateral directions at the boundaries between adjacent worms. These three branches share encoder features and are jointly optimized, enabling the network to learn complementary representational relationships between appearance texture, fluorescence distribution, and 3D morphology. During the segmentation stage, the contour distance map, orientation vector field, and depth morphology map are weighted and fused pixel-by-pixel to construct a joint edge cost map. The distance map provides a priori information on 2D individual boundaries, the orientation field provides directional abrupt change signals, and the depth map provides 3D geometric evidence for interlayer stacking and depressions. The synergy of these three elements significantly increases the boundary cost of adhesion regions compared to flat areas within the worm. Finally, using the local minima of the joint cost map and distance map as seeds, the labeling watershed algorithm is executed. The immersion front meets at the ridge line of the cost map to form a segmentation boundary, accurately cutting the adhered and stacked worms into independent single instances. By calculating the depth topography map from the existing focal stack as a third mode, and feeding it together with texture and fluorescence information into the multi-task segmentation network, and introducing an orientation vector field as a geometric prior to guide the watershed segmentation, the 3D topography information of the worms can be obtained without additional depth sensors, improving the segmentation accuracy of stacked and parallel adhered worms, and solving the problem of merging multiple worms when counting dense small worms.

[0026] Furthermore, in a preferred embodiment of the present invention, the step of extracting corresponding region blocks of each independent insect from visible light texture images, autofluorescence images, and depth topography images based on insect body segmentation examples, inputting them into a pre-constructed three-branch hybrid classification network to extract multimodal features, and using the extracted multimodal features to identify insect species and determine their liveness / death status through a multi-classification head, specifically includes: Based on the outline boundary of each insect segmentation instance, visible light texture region blocks, autofluorescence region blocks, and depth morphology region blocks with corresponding positions and sizes are cropped from the visible light texture image, autofluorescence image, and depth morphology image, respectively. Each region block contains only the pixel information of the independent insect itself. The visible light texture region block, autofluorescence region block, and depth morphology region block are respectively fed into the corresponding branches of the pre-constructed three-branch hybrid classification network. The three-branch hybrid classification network consists of a texture feature extraction branch, a fluorescence feature extraction branch, and a depth morphology feature extraction branch. The three branches are structurally independent of each other. The texture feature extraction branch extracts microscopic morphological features, including the direction of wing vein bifurcation, the arrangement of body surface punctures, and the outline of color spots, from the visible light texture region block by stacking multiple layers of small-sized convolution kernels. It also gradually compresses the spatial dimension through pooling operations to integrate the global texture representation of various parts of the insect body. The fluorescence feature extraction branch uses the same convolution depth as the texture branch but with independent weight parameters. It extracts the intensity distribution pattern of autofluorescence signal, the spatial concentration of fluorescent bright spots, and the decay gradient features of fluorescence from the center of the insect body to the edge from the autofluorescence region block layer by layer, forming a feature expression that reflects the physiological state of the insect body. The deep morphology feature extraction branch extracts three-dimensional geometric features from the deep morphology region block, including the curvature distribution of the worm's back, the height difference between segments, and the relative inclination angle between the worm and the attachment surface, to describe the worm's morphology features. Finally, the feature vectors extracted by the three branches are spliced ​​and fused in the fully connected layer to form a multimodal feature vector. On top of the fully connected layer, a fine-grained insect species classification head and a binary classification head for live / dead states are set up in parallel. The insect species classification head outputs the probability of the target insect segmentation instance belonging to each insect species, and the insect species with the highest probability is taken as the identification result. The live / dead state classification head outputs the binary classification probability value of live insects and dead insects.

[0027] It should be noted that traditional methods face a dilemma when identifying the species of independent insects segmented under a microscope. Thrips nymphs and whitefly nymphs have extremely similar body shapes and colors in two-dimensional textures, and the color patterns of two-spotted spider mites and aphid nymphs are easily confused under low magnification. It is difficult to achieve stable fine-grained classification based solely on visible light texture information. A further challenge lies in distinguishing between live and dead states—fresh dead insects, old and shriveled insects, and live insects that are not yet dead can coexist on sticky insect boards. They appear highly similar in static texture images, but the core indicator required for control decisions is the density of live insects. Including dead insects in the count would severely overestimate the risk of insect infestation. To overcome these challenges, this scheme, based on the segmentation results of independent insect instances, uses the contour boundaries of the instance to simultaneously crop three types of regions with corresponding positions and sizes from the visible light texture image, autofluorescence image, and depth topography image. Each region contains only the pixel information of the insect itself, excluding interference from adjacent individuals and the background, and uses this as input data for fine classification.

[0028] The three types of region blocks were then fed into the corresponding branches of the three-branch hybrid classification network. The texture feature extraction branch captured microscopic morphological features such as the bifurcation direction of wing veins, the density and arrangement of punctures on the body surface, and the contours of color spots by stacking multiple layers of small-sized convolutional kernels. It also gradually compressed the spatial dimension through pooling operations, integrating local details into a global texture representation with insect species recognition. The fluorescence feature extraction branch used the same convolution depth as the texture branch but with independent weight parameters. It extracted the intensity distribution pattern of fluorescence signals, the spatial concentration of bright spots, and the decay gradient characteristics of fluorescence from the center to the edge of the insect body from the autofluorescent region blocks layer by layer. The concentration of flavin metabolites in live insects is high and evenly distributed, and the fluorescence image is bright overall with a smooth decrease from the center to the outside. After death, the fluorescence intensity gradually decreases over time and the distribution tends to be mottled and uneven. The fluorescence branch formed a feature expression reflecting the physiological state of the insect body by abstracting these subtle differences layer by layer. The depth morphology feature extraction branch consists of shallower convolutional layers, specifically extracting three-dimensional geometric features from depth morphology regions, such as the curvature distribution of the insect's back, the height difference between body segments, and the relative inclination angle between the insect and the attachment surface. Flat spider mites and aphid nymphs with higher arches exhibit stable differences in depth morphology; these three-dimensional geometric clues cannot be obtained from pure two-dimensional textures. After the three branches are spliced ​​and fused in a fully connected layer, a fine-grained insect species classification head and a binary classification head for live / dead status are set up in parallel. The insect species classification head outputs the probability of the target insect belonging to each insect species, and the maximum value is taken as the identification result. The live / dead status classification head outputs the binary classification probability of live and dead insects. Through the classification architecture fused after independent encoding of the three branches, two-dimensional texture, autofluorescence physiological information, and three-dimensional morphological geometric clues are organically integrated at the feature level, thereby improving classification accuracy. Simultaneously, the ability to distinguish between live and fresh / dead insects allows the system to directly count the density of live insects, avoiding interference from dead insects in insect monitoring.

[0029] Figure 2 The flowchart of the second method of an intelligent monitoring method for tiny insects in fruits and vegetables based on microscopic magnification imaging is provided in an embodiment of the present invention. like Figure 2 As shown, the present invention provides a second method flowchart for an intelligent monitoring method for tiny insects in fruits and vegetables based on microscopic magnification imaging, comprising: S202, collect all single insect data collected multiple times within the preset monitoring period according to the collection timestamp and spatial location label to form a multimodal monitoring dataset. Each record contains the insect species label, live / dead status label, spatial coordinates, and the cumulative environmental temperature value corresponding to the collection of a single insect in a single collection. S204, the multimodal monitoring dataset is aggregated according to a preset statistical granularity, and the number of live insects, the number of dead insects, and the total insect population density of each insect species are counted in each time period. The average temperature and effective accumulated temperature are obtained from the environmental sensors deployed in the monitoring area in the corresponding time period to form the insect population density time series and the accumulated temperature time series. S206, construct a hybrid prediction framework that integrates the fusion mechanism model and the data-driven model. The first pathway calculates the accumulation rate of accumulated temperature based on the developmental starting temperature and effective accumulated temperature constant of each target insect species, combined with the daily accumulated temperature data. The rate and model output population trend prediction values ​​based on biological mechanisms are used. S208, the second pathway inputs the insect population density time series into a long short-term memory network with an attention mechanism. Through the gating mechanism, it selectively remembers and forgets the long-term trends, periodic fluctuations and short-term mutations in the historical insect population density series. The attention mechanism assigns weights to different historical moments and outputs a data-driven predicted value of insect population density. S210 introduces a Bayesian inference algorithm, which integrates the population trend prediction value and the insect population density prediction value under a unified probability framework, calculates the posterior probability distribution of the fused insect population density, and outputs the posterior mean as the fused prediction value and confidence interval simultaneously. Based on the posterior probability distribution, a three-level early warning judgment logic is constructed. S212: When the posterior mean of the fused predicted density is lower than the proliferation threshold, it is judged as low density level; when it exceeds the proliferation threshold and the lower limit of the confidence interval breaks through the rising judgment line, it is judged as proliferation period; when it exceeds the outbreak threshold and the posterior probability is greater than the confidence requirement, it is judged as outbreak period. Insect situation inference information is generated and early warning is pushed.

[0030] It should be noted that the fluctuations in field pest populations are influenced by a complex interplay of factors, including temperature-driven developmental rates, crop growth stages, the number of natural enemies, and human control measures. A single predictive model cannot fully capture the dynamic patterns of this complex system. While purely biological mechanism models, such as the accumulated temperature developmental rate method, can explain the causal relationship between temperature and insect development progress well, they lack the ability to perceive short-term population fluctuations caused by non-temperature factors such as rainfall, pesticide application, or the migration of natural enemies. On the other hand, while purely data-driven models can learn various implicit patterns from historical sequences, they may produce predictive results that violate biological common sense when the learning sample is limited, such as predicting a continued increase in pest populations during the cold winter months. The advantages and limitations of these two types of models are naturally complementary. Organically integrating them within a probabilistic framework is a reasonable path to improve the reliability of pest forecasting.

[0031] Specifically, this scheme is based on the identification and discrimination results of all single insects accumulated within the preset monitoring period. It collects the insect species tags, live / dead status tags, spatial coordinates and corresponding environmental accumulated temperature values ​​collected each time into a multimodal monitoring dataset according to timestamps and spatial location tags. Then, it aggregates the data according to the preset statistical granularity to obtain the time series of live insects, dead insects and total insect population density of each insect species. It is then aligned with the daily average temperature and effective accumulated temperature values ​​obtained synchronously from environmental sensors to form the basic data for modeling. In the prediction phase, a dual-pathway hybrid prediction framework is constructed. The first pathway continuously calculates the current accumulation rate of accumulated temperature based on the known developmental starting temperature and the effective accumulated temperature constant required to complete one generation for each target insect species, combined with actual accumulated temperature data from the field. It then uses the rate and model to estimate the potential occurrence window of the next generation of nymphs or adults, and outputs a population trend prediction value based on insect physiological laws. The second pathway inputs the insect population density time series into a long short-term memory network with an attention mechanism. This network uses gating mechanisms such as forgetting gates, input gates, and output gates to selectively remember and forget long-term trends, seasonal cyclical fluctuations, and sudden drops after short-term pesticide application in the historical sequence. The attention mechanism automatically identifies historical periods similar to the current situation in the time dimension and assigns higher weights to these periods, outputting a data-driven insect population density prediction value.

[0032] Furthermore, a Bayesian inference algorithm is introduced, treating the mechanism-based predictions from the accumulated temperature model and the data-driven predictions from the Long Short-Term Memory network as observational evidence from two independent information sources. Under a unified probabilistic framework, the fused posterior probability distribution is calculated, and the posterior mean is simultaneously output as the fused prediction value, along with the upper and lower bounds of the distribution as the prediction confidence interval. Based on this posterior distribution, a three-level early warning judgment logic is constructed: when the posterior mean of the fused predicted density is lower than the proliferation threshold, it is judged as a low-density level, maintaining only the regular monitoring frequency; when the posterior mean exceeds the proliferation threshold and the lower bound of the confidence interval breaks through the rising judgment line, it is judged as a proliferation period, triggering a primary early warning and prompting increased monitoring frequency; when the posterior mean exceeds the outbreak threshold and the posterior probability is greater than the preset confidence requirement, it is judged as an outbreak period, triggering a high-level early warning. Finally, the pest situation inference information, including predicted insect species, density range, early warning level, and control window suggestions, is pushed to the user terminal via a wireless communication module. This effectively reduces false alarms caused by random fluctuations and provides quantitative evidence with both biological interpretability and statistical reliability for control decisions.

[0033] Figure 3 An embodiment of the present invention provides an intelligent monitoring system 3 for fruit and vegetable micro-insects based on microscopic magnification imaging. The system includes: a memory 301, a processor 302, and a communication interface 303. The memory 301 contains a program for an intelligent monitoring method for fruit and vegetable micro-insects based on microscopic magnification imaging. When the program for the intelligent monitoring method for fruit and vegetable micro-insects based on microscopic magnification imaging is executed by the processor 302, it implements the steps of the intelligent monitoring method for fruit and vegetable micro-insects based on microscopic magnification imaging as described above.

[0034] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for intelligent monitoring of tiny insects in fruits and vegetables based on microscopic magnification imaging, characterized in that, include: The adjustable macro magnification acquisition module is placed close to the surface of fruit and vegetable leaves or sticky insect boards. White light and ultraviolet light are irradiated by time-division switching, and the focus is adjusted in equal steps within a preset depth range along the optical axis to acquire focus stacks. The focus stacks are fused by non-downsampling contour wave transformation to generate visible light texture images and autofluorescence images. A small target detection network is constructed. The visible light texture image and autofluorescence image are respectively input into the small target detection network. After the dual-modal features are fused by channel attention, separable convolution and spatial offset are performed to form an enhanced feature pyramid. Anchor box dense classification regression is performed in each layer of the pyramid to extract candidate regions of small insect bodies. For the candidate region of tiny insects, the pixel-by-pixel focus curve in the white light focus stack is fitted by Gaussian interpolation to generate a depth topography map, which is stacked with the visible light texture image and the autofluorescence image and fed into the multi-task instance segmentation network to obtain the foreground mask, contour distance map and orientation vector field and construct a joint edge cost map to perform densely adhered insect instance segmentation. Based on insect body segmentation examples, the corresponding region blocks of each independent insect body are extracted from visible light texture images, autofluorescence images and depth topography images. The regions are then input into a pre-constructed three-branch hybrid classification network to extract multimodal features. The extracted multimodal features are then used to identify insect species and determine their liveness or death status through a multi-classification head. Multimodal monitoring data of the target area within a preset period is collected. Based on the multimodal monitoring data, spatiotemporal modeling of population dynamics and prediction of reproductive cycle are performed. A Bayesian inference algorithm is introduced to infer the insect situation at future times based on the population dynamics evolution and the prediction results of reproductive cycle. Finally, insect situation inference information is generated and early warning is pushed out.

2. The intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging according to claim 1, characterized in that, The adjustable macro magnification acquisition module is placed close to the surface of fruit and vegetable leaves or sticky insect boards. White light and ultraviolet light are switched in a time-division manner, and focus is adjusted in equal steps along the optical axis within a preset depth range to acquire focus stack images. The focus stack is then fused using a non-downsampled contour wave transform to generate visible light texture images and autofluorescence images. Specifically, this includes: The adjustable macro magnification acquisition module, which integrates a variable focus liquid lens, a ring multi-band light source and an ultraviolet excitation filter, is placed close to the back of fruit and vegetable leaves or the surface of an insect sticky board. The microcontroller selects between white LED and ultraviolet LED in a time-division manner and switches the corresponding bandpass filter in front of the lens simultaneously. Under white light illumination, a stepped drive current signal is applied to the variable focus liquid lens to continuously change the diopter within a preset depth range in fixed steps. Each change in diopter triggers the CMOS image sensor to expose one frame, thereby obtaining a white light focus stack image sequence. Under ultraviolet illumination, the same driving current ladder sequence and exposure time sequence were used to repeatedly scan and obtain the autofluorescence of the insect's epidermal chitin and internal metabolites under ultraviolet excitation, and obtain the ultraviolet fluorescence focus stack image sequence registered with the spatial position of the white light focus stack. Multi-scale fusion processing based on non-downsampled contour wave transform was performed on the white light focus stacked image sequence and the ultraviolet fluorescence focus stacked image sequence respectively. The non-downsampled pyramid filter bank was used to decompose each frame of the image in the stack into a low-pass sub-band and several bandpass direction sub-bands, and the spatial size before downsampling was preserved at each decomposition level. Within the bandpass sub-band, the local phase consistency of pixels is used as a measure of focus sharpness. The corresponding sub-band coefficients of the frame with the largest phase consistency amplitude are selected for fusion. For the low-pass sub-band, the weighted average of the corresponding coefficients of all frames in the stack is taken for fusion. The fused sub-bands are then reconstructed in reverse according to the decomposition level to generate visible light texture images and autofluorescence images.

3. The intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging according to claim 1, characterized in that, The construction of the micro-target detection network involves inputting the visible light texture image and the autofluorescence image into the network, fusing dual-modal features through channel attention, performing separable convolution and spatial offset to form an enhanced feature pyramid, and performing anchor-bound dense classification regression at each layer of the pyramid to extract candidate regions for micro-insect bodies. Specifically, this includes: A small target detection network with a texture-fluorescence dual-stream shallow feature extraction structure is constructed. The texture stream performs continuous convolution operations on visible light texture images to extract morphological details of wing vein orientation, body surface punctures, and color spot boundaries. The fluorescence stream performs convolution operations of the same depth on autofluorescence images to extract bright spot contours and fluorescence intensity distribution features. The shallow feature maps output from the texture stream and fluorescence stream are concatenated along the channel dimension and fed into the channel attention module. The concatenated feature map is compressed into a channel description vector through global average pooling. The weighting coefficients of each channel are generated through a fully connected layer and activation function. The weighting coefficients of each channel are multiplied back into the original feature map for adaptive weighted fusion of dual-modal features. The feature map after channel attention fusion is fed into the dynamic receptive field enhancement module. Multiple sets of depth separable convolutional layers with different dilation rates are arranged in parallel within the module. A learnable spatial offset is introduced during the convolution sampling process so that the receptive field sampling points are adaptively concentrated on the discriminative local structure of the insect body. The outputs of each set are spliced ​​and then compressed and convolved to form a single-layer enhanced feature map. Using the single-layer enhanced feature map as the base layer of the feature pyramid, a multi-scale feature pyramid is constructed layer by layer through top-down upsampling paths and lateral connections. Each layer has different spatial resolution and semantic abstraction to adapt to the scale changes of tiny insects of different body lengths. At each level of the feature pyramid, anchor boxes of various scales and aspect ratios are densely pre-set around each spatial location. Convolutional predictions of classification and regression branches are performed in parallel for each anchor box. The classification branch outputs the confidence score containing the tiny insect, and the regression branch outputs the bounding box offset. The prediction results of all levels of the pyramid are aggregated, and a preliminary filter is performed using a preset relaxed confidence threshold. Non-maximum suppression is then applied to the retained anchor boxes to remove duplicate boxes on the same worm, ultimately generating candidate regions for tiny worms.

4. The intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging according to claim 1, characterized in that, For the candidate region of the tiny insect, the pixel-by-pixel focus curve in the white light focus stack is fitted with Gaussian interpolation to generate a depth topography map. This map is then stacked with the visible light texture image and the autofluorescence image and fed into a multi-task instance segmentation network to obtain the foreground mask, contour distance map, and orientation vector field, and to construct a joint edge cost map for segmenting densely clustered insect instances. Specifically, this includes: For each candidate region of a tiny insect, the improved Laplacian sum of each frame is calculated pixel by pixel in the white light focus stack image sequence with a preset window to generate a focus curve as the focus metric value. The focus curve is then fitted with a Gaussian function, and the depth index corresponding to the peak of the fitted curve is used as the height value of the insect surface corresponding to that pixel. Arrange the height values ​​of all pixels according to their spatial positions to form a depth topography map that is pixel-level aligned with the candidate region. The depth topography map shows height gradient changes at the junction depressions of the adhered worm bodies and at the interlayer transitions between stacked layers. The visible light texture image block, autofluorescence image block and the depth topography map corresponding to the same candidate region are stacked in the channel dimension to form a multi-channel input tensor, which is then imported into a multi-task instance segmentation network to obtain the foreground mask, contour distance map and orientation vector field. The multi-task instance segmentation network adopts a shared encoder and multi-branch decoder structure. The shared encoder extracts abstract features that fuse texture, fluorescence and depth morphology information from the multi-channel input. The foreground mask branch, contour distance map branch and orientation vector field branch are arranged in parallel. The three branches share the same encoder features and are jointly optimized during training. The contour distance map, orientation vector field and depth topography map are fused pixel by pixel to construct a joint edge cost map. The contour distance map provides the individual boundary prior on the two-dimensional plane of the worm, the orientation vector field provides the pixel orientation change signal at the adhesion boundary, and the depth topography map provides the interlayer stacking and surface depression information in the three-dimensional space of the worm. The effective processing area is constrained by the foreground mask. Local minimum points of the contour distance map within the foreground range are detected on the joint edge cost map as seed points for each potential individual. A marker watershed algorithm is introduced to simulate the water immersion process from various seed points. When the immersion fronts of different seed points meet at the ridge line of the cost map, a segmentation boundary is formed. Finally, the densely adhered and stacked worm bodies are cut into independent single instances.

5. The intelligent monitoring method for tiny insects in fruits and vegetables based on microscopic magnification imaging according to claim 1, characterized in that, The insect segmentation example extracts corresponding region blocks for each independent insect from visible light texture images, autofluorescence images, and depth topography images. These blocks are then input into a pre-constructed three-branch hybrid classification network to extract multimodal features. The extracted multimodal features are then used to identify the insect species and determine its liveness / death status through a multi-classification head. Specifically, this includes: Based on the outline boundary of each insect segmentation instance, visible light texture region blocks, autofluorescence region blocks, and depth morphology region blocks with corresponding positions and sizes are cropped from the visible light texture image, autofluorescence image, and depth morphology image, respectively. Each region block contains only the pixel information of the independent insect itself. The visible light texture region block, autofluorescence region block, and depth morphology region block are respectively fed into the corresponding branches of the pre-constructed three-branch hybrid classification network. The three-branch hybrid classification network consists of a texture feature extraction branch, a fluorescence feature extraction branch, and a depth morphology feature extraction branch. The three branches are structurally independent of each other. The texture feature extraction branch extracts microscopic morphological features, including the direction of wing vein bifurcation, the arrangement of body surface punctures, and the outline of color spots, from the visible light texture region block by stacking multiple layers of small-sized convolution kernels. It also gradually compresses the spatial dimension through pooling operations to integrate the global texture representation of various parts of the insect body. The fluorescence feature extraction branch uses the same convolution depth as the texture branch but with independent weight parameters. It extracts the intensity distribution pattern of autofluorescence signal, the spatial concentration of fluorescent bright spots, and the decay gradient features of fluorescence from the center of the insect body to the edge from the autofluorescence region block layer by layer, forming a feature expression that reflects the physiological state of the insect body. The deep morphology feature extraction branch extracts three-dimensional geometric features from the deep morphology region block, including the curvature distribution of the worm's back, the height difference between segments, and the relative inclination angle between the worm and the attachment surface, to describe the worm's morphology features. Finally, the feature vectors extracted by the three branches are spliced ​​and fused in the fully connected layer to form a multimodal feature vector. On top of the fully connected layer, a fine-grained insect species classification head and a binary classification head for live / dead states are set up in parallel. The insect species classification head outputs the probability of the target insect segmentation instance belonging to each insect species, and the insect species with the highest probability is taken as the identification result. The live / dead state classification head outputs the binary classification probability value of live insects and dead insects.

6. The intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging according to claim 1, characterized in that, The process involves collecting multimodal monitoring data of the target area within a preset period, performing spatiotemporal modeling of population dynamics and predicting the reproductive cycle based on the multimodal monitoring data, introducing a Bayesian inference algorithm to infer insect infestation at future times based on the population dynamics evolution and reproductive cycle prediction results, and finally generating insect infestation situation inference information for early warning push. Specifically, this includes: All single-insect data collected multiple times within a preset monitoring period are aggregated according to the collection timestamp and spatial location label to form a multimodal monitoring dataset. Each record contains the insect species label, live / dead status label, spatial coordinates, and the cumulative environmental temperature value corresponding to the collection of a single insect in a single collection. The multimodal monitoring dataset is aggregated according to a preset statistical granularity, and the number of live insects, the number of dead insects, and the total insect population density of each insect species are counted in each time period. The average temperature and effective accumulated temperature are obtained from the environmental sensors deployed in the monitoring area in the corresponding time period to form the insect population density time series and the accumulated temperature time series. A hybrid prediction framework combining a fusion mechanism model and a data-driven model is constructed. The first pathway calculates the accumulation rate of temperature based on the developmental starting temperature and effective accumulated temperature constant of each target insect species, combined with daily accumulated temperature data. The rate and model output population trend prediction values ​​based on biological mechanisms are then used. The second pathway inputs the insect population density time series into a long short-term memory network with an attention mechanism. Through a gating mechanism, it selectively memorizes and forgets the long-term trends, periodic fluctuations, and short-term mutations in the historical insect population density series. The attention mechanism assigns weights to different historical moments and outputs a data-driven predicted value of insect population density. A Bayesian inference algorithm is introduced to integrate the population trend prediction and insect population density prediction under a unified probability framework. The posterior probability distribution of the fused insect population density is calculated, and the posterior mean is output as the fused prediction value and confidence interval. A three-level early warning judgment logic is constructed based on the posterior probability distribution. When the posterior mean of the fused predicted density is lower than the proliferation threshold, it is judged as low density level; when it exceeds the proliferation threshold and the lower limit of the confidence interval breaks through the rising judgment line, it is judged as proliferation period; when it exceeds the outbreak threshold and the posterior probability is greater than the confidence level requirement, it is judged as outbreak period. Insect situation inference information is generated and early warning is pushed.

7. A smart monitoring system for tiny insects in fruits and vegetables based on microscopic magnification imaging, characterized in that, The system includes: a memory, a processor, and a communication interface. The memory contains a program for an intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging. When the processor executes the program for the intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging, it implements the steps of the intelligent monitoring method for microscopic insects in fruits and vegetables based on microscopic magnification imaging as described in any one of claims 1-6 above.