An image recognition-based termite alate adult flight monitoring method and system
Patent Information
- Application Number
- CN202611031980.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-13
AI Technical Summary
[0003]传统白蚁监测方式主要依赖人工定期巡查、埋设饵剂诱集装置等手段,不仅人力投入大、监测效率低,且受巡查周期限制,难以捕捉分飞行为的突发时段,基于图像识别的白蚁监测技术已有一定发展,通常采用诱捕灯等装置利用白蚁的趋光性进行诱集,再通过摄像头采集图像,利用图像识别算法对诱捕到的虫体进行识别与计数,然而,现有方案多采用深度学习模型对采集的图像进行逐帧全图检测,但是夜间监测场景中无目标的空背景时段占比极高,全时段全量推理会产生大量无效算力消耗,造成计算资源的浪费,且多数方案仅针对单只白蚁的形态特征进行识别,未充分利用白蚁分飞特有的群体行为规律,在蚊虫、飞蛾等多类飞虫混杂的夜间灯光环境下,误检率与漏检率较高,且当分飞群体密度较高时,虫体相互遮挡会导致个体识别与计数精度大幅下降;
本发明通过逐像素计算的亮度方差生成二值掩码以捕捉运动信号,在此基础上,采用预设尺寸的非重叠规则网格对掩码进行空间分块,计算各网格的运动像素激活率,进而筛选待分析网格和候选网格,能够从全画面中快速锁定疑似白蚁活动区域,之后基于筛选出至少N个邻接待分析网格的连续时间窗口,利用白蚁具有规律性盘旋行为学上的规律,进行加权质心追踪进而计算相对振荡频次,作为是否为疑似白蚁成虫分飞的条件之一,之后,以最新质心坐标为基准截取感兴趣区域,双通道条形卷积核提取区域内各像素的水平与垂直梯度分量,以获取边缘柔和度和主方向能量占比并进行分析,本地端即可初步判断是否为疑似白蚁成虫分飞区域,并保留低遮挡区域的候选网格图像,按需云端复核,实现低算力的白蚁有翅成虫分飞监测。
Smart Images

Figure CN122530958B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of termite swarming monitoring technology, specifically to a method and system for monitoring the swarming of winged adult termites based on image recognition. Background Technology
[0002] During the swarming period, the number of winged adult termites can range from several hundred to tens of thousands. Issuing an early warning when the colony size reaches a certain threshold has important engineering application value for guiding subsequent prevention and control work and reducing the risk of termite infestation.
[0003] Traditional termite monitoring methods mainly rely on regular manual inspections and the placement of bait traps. These methods are not only labor-intensive and inefficient, but also limited by the inspection cycle, making it difficult to capture sudden periods of swarming behavior. Image recognition-based termite monitoring technology has made some progress. Typically, devices such as trap lights are used to attract termites by taking advantage of their phototaxis. Images are then captured by cameras, and image recognition algorithms are used to identify and count the trapped insects. However, existing solutions often use deep learning models to perform frame-by-frame full-image detection on the captured images. But in nighttime monitoring scenarios, there is a high proportion of empty background periods without targets. Full-time, full-data inference will consume a lot of ineffective computing power, resulting in a waste of computing resources. Moreover, most solutions only identify the morphological characteristics of individual termites, failing to fully utilize the unique group behavior patterns of termite swarming. In nighttime lighting environments with a mix of mosquitoes, moths, and other flying insects, the false detection and false detection rates are high. Furthermore, when the density of swarming groups is high, the insects occlude each other, leading to a significant decrease in the accuracy of individual identification and counting. Therefore, there is a need for a method to detect termite swarming at a low cost and around the clock by identifying the behavioral and morphological characteristics of termite swarming when the termite colony size reaches a certain threshold.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for monitoring the swarming of winged adult termites based on image recognition, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for monitoring the swarming of winged adult termites based on image recognition, comprising the following steps: Step 1: Map the continuous frame images of the current time window to the luminance channel, calculate the luminance variance pixel by pixel and generate a motion binary mask by threshold segmentation, divide the mask using a preset non-overlapping grid, calculate the motion pixel activation rate of each grid, and filter the grid to be analyzed and candidate grids based on the motion pixel activation rate. Step 2: If at least N spatially adjacent grids to be analyzed are detected within M consecutive time windows, the mask centroid is solved by weighting the activation rate of the moving pixels of each grid, a temporal centroid sequence is constructed, and the total number of effective oscillation peaks and the relative oscillation frequency in the vertical direction of the centroid are calculated. If the total number of effective oscillation peaks is 2 and the relative oscillation frequency falls within the preset first frequency range, then proceed to step 3. Step 3: Extract the region of interest using the latest centroid coordinates, extract the horizontal and vertical gradient components within the region using a dual-channel strip convolution kernel, fuse them to obtain a gradient magnitude map, and use the result of global average pooling as the feather edge softness. Solve the edge direction angle based on the horizontal and vertical gradient components to calculate the energy proportion of the swarm's main flight direction. Step 4: After initially screening suspected targets by combining edge softness and main direction energy ratio on the local end, the candidate grid images that have been screened and are located in the region of interest are called and input into the trained deep learning recognition model. The binarized result is output to determine whether there are winged adult termites flying.
[0007] Furthermore, the method for calculating the motion pixel activation rate of each grid is as follows: Take any frame of image, and use the motion binary mask to count the total number of motion marker pixels in a single grid. The ratio of the total number of motion marker pixels in a single grid to the total number of pixels inside the grid is defined as the motion pixel activation rate of the frame image corresponding to that grid.
[0008] Furthermore, the method for filtering the grid to be analyzed and candidate grids based on the activation rate of moving pixels is as follows: A first single-frame activation rate interval and a second single-frame activation rate interval are preset, and the lower limit of the first single-frame activation rate interval is higher than the upper limit of the second single-frame activation rate interval. In the current time window, if the motion pixel activation rate of all frames of a single grid falls into the first single-frame activation rate interval, the grid is determined to be the grid to be analyzed. When there is a grid to be analyzed in the current time window, grids that satisfy the 8-connected domain relationship with the grid to be analyzed and whose motion pixel activation rate of at least one frame falls into the second single-frame activation rate interval are selected and marked as candidate grids.
[0009] Furthermore, the method for constructing a temporal centroid sequence by weighting the activation rates of the moving pixels in each grid to solve for the mask centroid is as follows: For all spatially adjacent grids to be analyzed within a single time window, the motion pixel activation rate corresponding to the latest frame of the image in the current time window is selected as the weighting weight. The geometric center coordinates of each grid to be analyzed are weighted and averaged to obtain the mask centroid coordinates corresponding to the current time window. The mask centroid coordinates obtained from M consecutive time windows are arranged in chronological order to complete the construction of the temporal centroid sequence.
[0010] Furthermore, the method for calculating the total number of effective oscillation peaks and the relative oscillation frequency in the vertical direction of the centroid is as follows: Extract the vertical coordinate components of each mask centroid coordinate in the time-series centroid sequence, and generate a centroid vertical displacement sequence in chronological order; perform smoothing filtering on the centroid vertical displacement sequence to obtain a smooth displacement sequence that retains the fluctuation trend; calculate the first-order forward difference on the smooth displacement sequence to obtain the vertical velocity sequence. Traverse the vertical velocity sequence, identify the intervals where the vertical velocity reverses from positive to negative, and the intervals where the vertical velocity reverses from negative to positive, and extract the coordinate components of the transition position of the vertical velocity in the interval from positive to negative or from negative to positive in the smooth displacement sequence. The adjacent coordinate components are extracted sequentially according to time, and the absolute value of the difference between them is calculated. If the absolute value of the difference is greater than the preset amplitude threshold, the time-series centroid sequence is marked as having a valid oscillation peak. If the absolute value of the difference is less than or equal to the preset amplitude threshold, it is not marked as a valid oscillation peak, and the total number of valid oscillation peaks is obtained. Extract the timestamps of the latest frame images corresponding to the first and last time windows of the time-series centroid sequence, and calculate the time difference between the two sets of timestamps; The sum of the time spans of the coordinate components corresponding to all effective oscillation peaks is calculated as the effective oscillation duration. Data points corresponding to the time sequence positions of the effective oscillation peaks are removed from the centroid vertical displacement sequence as candidate stationary sequences to be screened. For the candidate stationary sequences to be screened, the first data point is initialized as the first stationary segment. The mean of all data points in this segment is used as the segment reference vertical displacement. Subsequent data points in the sequence are read sequentially, and the absolute value of the deviation between the data point and the segment reference vertical displacement is calculated. If the absolute value of the deviation is not greater than the preset height fluctuation threshold, the data point is included in the current stationary segment. If the absolute value of the deviation is greater than the preset height fluctuation threshold, the stationary segment information is archived. A new stationary segment is generated with this data point as the boundary until all stationary segments are divided. The duration of all stationary segments is compared, and the maximum value is taken as the longest stationary flight duration. The ratio of the total number of effective oscillation peaks to the longest stationary flight duration is calculated, and the result is taken as the relative oscillation frequency in the vertical direction of the centroid.
[0011] Furthermore, the region of interest is truncated using the latest centroid coordinates, and the horizontal and vertical gradient components within the region are extracted using a dual-channel strip convolution kernel: Using the centroid coordinates of the mask obtained from the latest frame of the temporal centroid sequence at the end of the time window as the center, a rectangular region of fixed pixel size is delineated and cropped to obtain the region of interest (ROI) of the image. A dual-channel bar convolution kernel is set, with one channel configured with a horizontal bar convolution kernel to extract the horizontal gradient features of the ROI image, and the other channel configured with a vertical bar convolution kernel to extract the vertical gradient features of the ROI image. The ROI image is then convolved with the dual-channel bar convolution kernel to separate the horizontal and vertical gradient components of the output image.
[0012] Furthermore, the method for calculating the energy percentage of the swarm's main flight direction is as follows: The 0°~180° angle range is uniformly divided into several equal-width angle sub-intervals. Based on the horizontal and vertical gradient components of each pixel in the region of interest, the edge direction angle of each pixel is solved by the arctangent function. The edge direction angles of all pixels in the region of interest are counted, and the number of pixels falling in each angle sub-interval is collected. The angle sub-interval with the most pixels is the main direction interval. The ratio of the number of pixels in the main direction interval to the total number of pixels in the region of interest is used as the energy proportion of the main direction of the swarm flight.
[0013] Furthermore, the method for initial screening of suspected targets locally, combining edge smoothness and main direction energy ratio, is as follows: The edge smoothness threshold and the main direction energy ratio threshold are preset. The obtained wing edge smoothness and insect swarm flight main direction energy ratio are matched with the two preset thresholds respectively. When the wing edge smoothness is not higher than the corresponding smoothness threshold and the main direction energy ratio is not lower than the corresponding ratio threshold, the current target is determined to be a suspected target of termite swarm. Otherwise, it is determined to be an environmental interference target and is removed.
[0014] Furthermore, the method for determining whether winged adult termites swarm is as follows: For any candidate grid image, input it into the trained deep learning recognition model. If the output binarization result is 0, it is determined that the candidate grid image does not contain winged adult termites flying in formation; if the output binarization result is 1, it is determined that the candidate grid image contains winged adult termites flying in formation. When the binarization result of all candidate grid images is 0, it is determined that there are no winged adult termites flying in the monitoring area; conversely, as long as the binarization result of any candidate grid image is 1, it is determined that there are winged adult termites flying in the monitoring area.
[0015] Additionally, an image recognition-based monitoring system for swarming of winged adult termites is provided, characterized in that: the system is used to execute the image recognition-based monitoring method for swarming of winged adult termites described above, including: The grid classification module is used to map the continuous frame images of the current time window to the brightness channel, calculate the brightness variance pixel by pixel and generate a motion binary mask by threshold segmentation, use a preset non-overlapping grid to divide the mask, calculate the motion pixel activation rate of each grid, and filter the grid to be analyzed and candidate grids based on the motion pixel activation rate. The kinematics calculation module is used to solve the mask centroid by weighting the activation rate of the moving pixels of each grid if at least N spatially adjacent grids to be analyzed are detected within M consecutive time windows, construct a temporal centroid sequence, calculate the total number of effective oscillation peaks and the relative oscillation frequency in the vertical direction of the centroid, and if the total number of effective oscillation peaks is 2 and the relative oscillation frequency falls within the preset first frequency range, then enter the morphological calculation module. The gradient calculation module is used to extract the region of interest with the latest centroid coordinates, extract the horizontal and vertical gradient components within the region using a dual-channel strip convolution kernel, fuse them to obtain a gradient magnitude map, and use the result of global average pooling as the feather edge softness. Based on the horizontal and vertical gradient components, the edge direction angle is solved to calculate the energy proportion of the main direction of swarm flight. The termite detection module is used to initially screen suspected targets locally by combining edge softness and main direction energy ratio. Then, it calls the selected candidate grid images located in the region of interest, inputs them into the trained deep learning recognition model, and outputs a binarized result to determine whether there are winged adult termites flying.
[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention generates a binary mask by calculating the brightness variance pixel by pixel to capture motion signals. Based on this, the mask is spatially divided into blocks using a non-overlapping regular grid of a preset size. The activation rate of moving pixels in each grid is calculated, and then the grid to be analyzed and candidate grids are screened. This allows for the rapid identification of suspected termite activity areas from the entire image. Then, based on a continuous time window of at least N neighboring grids to be analyzed, the relative oscillation frequency is calculated by weighted centroid tracking, taking advantage of the regular hovering behavior of termites. This frequency is used as one of the conditions for determining whether a suspected termite adult swarming area is present. Subsequently, the region of interest is cropped based on the latest centroid coordinates. A dual-channel strip convolution kernel is used to extract the horizontal and vertical gradient components of each pixel within the region to obtain edge softness and the energy ratio of the main direction, which are then analyzed. The local device can preliminarily determine whether a suspected termite adult swarming area is present, and candidate grid images of low-occlusion areas are retained for on-demand cloud verification, achieving low-computing-power monitoring of winged adult termites swarming. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall method flow of the present invention; Figure 2 This is a schematic diagram of the temporal centroid sequence of the present invention; Figure 3 This is a schematic diagram of the overall system structure of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] Example: Please see Figures 1 to 2 The present invention provides a technical solution: A method for monitoring the swarming of winged adult termites based on image recognition, comprising the following steps: Step 1: Map the continuous frame images of the current time window to the luminance channel, calculate the luminance variance pixel by pixel and generate a motion binary mask by threshold segmentation, divide the mask using a preset non-overlapping grid, calculate the motion pixel activation rate of each grid, and filter the grid to be analyzed and candidate grids based on the motion pixel activation rate. The system converts multiple consecutive RGB color images within the current time window to a single-channel logarithmic luminance space. The frame number is determined based on the duration of the time window. To eliminate transient interference from insects and birds flying by quickly and to capture the brief dwelling behavior of termites during swarming, the duration of the time window is set to 0.5-1 seconds. The number of consecutive frames corresponding to a single time window is calculated based on the nominal frame rate of the monitoring camera. The time window adopts a frame-by-frame sliding update mechanism: each time a newest monitoring image is acquired, that image is added to the end of the frame sequence of the current time window, while the earliest frame at the beginning of the frame sequence is removed, so that the updated time window always maintains a fixed duration and number of frames. In converting consecutive frames of RGB color images to a single-channel logarithmic luminance space, for example, it can be based on the RGB three channels converted to linear luminance according to standard luminance weights, and then the natural logarithm of the luminance values is taken. The linear luminance value ranges from 0 to 255, and the effective numerical range of the logarithmic luminance is approximately 0 to 5.54. This method can compress the dynamic range of high-luminance areas and expand the contrast of low-luminance areas, achieving strong anti-interference capability in the display of features of trap lights and streetlights under nighttime lighting conditions. The specific method for calculating the luminance variance pixel by pixel and generating a motion binary mask by threshold segmentation is as follows: extract the logarithmic luminance value of the pixel position in all frames of the current time window to calculate the variance, and then obtain the luminance variance distribution map. Preset motion judgment A variance threshold is defined, and threshold segmentation is performed on the brightness variance distribution map. Pixels with variance not less than the motion determination variance threshold are marked as motion states, and pixels with variance less than the motion determination variance threshold are marked as background states. This serves as the motion binary mask for the latest frame. When winged adult termites flap their wings and crawl a short distance within the monitoring area, the brightness of the corresponding pixel position fluctuates continuously and periodically. Therefore, the variance threshold ranges from 0.03 to 0.08. This calculation method fully utilizes all frame data within the time window to complete motion determination, rather than relying on the difference between a single frame or two adjacent frames. It can effectively suppress random noise in a single frame, instantaneous light fluctuations, and interference from small targets that pass by quickly, so as to extract the persistent target area in the subsequent process. A preset non-overlapping grid partitioning mask is used. The size of the non-overlapping grid is set so that when the number of termites swarming reaches thousands and gathers in the area of the insect-attracting lamp, there are at least 4 grids to be analyzed. For example, in a typical monitoring scenario with a resolution of 1024×1024, a grid size of 64×64 pixels is selected. The dense movement area formed by thousands of adult termites can cover 4 to 9 consecutive adjacent grids to be analyzed. Take any frame image, and count the total number of motion marker pixels in a single grid based on the motion binary mask. The ratio of the total number of motion pixels in a single grid to the total number of pixels inside the grid is defined as the motion pixel activation rate of the frame image corresponding to that grid.
[0021] A first single-frame activation rate interval and a second single-frame activation rate interval are preset, and the lower limit of the first single-frame activation rate interval is higher than the upper limit of the second single-frame activation rate interval. In the current time window, if the motion pixel activation rate of all frames of a single grid falls into the first single-frame activation rate interval, the grid is determined to be the grid to be analyzed. When there is a grid to be analyzed in the current time window, grids that satisfy the 8-connected domain relationship with the grid to be analyzed and whose motion pixel activation rate of at least one frame falls into the second single-frame activation rate interval are selected and marked as candidate grids.
[0022] Step 2: If at least N spatially adjacent grids to be analyzed are detected within M consecutive time windows, the mask centroid is solved by weighting the activation rate of the moving pixels of each grid, a temporal centroid sequence is constructed, and the total number of effective oscillation peaks and the relative oscillation frequency in the vertical direction of the centroid are calculated. If the total number of effective oscillation peaks is 2 and the relative oscillation frequency falls within the preset first frequency range, then proceed to step 3. If at least N spatially adjacent grids to be analyzed are detected within M consecutive time windows, then for all spatially adjacent grids to be analyzed within a single time window, the motion pixel activation rate of each grid in the latest frame image of the current time window is selected as the weight, and the geometric center coordinates of these grids are weighted and averaged to solve for the centroid coordinates of the mask in the current time window. Since the latest frame represents the instantaneous observation state of the target in the monitoring scene, and the motion pixel activation rate directly represents the density of termite activity within each grid, using this as a weight can drive the mask centroid to adaptively shift towards the area where swarming termites are most densely concentrated and most violently moving, effectively avoiding interference from sporadic flying insects or discrete noise points at the edges.
[0023] The mask centroid coordinates obtained from solving M consecutive time windows are arranged in chronological order to construct the temporal centroid sequence. This process relies on weighted centroid extraction to transform the disordered movement of hundreds or thousands of termites into a single macroscopic spatial evolution trajectory, mapping the overall trend of termite swarming from a behavioral perspective. This avoids the problem of traditional single-target tracking failing in high-density, mutually occluding, and trajectories intersecting. The value of N must be no less than 4, aiming to use spatial connectivity logic to eliminate a few high-activation-rate grids formed by occasional isolated moths, large flying creatures, or local reflections. The value of M must ensure that the M consecutive time windows cover a complete swarming-specific up-and-down cycle of the termite colony, providing sufficient time span for subsequent extraction of effective continuous relative oscillation frequencies. A complete swarming-specific up-and-down cycle typically lasts 3-8 seconds. Figure 2 As shown, the up-and-down fluctuation period is about 3 seconds. For example, if the up-and-down fluctuation period of a complete separation is set to 8 seconds, then the timestamp difference between the latest frame image corresponding to the first time window and the last time window will be at least 8 seconds, and the corresponding M value will be calculated.
[0024] Termite adults with wings have relatively weak flight capabilities. When they gather in phototaxis, the centroid of the colony inevitably exhibits periodic low-frequency oscillations in the vertical direction. The vertical coordinate components of the centroid coordinates of each mask in the time-series centroid sequence are extracted, and a centroid vertical displacement sequence is generated in chronological order. Considering that the extracted point set is discrete and inevitably contains transient high-frequency noise caused by environmental perturbations, a smoothing filter based on local polynomial least squares fitting, such as Savitzky-Golay filtering, is used to denoise the sequence to obtain a smooth displacement sequence. Compared with traditional moving average filtering, polynomial smoothing filtering can effectively filter out small fluctuations under high frame rate sampling while preserving the extreme width and height of the macroscopic signal to the greatest extent. To remove transient high-frequency noise, the length of the filtering sliding window is set according to the camera frame rate, for example, to cover the number of consecutive frames corresponding to a duration of 0.3s-0.5s. The polynomial order is set to the 3rd order. While smoothing out disordered spikes between adjacent frames, the inflection point, peak and valley positions and amplitudes are completely preserved. Next, the first-order forward difference is calculated on the smoothed displacement sequence to obtain the vertical velocity sequence. The significance of calculating the first-order forward difference is that it mathematically transforms the complex comparison of physical coordinates into a direct determination of the trend of macroscopic motion direction changes. The vertical velocity sequence is traversed to identify intervals where the vertical velocity reverses from positive to negative, and intervals where it reverses from negative to positive. The coordinate components of the turning points where the vertical velocity changes from positive to negative or from negative to positive in each interval are extracted from the smoothed displacement sequence. Adjacent coordinate components are extracted sequentially according to time, and the absolute value of their difference is calculated. If the absolute value of the difference is greater than a preset amplitude threshold, the interval is marked. The temporal centroid sequence has a valid oscillation peak; if the absolute value of this difference is less than or equal to a preset amplitude threshold, it is not marked as a valid oscillation peak, thus obtaining the total number of valid oscillation peaks; wherein, the setting of the preset amplitude threshold needs to be based on the actual physical scale of the macroscopic stall fall of the termite colony, and combined with the non-overlapping grid size in step 1 for constraint mapping, based on the spatial calibration parameters of the monitoring camera, and based on the observation of the actual physical distance of the termite stall fall, for example, converting the physical drop of a typical single fall of 15cm into the corresponding image pixel difference as the amplitude threshold, wherein the formula for calculating the amplitude threshold is: In the formula, Indicates the amplitude threshold. Indicates the camera's focal length. This represents the physical drop a termite typically experiences in a single fall. This indicates the calibration distance from the camera to the center of the insect-attracting light. This indicates the size of a single pixel in the camera's image sensor. Pick , Pick , Pick , Pick At that time, the calculated amplitude threshold is 90 pixels; Extract the timestamps of the latest frame images corresponding to the first and last time windows of the time-series centroid sequence, and calculate the time difference between the two sets of timestamps; The sum of the time spans of the coordinate components corresponding to all effective oscillation peaks is calculated as the effective oscillation duration. Data points corresponding to the time positions of the effective oscillation peaks are removed from the centroid vertical displacement sequence, becoming candidate stationary sequences to be screened. For each candidate stationary sequence, the first data point is initialized as the first stationary segment. The mean of all data points within this segment is used as the segment's reference vertical displacement. Subsequent data points are read sequentially, and the absolute value of the deviation between the data point and the segment's reference vertical displacement is calculated. If the absolute value of the deviation is not greater than a preset height fluctuation threshold, the data point is included in the current stationary segment. If the absolute value of the deviation is greater than the preset height fluctuation threshold, the stationary segment information is archived, and a new stationary segment is generated using this data point as the boundary. This process continues until all stationary segments are divided. The duration of all stationary segments is compared, and the longest duration is selected. The maximum value is taken as the longest stable flight time. The ratio of the total number of effective oscillation peaks to the longest stable flight time is calculated, and the result is taken as the relative oscillation frequency in the vertical direction of the centroid. The altitude fluctuation threshold is the reasonable base noise generated when the termite colony hovers and stays at a high position around the light. Multiple frames of centroid vertical displacement data corresponding to the stable hovering and high-positioning stage of termites can be collected, and the standard deviation of this set of data can be calculated. Twice the standard deviation is taken as the altitude fluctuation threshold to accommodate the small fluctuations during normal stable flight and avoid incorrectly segmenting the continuous stable flight state. For example, for a 1080P resolution camera, multiple frames of termite images excluding the vertical fluctuation period are taken. After calculating the centroid vertical displacement data according to the above steps, 8 pixels are obtained by calculating twice the standard deviation. Then, the altitude fluctuation threshold is set to 8 pixels.
[0025] If the total number of effective oscillation peaks is 2 and the relative oscillation frequency falls within the preset first frequency range, then proceed to step 3. The logic for setting the preset first frequency range is as follows: Due to the characteristics of termites being physically weak and having a long time to hover and gather, the time taken for their single drop and return undulation action has a fixed theoretical boundary. Based on observations of multiple up-and-down undulation cycles, this theoretical boundary usually lasts for 3 to 8 seconds. The timestamp difference between the latest frame images corresponding to the first and last time windows is extracted as the total time span. The upper limit (8 seconds) and lower limit (3 seconds) of the fluctuation action time are subtracted from this total time span to derive the minimum and maximum theoretical values of the longest stable flight time of a termite colony. Under the condition that the total number of effective oscillation peaks is limited to 2, this value 2 is divided by the maximum and minimum theoretical values respectively. The resulting two frequency values constitute the lower and upper limits of the preset first frequency interval. This interval can be dynamically and adaptively adjusted according to the total video capture duration, excluding flying insects with extremely fast movements that cannot remain stable for long periods. For example, when the total time span of M consecutive time windows is 10 seconds, the upper limit (8 seconds) and lower limit (3 seconds) of the fluctuation action time are subtracted from this total time span to obtain the minimum theoretical value 2 seconds and the maximum theoretical value 7 seconds. The value 2 is then divided by the maximum and minimum theoretical values respectively to obtain the first frequency interval. .
[0026] Step 3: Extract the region of interest using the latest centroid coordinates, extract the horizontal and vertical gradient components within the region using a dual-channel strip convolution kernel, fuse them to obtain a gradient magnitude map, and use the result of global average pooling as the feather edge softness. Combine this with the edge direction angle calculated based on the horizontal and vertical gradient components to calculate the energy proportion of the swarm's main flight direction; Using the centroid coordinates of the mask obtained by solving the latest frame of the temporal centroid sequence at the end of the time window as the center, a rectangular area of fixed pixel size is delineated and cropped to obtain the region of interest in the image. At this time, the target group has just experienced up and down movement, the spatial density is most concentrated and it is in a transient stagnation period, suppressing motion blur. The setting of the fixed pixel size needs to be constrained and mapped according to the spatial resolution of the camera and the typical physical aggregation radius of the termite colony. For example, it is set to a pixel matrix size that can cover a real physical space of 20cm×20cm to ensure that the area contains a sufficient number of termite individuals. Due to the unique biological morphology of termites, which have elongated bodies and narrow, semi-transparent membranous wings, a dual-channel strip convolution kernel was used to replace the traditional square isotropic convolution kernel. One channel is configured with a horizontal strip convolution kernel, specifically for extracting horizontal gradient features from the region of interest (ROI) image, while the other channel is configured with a vertical strip convolution kernel for extracting vertical gradient features. This dual-channel strip structure can spatially decouple the horizontal and vertical slender structural energy of the target, improving the sensitivity to the perception of the narrow wing boundaries of termites. The length of the strip convolution kernel is limited to an odd number slightly larger than the length of a single unit to effectively integrate the continuous line features of the narrow wings. For example, when the ROI region is 256×256 pixels, the dimensions of the horizontal and vertical strip convolution kernels are set to 1×21 and 21×1, respectively. The ROI image is then convolved with the dual-channel strip convolution kernel, and the horizontal and vertical gradient components of the output image are separated accordingly.
[0027] The 0°~180° angle range is uniformly divided into several equal-width sub-intervals, for example, each sub-interval is 15°. Based on the horizontal and vertical gradient components of each pixel within the region of interest, the edge direction angle of each pixel is calculated using the arctangent function. The edge direction angles of all pixels within the region of interest are counted, and the number of pixels falling within each angle sub-interval is collected. The angle sub-interval with the largest number of pixels is designated as the main direction interval. The ratio of the number of pixels in the main direction interval to the total number of pixels within the region of interest is taken as the main direction energy proportion of the swarm flight. The calculation of the main direction energy proportion of the swarm flight... The significance lies in the fact that ordinary flying insects hover randomly within a region, and their edge directional angles are evenly and discretely distributed in each interval; while termite colonies, constrained by the wind resistance effect of their large wings, exhibit a highly consistent streamlined spatial arrangement of their bodies and wings during swarming or phototropic sprints due to aerodynamics. Therefore, the angle sub-interval with the most pixels is selected as the main direction interval, and the ratio of the number of pixels in this main direction interval to the total number of edge pixels in the region is taken as the energy proportion of the main direction of the swarm's flight. The higher this proportion, the stronger the morphological consistency of the colony under aerodynamic constraints. The gradient magnitude map of the region of interest is obtained by taking the square root of the sum of the squares of the horizontal and vertical gradient components, pixel by pixel. Since termite wings have the biological characteristics of being translucent and membranous, they appear as low-contrast, gently fading edges that blend into the background in the image, which is significantly different from the high-frequency sharp edges of flying insects with hard shells. Therefore, a global average pooling operation is performed on the gradient magnitude map, that is, the arithmetic mean of the gradient magnitudes of all pixels in the region is calculated. This average value is used as the softness of the wing edge. The magnitude of this value inversely represents the sharpness of the target edge, thereby identifying whether the target has a translucent, membranous, soft edge.
[0028] Step 4: After initially screening suspected targets by combining edge softness and main direction energy ratio on the local end, the candidate grid images that are selected and located in the region of interest are called and input into the trained deep learning recognition model, and the binarized result is output to determine whether there are winged adult termites flying. The wing edge smoothness is the arithmetic mean of the gradient magnitudes of all pixels within the region of interest. A lower value indicates a smoother edge transition and lower sharpness, corresponding to the optical characteristics of translucent, membranous feathers. A higher value indicates a sharper edge, corresponding to keratinized, hard wings, insects with rigid bodies, or strong reflective interference. Based on the global average gradient magnitude distribution of statistical samples, the wing edge smoothness typically falls between 0.07 and 0.16. The edge smoothness threshold is set to 0.16. The main direction energy ratio is the ratio of the number of pixels in the main interval of the edge direction angle to the total number of edge pixels in the region of interest. The higher the value, the stronger the consistency of the spatial arrangement of the insect colony's body and wings, corresponding to the streamlined collective movement characteristics of the termite colony under aerodynamic constraints. The lower the value, the more disordered the individual flight direction, corresponding to the random hovering behavior of ordinary flying insects. By statistically analyzing the samples of termite swarms during the collective take-off and landing phase, the main direction ratio can reach 25%~40%. Therefore, the main direction energy ratio threshold is set to 0.25. The obtained wing edge smoothness and insect swarm flight main direction energy ratio are matched with two sets of preset thresholds. When the wing edge smoothness is not higher than the corresponding smoothness threshold and the main direction energy ratio is not lower than the corresponding ratio threshold, the current target is determined to be a suspected target of the termite swarm. Otherwise, it is determined to be an environmental interference target and is removed.
[0029] The grid to be analyzed is the core aggregation area of the termite colony, with extremely high density of individual termites, severe overlap and occlusion, and fragmented individual morphological features. The deep learning model has difficulty extracting effective identification features and is prone to misclassification. In contrast, the candidate grid is located in the outer transition area of the colony. It is confirmed to be adjacent to the grid to be analyzed through the 8-connectivity rule. The termite individuals are relatively independent, with complete morphology, and key identification features such as wings, body, and antennae are clear. The model's identification accuracy is significantly higher. Therefore, for any candidate grid image selected and located within the region of interest, the deep learning recognition model is input to perform binary classification and output the result. If the model outputs a positive result indicating winged adult termites in all candidate grid images, it is determined that winged adult termites are swarming; if the model outputs a negative result indicating environmental interference, it is determined that winged adult termites are not swarming. The local end only uploads a small number of extracted candidate grid images to the cloud and uses the trained deep learning recognition model for binary classification. Since reasoning is only performed on local slices as needed, the concurrent computation pressure on the model is reduced. Based on the above embodiments, the deep learning recognition model is constructed using a deep learning network based on a multilayer perceptron. The deep neural network of the multilayer perceptron includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. The first hidden layer, the second hidden layer, and the third hidden layer each have at least two neurons and all use ReLU as the activation function. Historical monitoring candidate grid images are collected, converted to grayscale, and flattened into one-dimensional pixel feature vectors. Each image is manually labeled to indicate whether it contains winged adult termites (1 indicates presence, 0 indicates absence), thus constructing a binary classification sample set. The samples are divided into training, validation, and test sets in a 7:2:1 ratio. The training set is used to learn the model parameters; the validation set is used to adjust hyperparameters during training to prevent overfitting; and the test set is used to evaluate the generalization ability of the model after training.
[0030] The structure of a deep learning network with a multilayer perceptron is as follows: Input layer: Used to receive the pixel feature vectors of the candidate grid image after flattening, with a dimension of 1024; The first hidden layer has 128 neurons and uses ReLU as the activation function. The second hidden layer has 64 neurons and also uses the ReLU activation function; The third hidden layer has 32 neurons and uses the ReLU activation function; Output layer: It has 1 neuron, uses the Sigmoid activation function, and outputs the probability that there is a winged adult termite in the candidate grid.
[0031] The process of training a deep learning model is as follows: The flattened vectors of candidate grid images are used as model input, and the corresponding manually labeled binary labels are used as ground truth supervision labels. A binary cross-entropy loss function is selected. The loss value is calculated between the existence probability of the deep learning model's predicted output and the ground truth label. The deep learning model parameters are iteratively updated using the backpropagation algorithm. When the validation set loss value stabilizes and there is no significant decrease in loss over 20 consecutive training epochs, the deep learning model is considered converged and training is stopped. For example, in actual inference, if the output probability is greater than 0.5, each image contains a winged adult termite and is marked as 1; otherwise, it is marked as 0.
[0032] Please see Figure 3 The present invention also provides an image recognition-based monitoring system for swarming of winged adult termites, used to execute the image recognition-based monitoring method for swarming of winged adult termites described in any of the above claims, comprising: The grid classification module is used to map the continuous frame images of the current time window to the brightness channel, calculate the brightness variance pixel by pixel and generate a motion binary mask by threshold segmentation, use a preset non-overlapping grid to divide the mask, calculate the motion pixel activation rate of each grid, and filter the grid to be analyzed and candidate grids based on the motion pixel activation rate. The kinematics calculation module is used to solve the mask centroid by weighting the activation rate of the moving pixels of each grid if at least N spatially adjacent grids to be analyzed are detected within M consecutive time windows, construct a temporal centroid sequence, calculate the total number of effective oscillation peaks and the relative oscillation frequency in the vertical direction of the centroid, and if the total number of effective oscillation peaks is 2 and the relative oscillation frequency falls within the preset first frequency range, then enter the morphological calculation module. The gradient calculation module is used to extract the region of interest with the latest centroid coordinates, extract the horizontal and vertical gradient components within the region using a dual-channel strip convolution kernel, fuse them to obtain a gradient magnitude map, and use the result of global average pooling as the feather edge softness. Based on the horizontal and vertical gradient components, the edge direction angle is solved to calculate the energy proportion of the main direction of swarm flight. The termite detection module is used to initially screen suspected targets locally by combining edge softness and main direction energy ratio. Then, it calls the selected candidate grid images located in the region of interest, inputs them into the trained deep learning recognition model, and outputs a binarized result to determine whether there are winged adult termites flying.
[0033] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0034] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0035] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0036] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for monitoring the swarming of winged adult termites based on image recognition, characterized in that, The specific steps include: Step 1: Map the continuous frame images of the current time window to the luminance channel, calculate the luminance variance pixel by pixel and generate a motion binary mask by threshold segmentation, divide the mask using a preset non-overlapping grid, calculate the motion pixel activation rate of each grid, and filter the grid to be analyzed and candidate grids based on the motion pixel activation rate. Step 2: If at least N spatially adjacent grids to be analyzed are detected within M consecutive time windows, the mask centroid is solved by weighting the activation rate of the moving pixels of each grid, a temporal centroid sequence is constructed, and the total number of effective oscillation peaks and the relative oscillation frequency in the vertical direction of the centroid are calculated. If the total number of effective oscillation peaks is 2 and the relative oscillation frequency falls within the preset first frequency range, then proceed to step 3. Step 3: Extract the region of interest using the latest centroid coordinates, extract the horizontal and vertical gradient components within the region using a dual-channel strip convolution kernel, fuse them to obtain a gradient magnitude map, and use the result of global average pooling as the feather edge softness. Solve the edge direction angle based on the horizontal and vertical gradient components to calculate the energy proportion of the swarm's main flight direction. Step 4: After initially screening suspected targets by combining edge softness and main direction energy ratio on the local end, the candidate grid images that are selected and located in the region of interest are called and input into the trained deep learning recognition model, and the binarized result is output to determine whether there are winged adult termites flying. The method for calculating the motion pixel activation rate of each grid is as follows: Take any frame image, and use the motion binary mask to count the total number of motion marker pixels in a single grid. The ratio of the total number of motion marker pixels in a single grid to the total number of pixels inside the grid is defined as the motion pixel activation rate of the frame image corresponding to that grid. The method for filtering the grid to be analyzed and candidate grids based on the activation rate of moving pixels is as follows: A first single-frame activation rate interval and a second single-frame activation rate interval are preset, and the lower limit of the first single-frame activation rate interval is higher than the upper limit of the second single-frame activation rate interval. In the current time window, if the motion pixel activation rate of all frames of a single grid falls into the first single-frame activation rate interval, the grid is determined to be the grid to be analyzed. When there is a grid to be analyzed in the current time window, the grids that satisfy the 8-connected domain relationship with the grid to be analyzed and whose motion pixel activation rate of at least one frame falls into the second single-frame activation rate interval are selected and marked as candidate grids. The method for calculating the total number of effective oscillation peaks and the relative oscillation frequency in the vertical direction of the centroid is as follows: Extract the vertical coordinate components of each mask centroid coordinate in the time-series centroid sequence, and generate a centroid vertical displacement sequence in chronological order; perform smoothing filtering on the centroid vertical displacement sequence to obtain a smooth displacement sequence that retains the fluctuation trend; calculate the first-order forward difference on the smooth displacement sequence to obtain the vertical velocity sequence. Traverse the vertical velocity sequence, identify the intervals where the vertical velocity reverses from positive to negative, and the intervals where the vertical velocity reverses from negative to positive, and extract the coordinate components of the transition position of the vertical velocity in the interval from positive to negative or from negative to positive in the smooth displacement sequence. The adjacent coordinate components are extracted sequentially according to time, and the absolute value of the difference between them is calculated. If the absolute value of the difference is greater than the preset amplitude threshold, the time-series centroid sequence is marked as having a valid oscillation peak. If the absolute value of the difference is less than or equal to the preset amplitude threshold, it is not marked as a valid oscillation peak, and the total number of valid oscillation peaks is obtained. Extract the timestamps of the latest frame images corresponding to the first and last time windows of the time-series centroid sequence, and calculate the time difference between the two sets of timestamps; The sum of the time spans of the coordinate components corresponding to all effective oscillation peaks is calculated as the effective oscillation duration. Data points corresponding to the time sequence positions of the effective oscillation peaks are removed from the centroid vertical displacement sequence as candidate stationary sequences to be screened. For the candidate stationary sequences to be screened, the first data point is initialized as the first stationary segment. The mean of all data points in this segment is used as the segment reference vertical displacement. Subsequent data points in the sequence are read sequentially, and the absolute value of the deviation between the data point and the segment reference vertical displacement is calculated. When the absolute value of the deviation is not greater than the preset height fluctuation threshold, the data point is included in the current stationary segment. When the absolute value of the deviation is greater than the preset height fluctuation threshold, the stationary segment information is archived. A new stationary segment is generated with this data point as the boundary until all stationary segments are divided. The duration of all stationary segments is compared, and the maximum value is taken as the longest stationary flight duration. The ratio of the total number of effective oscillation peaks to the longest stationary flight duration is calculated, and the result is taken as the relative oscillation frequency in the vertical direction of the centroid. The method of extracting the region of interest using the latest centroid coordinates and then using a dual-channel strip convolution kernel to extract the horizontal and vertical gradient components within the region is as follows: Using the centroid coordinates of the mask obtained from the latest frame of the temporal centroid sequence at the end of the time window as the center, a rectangular region of fixed pixel size is delineated and cropped to obtain the region of interest (ROI) of the image. A dual-channel bar convolution kernel is set, with one channel configured with a horizontal bar convolution kernel to extract the horizontal gradient features of the ROI image, and the other channel configured with a vertical bar convolution kernel to extract the vertical gradient features of the ROI image. The ROI image is then convolved with the dual-channel bar convolution kernel to separate the horizontal and vertical gradient components of the output image. The method for calculating the energy percentage of the swarm's main flight direction is as follows: The 0°~180° angle range is uniformly divided into several equal-width angle sub-intervals. Based on the horizontal and vertical gradient components of each pixel in the region of interest, the edge direction angle of each pixel is solved by the arctangent function. The edge direction angles of all pixels in the region of interest are counted, and the number of pixels falling in each angle sub-interval is collected. The angle sub-interval with the most pixels is the main direction interval. The ratio of the number of pixels in the main direction interval to the total number of pixels in the region of interest is used as the energy proportion of the main direction of the swarm flight.
2. The method for monitoring the swarming of winged adult termites based on image recognition according to claim 1, characterized in that: The method for constructing a temporal centroid sequence by weighting the activation rates of the moving pixels in each grid as the weights is as follows: For all spatially adjacent grids to be analyzed within a single time window, the motion pixel activation rate corresponding to the latest frame of the image in the current time window is selected as the weighting weight. The geometric center coordinates of each grid to be analyzed are weighted and averaged to obtain the mask centroid coordinates corresponding to the current time window. The mask centroid coordinates obtained from M consecutive time windows are arranged in chronological order to complete the construction of the temporal centroid sequence.
3. The method for monitoring the swarming of winged adult termites based on image recognition according to claim 1, characterized in that: The local method for initial screening of suspected targets, combining edge smoothness and main direction energy ratio, is as follows: The edge smoothness threshold and the main direction energy ratio threshold are preset. The obtained wing edge smoothness and insect swarm flight main direction energy ratio are matched with the two preset thresholds respectively. When the wing edge smoothness is not higher than the corresponding smoothness threshold and the main direction energy ratio is not lower than the corresponding ratio threshold, the current target is determined to be a suspected target of termite swarm. Otherwise, it is determined to be an environmental interference target and is removed.
4. The method for monitoring the swarming of winged adult termites based on image recognition according to claim 1, characterized in that: The method for determining whether winged adult termites swarm is as follows: For any candidate grid image, input it into the trained deep learning recognition model. If the output binarization result is 0, it is determined that the candidate grid image does not contain winged adult termites flying apart. If the output binarization result is 1, it is determined that the candidate grid image contains winged adult termites flying in formation. When the binarization result of all candidate grid images is 0, it is determined that there are no winged adult termites flying in the monitoring area; conversely, as long as the binarization result of any candidate grid image is 1, it is determined that there are winged adult termites flying in the monitoring area.
5. A monitoring system for swarming of winged adult termites based on image recognition, characterized in that: The system is used to perform the image recognition-based method for monitoring the swarming of winged adult termites as described in any one of claims 1-4: The grid classification module is used to map the continuous frame images of the current time window to the brightness channel, calculate the brightness variance pixel by pixel and generate a motion binary mask by threshold segmentation, use a preset non-overlapping grid to divide the mask, calculate the motion pixel activation rate of each grid, and filter the grid to be analyzed and candidate grids based on the motion pixel activation rate. The kinematics calculation module is used to solve the mask centroid by weighting the activation rate of the moving pixels of each grid if at least N spatially adjacent grids to be analyzed are detected within M consecutive time windows, construct a temporal centroid sequence, calculate the total number of effective oscillation peaks and the relative oscillation frequency in the vertical direction of the centroid, and if the total number of effective oscillation peaks is 2 and the relative oscillation frequency falls within the preset first frequency range, then enter the morphological calculation module. The gradient calculation module is used to extract the region of interest with the latest centroid coordinates, extract the horizontal and vertical gradient components within the region using a dual-channel strip convolution kernel, fuse them to obtain a gradient magnitude map, and use the result of global average pooling as the feather edge softness. Based on the horizontal and vertical gradient components, the edge direction angle is solved to calculate the energy proportion of the main direction of swarm flight. The termite detection module is used to initially screen suspected targets locally by combining edge softness and main direction energy ratio. Then, it calls the selected candidate grid images located in the region of interest, inputs them into the trained deep learning recognition model, and outputs a binarized result to determine whether there are winged adult termites flying.
Citation Information
Patent Citations
On-line monitoring device for winged adult termites based on image recognition
CN117671735A
Flying insect monitoring system and method
EP4039089A1