A distribution early warning method based on termite distribution inference model

By classifying termite diet groups and using a group-specific prediction model, the problem of insufficient prediction accuracy caused by ignoring the differences in environmental response between groups in existing technologies has been solved, enabling accurate distribution early warning and targeted prevention and control of different termite groups.

CN120950863BActive Publication Date: 2026-05-12HUNAN JIANGSHANMEI ECOLOGICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN JIANGSHANMEI ECOLOGICAL TECH CO LTD
Filing Date
2025-07-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies fail to effectively differentiate the responses of different dietary groups to environmental factors in termite control, resulting in insufficient prediction accuracy and an inability to provide targeted control guidance.

Method used

By using a distribution early warning method based on termite diet group classification, multi-source data is obtained, a group-specific prediction model is established, a sub-group distribution probability map is generated, and targeted early warning information is output.

Benefits of technology

It enables accurate prediction of the distribution of different termite feeding groups, improves the targeting and efficiency of prevention and control, and reduces resource waste and ecological disturbance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950863B_ABST
    Figure CN120950863B_ABST
Patent Text Reader

Abstract

The application discloses a termite distribution early warning method based on a termite distribution inference model, and relates to the technical field of termite prevention and control and ecological monitoring. The method comprises the following steps: obtaining multi-source data associated with a target geographic area, wherein the multi-source data comprises historical termite activity records, environmental factor data and engineering structure parameters; assigning termite samples in the historical termite activity records to one or more specific dietary groups according to a preset termite dietary group classification system, and generating a group activity data set. The application solves the problem of insufficient prediction accuracy caused by ignoring the differences in environmental response among groups in the prior art by classifying termites according to dietary groups and constructing a dedicated prediction model. The application can accurately capture the specific correlation between different groups and environmental factors, and the generated group distribution probability map can clearly show the spatial distribution state of each group, thereby avoiding overlapping of distribution characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of termite control and ecological monitoring technology, specifically a distribution early warning method based on a termite distribution inference model. Background Technology

[0002] In the field of termite control, current technologies can collect data on termite surface activity characteristics, biological characteristics, and environmental preferences through multiple channels. This data is then integrated with key parameters such as engineering structure parameters and soil type. Standardization is performed using statistical analysis or machine learning algorithms, and appropriate mathematical models are selected for training. By combining GIS technology and remote sensing imagery for spatial distribution modeling, active termite areas can be visualized, and dynamic monitoring systems can be established to adjust predictions and early warnings based on new information.

[0003] However, existing technologies have significant limitations in refining the responses of different termite populations to environmental factors. Different termite populations, including wood-dwelling, soil-dwelling, soil-and-wood-amphibian, and fungal-infested termite colonies, exhibit significant differences in their environmental requirements for survival and reproduction. Current models mostly treat the termite as a whole, failing to deeply distinguish the subtle differences in environmental preferences among different population groups, particularly lacking technical solutions for developing dedicated predictive models for different populations. This limits the accuracy of models in predicting the distribution of different termite populations within a specific area, hindering the provision of more differentiated guidance for targeted control and failing to meet the needs of precise location and control of different termite types in practical pest management work. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a distribution early warning method and system based on termite diet group classification, which improves the accuracy of predicting the distribution of different termite diet groups and provides differentiated guidance for targeted prevention and control.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a distribution early warning method based on a termite distribution inference model, applied to a target geographical area, the method comprising the following steps:

[0006] Acquire multi-source data associated with the target geographic area, including historical termite activity records, environmental factor data, and engineering structure parameters;

[0007] Based on a pre-defined termite diet group classification system, termite samples from the historical termite activity records are assigned to one or more specific diet groups to generate subgroup activity datasets.

[0008] For each of the one or more specific dietary groups, a group-specific prediction model is created;

[0009] Based on the group-specific prediction model, a sub-group distribution probability map is generated to characterize the spatial distribution status of each specific dietary group within the target geographical area;

[0010] Based on the group distribution probability map and the preset early warning triggering logic, distribution early warning information for specific dietary groups is generated and output.

[0011] Furthermore, the step of creating a group-specific prediction model includes:

[0012] For each specific dietary group, extract the corresponding activity point data from the subgroup activity dataset;

[0013] By integrating the environmental factor data and the engineering structure parameters corresponding to the activity point data in the spatiotemporal dimension, a group-specific original feature dataset is constructed.

[0014] The original feature dataset specific to the group is standardized to generate a group-specific training set for model training.

[0015] Using a pre-defined machine learning algorithm, a group-specific prediction model is obtained by training on the group-specific training set. This model internalizes the specific association rules between the specific dietary group and environmental factors.

[0016] Furthermore, the preset termite diet group classification system is based on the termites' main food sources, and the specific diet groups include:

[0017] Wood-dwelling termite colonies primarily feed on woody structures and cellulose;

[0018] Soil-dwelling termite colonies primarily feed on humus and organic matter in the soil;

[0019] A type of termite colony that is both terrestrial and soil-dwelling, exhibiting feeding characteristics that combine terrestrial and soil-dwelling characteristics;

[0020] In addition, there are termite colonies that cultivate fungi in their nests as a food source.

[0021] Furthermore, the environmental factor data includes:

[0022] Climate data, including annual average temperature, monthly extreme temperature, annual average rainfall, and relative humidity data;

[0023] Soil data, which includes soil type, soil pH, soil organic matter content, and soil moisture content;

[0024] Vegetation data, which includes vegetation coverage index, vegetation type and surface litter layer thickness data;

[0025] Topographic data, which includes elevation, slope and aspect data.

[0026] Furthermore, the step of constructing a group-specific original feature dataset also includes:

[0027] For each specific dietary group, an initial subset of environmental factors is selected from the environmental factor data based on a pre-defined group preference knowledge base;

[0028] The completeness and redundancy of the initial environmental factor subset in characterizing the living environment of this specific dietary group are evaluated to obtain the evaluation results;

[0029] If the evaluation results indicate insufficient completeness or excessive redundancy, the initial subset of environmental factors will be adjusted to form an optimized subset of environmental factors.

[0030] The optimized environmental factor subset or the initial environmental factor subset is used as the environmental factor input for constructing the original feature dataset specific to the group.

[0031] Furthermore, the step of evaluating the completeness and redundancy of the initial subset of environmental factors in characterizing the living environment of the specific dietary group includes:

[0032] A feature selection algorithm is applied to calculate the correlation score between each environmental factor in the initial subset of environmental factors and the distribution of activity points of the specific dietary group;

[0033] Calculate the mutual information score or correlation coefficient among the environmental factors in the initial subset of environmental factors;

[0034] Based on preset scoring thresholds and mutual information thresholds, low-relevance factors and high-redundancy factors are identified, and the evaluation results are generated.

[0035] Furthermore, the step of generating a sub-group distribution probability map representing the spatial distribution status of each specific dietary group within the target geographical area includes:

[0036] The target geographic region is gridded to generate a set of geographic cells;

[0037] For each geographic cell in the set of geographic cells, extract its corresponding environmental factor data and engineering structure parameters to form a feature vector to be predicted.

[0038] For each specific dietary group, the corresponding group-specific prediction model is applied to the feature vector to be predicted, and the probability value of the occurrence of the dietary group in the geographic cell is calculated.

[0039] The probability values ​​of all geographic cells are combined to form a spatial distribution probability layer for this specific dietary group;

[0040] The spatial distribution probability layers of all specific dietary groups are overlaid to form the subgroup distribution probability map.

[0041] Furthermore, the warning triggering logic includes:

[0042] For each specific dietary group, a probability threshold and a spatial clustering threshold are pre-set;

[0043] When the distribution probability of a specific dietary group in the subgroup distribution probability map is continuously higher than its corresponding probability threshold in a certain continuous area, and the area or number of cells in that area reaches the spatial clustering threshold, an early warning is triggered.

[0044] The warning information includes the type of dietary group being warned, the geographical coordinates of the warning area, and the average probability of occurrence within that area.

[0045] Furthermore, the method also includes the following dynamic update step:

[0046] Receive newly added termite activity monitoring data, which includes the location of new activity points and their corresponding food group types;

[0047] The newly added termite activity monitoring data and its corresponding environmental factor data are appended to the corresponding sub-group activity dataset and the original feature dataset specific to the group.

[0048] According to a preset update strategy, a retraining process is triggered on one or more affected group-specific prediction models to generate an updated group-specific prediction model.

[0049] The update strategies include: a trigger strategy based on data increments reaching a predetermined number, or a timed trigger strategy based on a preset time period.

[0050] Furthermore, the step of triggering the retraining process for one or more affected group-specific prediction models further includes:

[0051] Before performing retraining, the performance metrics of the current population-specific prediction model are evaluated using an independent validation dataset.

[0052] Perform the retraining process to generate a candidate updated prediction model;

[0053] Using the independent validation dataset, evaluate the performance metrics of the candidate updated prediction models;

[0054] Compare the performance metrics of the current population-specific prediction model with the performance metrics of the candidate updated prediction model. If the latter's performance metrics are better than the former's, then the candidate updated prediction model is used to replace the current population-specific prediction model.

[0055] Record the model replacement operation, the corresponding performance metric changes, and the timestamp information that triggered the operation.

[0056] Beneficial effects

[0057] This invention addresses the problem of insufficient prediction accuracy in existing technologies due to neglecting differences in environmental responses between termite groups by classifying termites according to their dietary groups and constructing a dedicated prediction model. It accurately captures the specific correlations between different groups and environmental factors, and the generated sub-group distribution probability maps clearly show the spatial distribution of each group, avoiding overlapping distribution characteristics. Based on this, early warning information can provide precise distribution warnings for specific groups, offering differentiated guidance for prevention and control, and improving the targeting and efficiency of prevention and control. Simultaneously, the dynamic update mechanism retrains the model with new data, ensuring that the model continuously adapts to environmental changes and the evolution of termite behavior, avoiding prediction bias caused by outdated data, guaranteeing the reliability of early warning information, and reducing resource waste and ecological disturbance. Attached Figure Description

[0058] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] While multi-source data fusion and machine learning achieve basic visualization, they fail to decouple the different responses of dietary groups to environmental factors. When multiple groups coexist, the system trains a global model using a unified feature set, masking group-specific characteristics and resulting in overlapping distribution features in the predictions, failing to provide differentiated early warning coordinates. This problem has serious consequences: prevention requires broad-spectrum pest control, leading to resource waste and ecological disturbance, such as mismatched protective measures and failure of biological control. Simultaneously, due to the lack of group type identification, new data cannot trigger incremental learning of the corresponding model, creating a negative feedback loop of model degradation and niche migration. This invention addresses the core contradiction. Traditional models treat termites as homogeneous objects, solving the problem by starting with data partitioning and model decoupling: introducing group classification in the preprocessing stage, splitting data according to feeding characteristics; and building independent models for each subset. Ultimately, a combination of these two approaches is chosen, using sub-group datasets to train dedicated models, and then overlaying them to generate a composite distribution map.

[0061] For this, please refer to Figure 1 This invention proposes a distribution early warning method based on a termite distribution inference model, applied to a target geographical area. The method includes the following steps:

[0062] Acquire multi-source data associated with the target geographic area, including historical termite activity records, environmental factor data, and engineering structure parameters;

[0063] Based on a pre-defined termite diet group classification system, termite samples from historical termite activity records are assigned to one or more specific diet groups to generate subgroup activity datasets.

[0064] For each dietary group within one or more specific dietary groups, a group-specific prediction model is created; based on the group-specific prediction model, a sub-group distribution probability map representing the spatial distribution status of each specific dietary group within the target geographic area is generated;

[0065] Based on the probability distribution map of different groups and the preset early warning triggering logic, distribution early warning information for specific dietary groups is generated and output.

[0066] Multi-source data refers to historical termite activity records, environmental factor data, and engineering structure parameters related to the target geographic area. Specifically, it can be obtained through sensor networks, manual survey records, or geographic information system databases to comprehensively reflect the correlation between termite activity and environmental conditions.

[0067] The termite diet group classification system refers to the classification criteria established based on the differences in the main food objects of termites. Specifically, it can be defined through biological characteristic analysis or expert experience to distinguish the differences in the response of different termite groups to environmental factors.

[0068] Among them, the sub-group activity dataset refers to the data set formed by classifying historical termite activity records according to their dietary groups. This can be achieved through data annotation and classification algorithms, and is used to build an independent analysis basis for different groups.

[0069] Among them, the group-specific prediction model refers to the machine learning model built for a specific dietary group. Specifically, it can be trained using algorithms such as random forest and support vector machine to capture the specific association rules between the group and environmental factors.

[0070] Among them, the subgroup distribution probability map refers to a visual layer that represents the spatial distribution probability of different dietary groups. Specifically, it can be generated by overlaying geographic gridding and probability prediction to identify high-risk areas.

[0071] Among them, the early warning triggering logic refers to the early warning condition judgment rules based on probability thresholds and spatial clustering. Specifically, it can be implemented by setting threshold parameters and spatial statistical algorithms to dynamically trigger targeted early warning information.

[0072] The core innovation of this invention lies in solving the problem of insufficient prediction accuracy caused by ignoring the differences in environmental response between termite groups in the prior art by dividing termite feeding groups and constructing a group-specific prediction model, thereby achieving accurate distribution early warning for different termite groups.

[0073] The working process and principle of this invention are as follows: First, multi-source data associated with the target geographical area are acquired, including historical termite activity records, environmental factor data, and engineering structure parameters. This data provides the foundation for subsequent analysis. Then, based on a preset termite diet group classification system, termite samples from historical termite activity records are assigned to specific diet groups, generating sub-group activity datasets. This step achieves the differentiation of different diet groups. Next, a group-specific prediction model is created for each diet group. By establishing an independent model for each group, the specific responses of each group to environmental factors can be captured. Based on these group-specific prediction models, a sub-group distribution probability map representing the spatial distribution status of each specific diet group within the target geographical area is generated. Finally, based on the sub-group distribution probability map and a preset early warning triggering logic, distribution early warning information for specific diet groups is generated and output. This process realizes a complete workflow from data collection, group classification, model building to early warning output.

[0074] As a preferred embodiment, the present invention is implemented as follows: First, multi-source data of the target geographical area is obtained through government departments, research institutions, and termite control companies. Historical termite activity records include the location coordinates and activity times of termite activity points recorded over the past 5 years. Environmental factor data covers temperature and humidity data from meteorological stations, soil parameters from soil monitoring stations, and vegetation indices interpreted from remote sensing images. Engineering structure parameters include information such as building type and service life. Next, using a termite diet group classification system based on morphological features and DNA barcoding technology, termite samples from historical activity records are classified into wood-dwelling, soil-dwelling, soil-and-wood amphibious, and fungus-growing termite groups, generating sub-group activity datasets. Then, for each diet group, a group-specific prediction model is created using the random forest algorithm. The model inputs are environmental factors and engineering structure parameters, and the output is the probability of the group appearing at a specific location. Based on the trained model, the target area is gridded for prediction, generating spatial distribution probability layers for each group, which are then overlaid to form a sub-group distribution probability map. Finally, set up warning triggering logic. When the distribution probability of a certain group in a continuous area exceeds a preset threshold and the area reaches a specified size, generate warning information including the group type, warning area range, and average occurrence probability.

[0075] Through the above-described scheme, this invention achieves refined distribution prediction and early warning for termites with different dietary groups. By classifying termite samples into specific dietary groups and creating a dedicated prediction model for each group, it effectively solves the problem of neglecting the differences in the responses of different groups to environmental factors in traditional methods. The generated sub-group distribution probability map clearly shows the distribution status of each dietary group within the target area, avoiding overlapping group distribution characteristics. Based on this, the output early warning information can provide accurate distribution warnings for specific dietary groups, providing a basis for pest control personnel to formulate differentiated control strategies and improving the targeting and efficiency of termite control.

[0076] The present invention further proposes the following steps for creating a group-specific prediction model: for each specific dietary group, extracting corresponding activity point data from the sub-group activity dataset; integrating environmental factor data and engineering structure parameters corresponding to the activity point data in the spatiotemporal dimension to construct a group-specific original feature dataset; standardizing the group-specific original feature dataset to generate a group-specific training set for model training; and using a preset machine learning algorithm to train the model based on the group-specific training set to obtain a group-specific prediction model, which internalizes the specific association rules between the specific dietary group and environmental factors.

[0077] The extraction of activity point data was achieved by filtering records in the subgroup activity dataset that were associated with the dietary group, ensuring a strict correspondence between the training data and the biological characteristics of the target group. The integration of environmental factor data and engineering structure parameters employed a spatiotemporal alignment method, matching multi-source data within the same geographical location and time range to form a multi-dimensional feature vector. Standardization was performed using Z-score normalization to eliminate the influence of different units on model training. Machine learning algorithms, such as random forests or gradient boosting decision trees, were selected to establish a mapping between the population and environmental factors through feature importance assessment and rule extraction.

[0078] Specifically, after screening, the activity point data is spatiotemporally correlated with corresponding environmental factors and engineering structural parameters to form a dataset containing features such as temperature, soil type, and building structure. Standardization transforms each feature value into a distribution with a mean of 0 and a variance of 1, avoiding model weight shifts caused by differences in numerical ranges. During training, the machine learning algorithm iteratively optimizes the feature weight allocation through multiple rounds, identifying environmental factors that significantly affect the distribution of specific dietary groups; for example, wood-dwelling groups are more sensitive to wood structure parameters than earth-dwelling groups. The final predictive model can quantify the degree of influence of different environmental factors on the distribution of the target group, achieving accurate probabilistic predictions.

[0079] As a preferred embodiment, the solution of the present invention is implemented as follows:

[0080] When creating a population-specific prediction model, the first step is to extract corresponding activity point data from the sub-population activity dataset for each specific dietary group. For example, for a wood-dwelling termite colony, all activity record points marked as wood-dwelling are extracted.

[0081] Next, environmental factor data and engineering structure parameters corresponding to the activity point data in the spatiotemporal dimensions are integrated to construct a raw feature dataset specific to the group. Specifically, for each activity point, information such as climate data (e.g., temperature, rainfall), soil data (e.g., pH value, moisture content), vegetation data (e.g., cover), and surrounding building types are collected corresponding to its time and location.

[0082] Then, the original feature dataset specific to the group is standardized to generate a group-specific training set for model training. The standardization process includes normalizing numerical features and converting categorical features into one-hot encodings.

[0083] Finally, a group-specific prediction model is obtained by training a pre-defined machine learning algorithm on a group-specific training set. For example, a random forest algorithm can be used, and through multiple iterations of training, a model capable of accurately predicting the distribution probability of a specific dietary group can be obtained. This model internalizes the specific association rules between a particular dietary group and environmental factors.

[0084] Through the above technical solution, this invention achieves refined modeling for different termite dietary groups. Therefore, the model can capture the unique response patterns of each dietary group to environmental factors, improving the accuracy of predicting the distribution of different termite dietary groups within a specific area. Furthermore, this differentiated prediction result provides more targeted guidance for targeted control, effectively meeting the needs of precise location and control of different termite types in practical control work.

[0085] This invention further proposes a termite diet group classification system based on the termites' main food sources. Specific diet groups include wood-dwelling termite groups, soil-dwelling termite groups, soil-and-wood amphibious termite groups, and fungus-growing termite groups.

[0086] Among them, wood-dwelling termite colonies mainly feed on woody structures and cellulose, soil-dwelling termite colonies mainly feed on humus and organic matter in the soil, amphibious termite colonies exhibit both wood-dwelling and soil-dwelling feeding characteristics, and fungus-growing termite colonies cultivate fungi in their nests as a food source. This classification system establishes colony division criteria based on the biological differences in their feeding targets, providing a clear basis for grouping activity data. For example, the activity data of wood-dwelling colonies correspond to the distribution information of woody materials in engineering structural parameters, while the activity data of soil-dwelling colonies correspond to the soil organic matter content parameter.

[0087] Specifically, based on the classification criteria of feeding targets, samples from historical termite activity records can be accurately categorized into their corresponding feeding groups. When creating a group-specific prediction model, training data for wood-dwelling groups will focus on associating with wood material parameters in engineering structures, while training data for soil-dwelling groups will be associated with soil humus content parameters. Activity data for termite colonies in fungal nurseries needs to be combined with vegetation cover index and the thickness of the surface litter layer, as their fungal cultivation depends on a specific vegetation environment. By precisely matching group activity data with corresponding environmental factors, the model can internalize the specific survival rules of different groups. For example, wood-dwelling groups show a significantly increased probability of occurrence in areas where the moisture content of wood structures is higher than 15%, while soil-dwelling groups show increased activity frequency when the soil pH is between 6.0 and 7.5. This classification system gives the sub-group activity datasets biological rationality, thereby improving the spatial resolution of subsequent probability distribution maps and the accuracy of early warning information.

[0088] As a preferred embodiment, the solution of the present invention is implemented as follows:

[0089] The pre-defined termite diet classification system is based on the termites' main food sources. Specific diet groups include: wood-dwelling termite colonies, which mainly feed on wood structures and cellulose; soil-dwelling termite colonies, which mainly feed on humus and organic matter in the soil; amphibious termite colonies, which have both wood-dwelling and soil-dwelling feeding characteristics; and fungus-growing termite colonies, which cultivate fungi in the nest as a food source.

[0090] Specifically, wood-dwelling termite colonies mainly include some species from the genera *Dryopteris* and *Dactylogyrus*. These termites tend to nest inside or near wood, feeding primarily on wood fibers. Soil-dwelling termite colonies typically include most species from the genus *Odontotermes*, which nest in the soil and feed on organic matter. Amphibious termite colonies include species such as *Odontotermes maculatus*, which can move freely between soil and wood, offering a wider range of food sources. Fungus-growing termite colonies mainly refer to species from the genus *Macrotermes*, which cultivate specific fungi in their nests and feed primarily on fungi.

[0091] Furthermore, corresponding environmental preference indicators can be established for each dietary group. For example, wood-dwelling termite colonies may rely more on the abundance of wood resources and humidity conditions in the environment; soil-dwelling termite colonies may pay more attention to soil pH and organic matter content; amphibious termite colonies need to consider both wood resources and soil conditions; and fungus-growing termite colonies may have higher requirements for temperature and humidity stability.

[0092] Therefore, when predicting termite distribution, specialized prediction models can be built for different dietary groups, thereby improving the accuracy and relevance of the prediction.

[0093] Through the above technical solution, this invention achieves refined classification of termite colonies, providing a more accurate foundation for subsequent distribution prediction and control measures. By distinguishing the ecological needs and environmental preferences of different dietary groups, the prediction model can better capture the distribution characteristics of each colony, thereby improving prediction accuracy. Furthermore, this classification method also provides guidance for developing targeted control strategies, enabling control measures to be more precisely targeted at specific types of termite colonies, improving control efficiency and reducing unnecessary resource waste.

[0094] This invention further proposes environmental factor data including climate data, which includes annual average temperature, monthly extreme temperature, annual average rainfall, and relative humidity data; soil data, which includes soil type, soil pH, soil organic matter content, and soil moisture content data; vegetation data, which includes vegetation cover index, vegetation type, and surface litter layer thickness data; and topographic data, which includes altitude, slope, and aspect data.

[0095] Among the data, climate data reflects the thermal conditions of termite activity areas through annual average temperature, monthly extreme temperatures assess the impact of temperature fluctuations on termite survival, and annual average rainfall and relative humidity jointly characterize the regional humidity environment. Soil data includes soil type and pH to determine soil suitability for termite nesting, and organic matter content and moisture content correlate with termite food resources and habitat humidity. Vegetation data identifies vegetation abundance and species through cover index and type, and the thickness of the surface litter layer reflects the level of humus accumulation. Topographic data shows that altitude affects temperature and air pressure distribution, while slope and aspect determine surface runoff and sunlight conditions.

[0096] Specifically, the annual average temperature in the climate data is quantified in degrees Celsius. For example, the annual average temperature for a certain region is 18-25 degrees Celsius, with monthly extreme temperatures recorded as a summer high of 35 degrees Celsius and a winter low of 5 degrees Celsius. The annual average rainfall is set at 1200-1500 mm, and the relative humidity is maintained at 70%-85%. Soil types are classified into clay, sandy soil, and loam. The pH range is set as neutral (6.5-7.5). Organic matter content is expressed as a percentage, and soil moisture content is determined by gravimetric analysis. The vegetation cover index uses the NDVI value, and vegetation types are classified into broadleaf forest, coniferous forest, and shrubland. The thickness of the surface litter layer is measured in centimeters. In the topographic data, elevation is recorded in meters, slope is expressed as an angle, and slope aspect is divided into eight directions, including north, northeast, and east. These data are standardized and transformed into numerical feature vectors, which are then input into a population-specific prediction model. This allows the model to capture the differentiated responses of different dietary groups to environmental factors. For example, wood-dwelling termites are less sensitive to soil moisture content than soil-dwelling termites, and the distribution of fungal termites is positively correlated with the thickness of the leaf litter layer. Thus, the environmental factor dataset, through structured classification and quantitative indicators, provides highly discriminative feature inputs for model training, improving the prediction accuracy of the probability distribution map.

[0097] As a preferred embodiment, the solution of the present invention is implemented as follows:

[0098] Environmental data includes climate, soil, vegetation, and topographic data. Climate data includes annual average temperature, monthly extreme temperatures, annual average rainfall, and relative humidity. Soil data includes soil type, soil pH, soil organic matter content, and soil moisture content. Vegetation data includes vegetation cover index, vegetation type, and surface litter layer thickness. Topographic data includes elevation, slope, and aspect.

[0099] Specifically, climate data is obtained from weather stations. Annual average temperature and monthly extreme temperatures are recorded in degrees Celsius. Annual average rainfall is measured in millimeters. Relative humidity is expressed as a percentage.

[0100] Soil data were obtained through field sampling and laboratory analysis. Soil types were classified according to the International Soil Classification System. Soil pH was expressed as a percentage of soil pH. Soil organic matter content was measured in grams per kilogram (g / kg). Soil moisture content was expressed as a percentage by volume.

[0101] Vegetation data were obtained through remote sensing imagery and field surveys. The Normalized Difference Vegetation Index (NDVI) was used as the vegetation cover index. Vegetation types were classified according to the dominant vegetation communities. The thickness of the litter layer was measured in centimeters.

[0102] Topographic data was obtained using a digital elevation model (DEM). Elevation is recorded in meters. Slope is expressed in degrees. Aspect is expressed as a 360-degree azimuth.

[0103] Through the above technical solution, this invention comprehensively covers the key environmental factors affecting termite distribution, providing a comprehensive data foundation for subsequently establishing accurate termite distribution prediction models. Climate data reflects regional temperature and humidity conditions, soil data reflects the characteristics of the termite's survival substrate, vegetation data reflects food sources and habitat conditions, and topographic data reflects changes in the microenvironment. This multi-dimensional environmental factor data can more accurately characterize the survival environment preferences of different termite dietary groups, thereby improving the accuracy and reliability of termite distribution prediction.

[0104] This invention further proposes that, for each specific dietary group, an initial subset of environmental factors is selected from environmental factor data based on a pre-defined group preference knowledge base; the completeness and redundancy of the initial environmental factor subset in characterizing the living environment of the specific dietary group are evaluated to obtain an evaluation result; if the evaluation result indicates insufficient completeness or excessive redundancy, the initial environmental factor subset is adjusted to form an optimized environmental factor subset; the optimized environmental factor subset or the initial environmental factor subset is used as the environmental factor input for constructing the original feature dataset specific to the group.

[0105] The group preference knowledge base stores the association rules between different dietary groups and environmental factors. For example, wood-dwelling termite groups correspond to factors such as lignocellulose content and wood moisture content, while soil-dwelling termite groups correspond to factors such as soil organic matter content and humus layer thickness. During the evaluation process, the Pearson correlation coefficient is used to calculate the spatial correlation between each factor and the distribution of activity points. Factors with a correlation coefficient lower than 0.3 are judged as low-correlation factors. At the same time, the Spearman rank correlation coefficient between factors is calculated, and factor pairs with a correlation coefficient higher than 0.7 are considered as highly redundant factors. Adjustment operations include deleting low-correlation factors and retaining a single representative of high-correlation factors. For example, when the correlation coefficient between soil pH and humus content reaches 0.8, only the humus content factor is retained.

[0106] Specifically, taking a wood-dwelling termite colony as an example, the initial environmental factor subset indicated by the colony preference knowledge base includes wood moisture content, leaf litter layer thickness, and soil type. Evaluation revealed a correlation coefficient of 0.15 between soil type and activity point distribution, while the correlation coefficient between wood moisture content and leaf litter layer thickness was 0.75. After adjustment, the soil type factor was removed, and the wood moisture content factor was retained between the wood moisture content and leaf litter layer thickness factors, forming an optimized environmental factor subset. When this subset was used as input to construct the training set, model training time was reduced by 22%, and prediction accuracy improved by 8.6%. By dynamically optimizing the combination of environmental factors, the influence of irrelevant and redundant features on the model was effectively eliminated, improving the spatial resolution of the probability distribution map.

[0107] As a preferred embodiment, the solution of the present invention is implemented as follows:

[0108] For wood-dwelling termite colonies, a pre-defined population preference knowledge base includes key environmental factors such as temperature, humidity, and wood structure density. Annual average temperature, relative humidity, and vegetation cover index are selected as the initial subset of environmental factors from the environmental factor data.

[0109] When evaluating the initial subset of environmental factors, a random forest algorithm was applied to calculate the correlation score between each factor and the distribution of activity points of wood-dwelling termites. Mutual information scores between factors were calculated to identify low-correlation factors (correlation below 0.3) and high-redundancy factors (mutual information above 0.8).

[0110] The assessment results showed that the vegetation cover index had a low correlation with termite distribution, while annual average temperature and relative humidity showed high redundancy. The initial environmental factor subset was adjusted by removing the vegetation cover index and increasing soil moisture content to form an optimized environmental factor subset.

[0111] An optimized subset of environmental factors, including annual average temperature and soil moisture content, was used as the environmental factor input for constructing the population-specific original feature dataset.

[0112] Through the above technical solution, this invention achieves optimized selection of environmental factors for specific termite diet groups. By evaluating the completeness and redundancy of the initial environmental factor subset and making corresponding adjustments, the accuracy of environmental factors in representing the living environment of specific diet groups is improved. The optimized environmental factor subset better reflects the environmental preferences of wood-dwelling termite colonies, providing more accurate input data for subsequently constructing colony-specific prediction models, thereby improving the accuracy of termite distribution prediction.

[0113] This invention further proposes to use a feature selection algorithm to calculate the correlation score between each environmental factor in the initial environmental factor subset and the activity point distribution of the specific dietary group; calculate the mutual information score or correlation coefficient between each environmental factor in the initial environmental factor subset; and identify low-correlation factors and high-redundancy factors based on preset score thresholds and mutual information thresholds to generate evaluation results.

[0114] The feature selection algorithm, employing random forests, chi-square tests, or information gain methods, quantifies the influence of environmental factors on termite activity distribution. Mutual information scores measure the degree of information overlap between two environmental factors, while correlation coefficients are calculated using Pearson or Spearman methods to determine linear or nonlinear relationships. A normalized threshold of 0.3 to 0.5 is set for the score, and a normalized threshold of 0.4 to 0.6 is set for the mutual information threshold to filter out factors with insufficient contribution or redundant information.

[0115] Specifically, in the assessment of wood-dwelling termite colonies, the initial environmental factor subset, including soil pH, vegetation cover index, and monthly extreme temperature, had its correlation scores with the distribution of activity points calculated. Soil pH scores below 0.3 were identified as low-relevance factors. Simultaneously, the Pearson correlation coefficient between soil pH and vegetation cover index was calculated to be 0.82, exceeding the mutual information threshold of 0.6, thus identifying it as a highly redundant factor. By removing the soil pH factor and retaining the vegetation cover index and monthly extreme temperature, an optimized environmental factor subset was formed. After this optimized subset was input into the colony-specific prediction model, model training time was reduced by 18%, and prediction accuracy improved by 12%.

[0116] As a preferred embodiment, the solution of the present invention is implemented as follows:

[0117] Feature selection algorithms are applied to calculate the correlation scores between each environmental factor in the initial subset of environmental factors and the distribution of activity points for a specific dietary group. For example, the random forest algorithm is used to calculate the importance score of each environmental factor, with a score range of 0-1.

[0118] Calculate the mutual information score or correlation coefficient between each environmental factor in the initial subset of environmental factors. For example, the Pearson correlation coefficient can be used to calculate the pairwise correlation between environmental factors, with the correlation coefficient ranging from -1 to 1.

[0119] Based on preset scoring and mutual information thresholds, low-relevance factors and high-redundancy factors are identified, generating evaluation results. Specifically, the correlation score threshold can be set to 0.3, and the mutual information threshold to 0.8. Factors with a correlation score below 0.3 are marked as low-relevance factors, and factor pairs with a mutual information score above 0.8 are marked as high-redundancy factor pairs. The final evaluation results include the correlation score and mutual information score of each environmental factor, as well as the labeling information for low-relevance and high-redundancy factors.

[0120] Through the above technical solution, this invention can objectively evaluate the characterization ability of an initial subset of environmental factors on the living environment of a specific dietary group. This allows for the identification of environmental factors with low contribution to the prediction of the target termite colony distribution, as well as factor combinations with information redundancy. Furthermore, this provides a basis for subsequent optimization of the environmental factor subset, helping to improve the model's prediction accuracy and computational efficiency.

[0121] This invention further proposes a method for generating a sub-group distribution probability map that characterizes the spatial distribution of specific dietary groups within a target geographic area. The method includes: gridding the target geographic area to generate a set of geographic cells; extracting corresponding environmental factor data and engineering structure parameters for each geographic cell to form a feature vector to be predicted; applying a corresponding group-specific prediction model to calculate the probability value of each dietary group; combining the probability values ​​of all cells to form a spatial distribution probability layer; and overlaying the probability layers of all groups to form a sub-group distribution probability map.

[0122] The gridding operation employs a regular geographic grid division method, with cell sizes ranging from 10m x 10m to 100m x 100m based on the target area and data resolution. During the construction of the feature vector to be predicted, spatial alignment between environmental factor data and engineering structure parameters is achieved through coordinate matching via a geographic information system. In the probability layer generation stage, spatial interpolation algorithms are used to smooth the probability values ​​of discrete cells, and the overlay operation utilizes a layer transparency overlay algorithm to achieve the visual fusion of multiple population distributions.

[0123] Specifically, by dividing the target area into equal-area geographical cells, the environmental parameters of each spatial unit are ensured to be uniform, eliminating the interference of spatial heterogeneity on the prediction model. The feature vectors to be predicted are standardized to maintain consistency with the data format during model training, ensuring the reliability of the prediction results. The probability of occurrence of each dietary group is calculated independently for each cell, and a continuous probability surface is generated through spatial interpolation to accurately characterize the hotspots of group distribution. After overlaying multiple group probability layers, overlapping and independent distribution areas of different dietary groups can be intuitively identified, providing data support for the formulation of differentiated prevention and control strategies. This technical solution significantly improves the spatial continuity of distribution prediction results through spatial discretization calculation and probability aggregation mechanisms, enabling the early warning triggering logic to accurately identify clustered areas that meet the area and probability thresholds.

[0124] As a preferred embodiment, the solution of the present invention is implemented as follows:

[0125] The target geographic region is gridded into a set of geographic cells. Each geographic cell is set to 100 meters x 100 meters in size. For each geographic cell, corresponding climate, soil, vegetation, and topographic data are extracted from an environmental factor database, while parameters such as building type and age within that cell are extracted from an engineering structure database. These data are combined into a feature vector to be predicted.

[0126] For the four termite feeding groups—wood-dwelling, soil-dwelling, soil-wood-amphibious, and fungal-growing—a dedicated prediction model was applied to process the feature vectors to be predicted. Each model outputs the probability value of the feeding group appearing in the current geographic cell, ranging from 0 to 1.

[0127] Furthermore, the probability values ​​of all geographic cells are combined to generate a spatial distribution probability layer for each dietary group. For example, in the probability layer for wood-dwelling termites, each cell has a value representing the probability of that type of termite appearing.

[0128] Finally, the probability layers of the four dietary groups were overlaid to form the final sub-group distribution probability map. This map uses different shades of color to represent the distribution probability of each type of termite in different areas, providing a basis for subsequent early warning analysis.

[0129] Through the above technical solution, this invention achieves a precise quantitative characterization of the distribution of different termite feeding groups within a target geographical area. This allows for a direct visualization of the spatial distribution characteristics and differences among various termite species, providing data support for developing targeted control strategies. Furthermore, by integrating the outputs of multiple dedicated prediction models, this method improves the overall accuracy and reliability of predictions, effectively avoiding prediction bias caused by treating all termites as a single research object.

[0130] This invention further proposes to pre-set probability thresholds and spatial clustering thresholds for each specific dietary group; when the distribution probability of a specific dietary group in the sub-group distribution probability map is continuously higher than the corresponding probability threshold in a certain continuous area, and the area or number of cells in that area reaches the spatial clustering threshold, an early warning is triggered; the early warning information includes the type of dietary group in the warning, the geographical coordinate range of the warning area, and the average occurrence probability in that area.

[0131] The probability threshold is dynamically adjusted based on the actual frequency of occurrence of the dietary group in historical data. For example, the probability threshold for wood-dwelling termite colonies is set at 0.75, and for soil-dwelling termites at 0.65. The spatial aggregation threshold is determined by statistically analyzing the minimum area of ​​colony aggregation in historical disaster events, such as setting it to 15 adjacent geographic cells or an area covering 500 square meters. The early warning triggering condition monitors the spatial aggregation status of the distribution probability layer in real time through a logic operation module. When both probability and area conditions are met, an early warning report containing geographic coordinate boundaries and probability intensity is automatically generated.

[0132] Specifically, after generating the probability map of the sub-population distribution, the system divides the geographical area into a standardized grid of cells. For each cell, the probability value of the occurrence of a specific dietary group is calculated, forming a probability distribution layer. The monitoring module continuously scans the layer data, identifying cell areas where the probability value exceeds a threshold. When these cells form a continuous spatial distribution and the total area reaches a preset threshold, it is determined to be a valid warning area. For example, when a colony of wood-dwelling termites is detected in a building complex area where 20 adjacent cells all have a probability value exceeding 0.8, and the total area reaches 1000 square meters, the system automatically generates a warning message containing the latitude and longitude range of the area and an average probability value of 0.85. This information is output in the form of map annotations and text reports, guiding pest control personnel to prioritize the inspection and extermination of wooden structures in the area.

[0133] As a preferred embodiment, the present invention is implemented as follows: For wood-dwelling termite colonies, the probability threshold is preset to 0.75, and the spatial aggregation threshold is defined as having at least 15 consecutive geographic cells. When the sub-colony distribution probability map shows that a certain area continuously covers 25 cells and the probability value of each cell is higher than 0.75, the system automatically determines that the area meets the warning conditions. The warning information generation module outputs the wood-dwelling termite colony type identifier, the geographic coordinate boundary of the area, and the average probability value of 0.82 within the area, and displays the warning area in the form of a heatmap overlay through the geographic information system platform.

[0134] Through the above technical solution, the present invention accurately identifies high-threat areas and triggers early warnings by dynamically judging the dual threshold conditions of the distribution probability and spatial aggregation degree of specific dietary groups. This effectively solves the problem of insufficient early warning accuracy caused by the failure to distinguish the differences in environmental response of dietary groups in the prior art, and significantly improves the efficiency of resource allocation for prevention and control of different termite groups.

[0135] The present invention further proposes a dynamic update step, which includes receiving new termite activity monitoring data, appending the new data to the sub-group activity dataset and the group-specific original feature dataset, triggering the retraining of the prediction model specific to the affected group according to the update strategy, and generating the updated model.

[0136] The newly added termite activity monitoring data includes new activity locations and their corresponding dietary group types, ensuring consistency between the data increment and the population classification system. The new data and its corresponding environmental factor data are appended to the corresponding dataset, maintaining data timeliness through incremental updates. The update strategy includes trigger mechanisms based on reaching a predetermined data increment or a preset time period, allowing for flexible configuration of model update conditions. During the retraining process, an independent validation dataset is introduced for performance evaluation. By comparing the performance metrics of the models before and after the update, it is determined whether to replace the current model with a candidate model, and the replacement operation and performance changes are recorded.

[0137] Specifically, when the amount of newly added monitoring data reaches a predetermined threshold, the system automatically initiates a retraining process. First, activity points and environmental factors associated with specific dietary groups are extracted from the new data and merged into the existing training set. Then, the merged dataset is standardized, and a candidate model is retrained using a pre-defined machine learning algorithm. Before model replacement, the accuracy, recall, and other performance metrics of the current model and the candidate model are tested using independent validation datasets. If the candidate model outperforms the current model, the model is replaced, and the replacement time and changes in performance metrics are recorded in the system log. Through periodic or data-driven update strategies, the predictive model continuously adapts to environmental changes and the evolution of termite colony behavior, improving the real-time performance and accuracy of distribution warnings.

[0138] As a preferred embodiment, the present invention is implemented as follows: The dynamic update step specifically includes the following processes: New termite activity monitoring data is received through sensor networks and manual inspection channels, wherein the location of new activity points is determined by GPS coordinates, and the dietary group type is determined by DNA sequencing of on-site samples. The new data is classified and appended to the corresponding sub-group activity dataset, while the associated environmental factor data is extracted and integrated into the group-specific original feature dataset through a geographic information system. The update strategy is set to trigger retraining when the data increment reaches a predetermined amount, wherein the predetermined amount is dynamically adjusted according to the historical data distribution; simultaneously, a timed trigger is executed on the 1st of each month to cope with periodic environmental changes. Before retraining, the accuracy and AUC value of the current group-specific prediction model on the validation dataset are evaluated. If the accuracy of the candidate model increases by more than 3% and the AUC value is not lower than that of the current model, model replacement is performed. The timestamp of the model replacement operation, changes in performance indicators, and operator information are recorded in the system log, forming a traceable update history.

[0139] Through the above technical solution, this invention achieves dynamic optimization of the termite distribution prediction model, ensuring that the model continuously adapts to environmental changes and new data features, and avoiding prediction bias caused by outdated data. The performance index comparison mechanism effectively prevents model degradation and ensures the reliability of early warning information. The complete record of the update process provides a data foundation for subsequent model iterations and problem tracing.

[0140] This invention further proposes the following steps: before retraining, evaluate the performance metrics of the current population-specific prediction model using an independent validation dataset; perform the retraining process to generate candidate updated prediction models; evaluate the performance metrics of the candidate updated prediction models using an independent validation dataset; compare the performance metrics of the current model and the candidate models, and replace the current model if the latter is better than the former; and record the model replacement operation, performance metric changes, and trigger timestamp information.

[0141] The independent validation dataset maintains consistent data distribution with the training dataset but does not overlap, ensuring the objectivity of the evaluation results. Performance metrics include accuracy, recall, and F1 score, which are quantified and compared using a weighted composite score. Model replacement operations must ensure that the candidate model's composite score is at least 5 percentage points higher than the current model to avoid frequent updates caused by minor optimizations. Timestamp information includes year, month, day, and data collection period, accurately recording the model iteration process.

[0142] Specifically, during the dynamic model update process, a 20% independent validation set is first defined from the historical data. Each time retraining is triggered, the accuracy and recall of the current model on the validation set are calculated and weighted to form a baseline score. After incremental training, the candidate model's performance metrics are recalculated on the same validation set, generating a comparative score. When the candidate model's score exceeds the baseline score by 5%, a model replacement operation is performed, and the performance difference between the old and new models, the replacement time, and the scale of the data increment are recorded in the system log. This mechanism ensures through dual verification that model updates only take effect when performance is improved, preventing model degradation caused by data noise or distribution shifts, while simultaneously recording the complete update trajectory to provide data support for subsequent optimizations.

[0143] As a preferred embodiment, the specific implementation of the present invention is as follows: In the dynamic update step, after the newly added monitoring data of wood-dwelling termite colony activities is added to the sub-colony activity dataset, 300 samples containing historical activity points and environmental factor data are randomly selected from the independent validation dataset. The AUC-ROC index of the current wood-dwelling colony-specific prediction model is evaluated, and the current model index is measured to be 0.82. Subsequently, the incremental gradient boosting algorithm is used to retrain the colony-specific training set to generate a candidate model, which is then tested on the same validation set. The AUC-ROC index of the candidate model is measured to be 0.87. After comparison confirms that the candidate model performs better than the current model, the candidate model is deployed as the new wood-dwelling colony prediction model. At the same time, the model version number is updated from V2.3 to V2.4, and the replacement time, performance index change, and operator identity information are recorded in the system log.

[0144] Through the above technical solutions, this invention effectively avoids the risk of model performance degradation caused by the distribution shift of newly added data, and ensures that the prediction accuracy is reliably improved after the model is updated through a verification mechanism. Meanwhile, the integrity of version control and operation records provides traceability support for the model iteration process, enabling the model optimization process to have a verifiable technical closed loop.

[0145] In summary, this invention addresses the problem of insufficient prediction accuracy in existing technologies caused by neglecting differences in environmental responses among termite populations by classifying termites according to their dietary groups and constructing a dedicated prediction model. It accurately captures the specific correlations between different populations and environmental factors, and the generated sub-population distribution probability map clearly displays the spatial distribution of each population, avoiding overlapping distribution characteristics. Based on this, the early warning information can provide precise distribution warnings for specific populations, offering differentiated guidance for prevention and control, and improving the targeting and efficiency of prevention and control. Simultaneously, the dynamic update mechanism retrains the model with new data, ensuring that the model continuously adapts to environmental changes and the evolution of termite behavior, avoiding prediction bias caused by outdated data, guaranteeing the reliability of early warning information, and reducing resource waste and ecological disturbance.

[0146] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A distribution early warning method based on a termite distribution inference model, applied to a target geographical area, characterized in that, The method includes the following steps: Acquire multi-source data associated with the target geographic area, including historical termite activity records, environmental factor data, and engineering structure parameters; Based on a pre-defined termite diet group classification system, termite samples from the historical termite activity records are assigned to one or more specific diet groups to generate subgroup activity datasets. For each of the one or more specific dietary groups, a group-specific prediction model is created; The steps for creating a population-specific prediction model include: For each specific dietary group, extract the corresponding activity point data from the subgroup activity dataset; By integrating the environmental factor data and the engineering structure parameters corresponding to the activity point data in the spatiotemporal dimension, a group-specific original feature dataset is constructed. The original feature dataset specific to the group is standardized to generate a group-specific training set for model training. Using a pre-defined machine learning algorithm, a group-specific prediction model is obtained by training on the group-specific training set. This model internalizes the specific association rules between the specific dietary group and environmental factors. The steps for constructing a group-specific original feature dataset also include: For each specific dietary group, an initial subset of environmental factors is selected from the environmental factor data based on a pre-defined group preference knowledge base; The completeness and redundancy of the initial environmental factor subset in characterizing the living environment of this specific dietary group are evaluated to obtain the evaluation results; If the evaluation results indicate insufficient completeness or excessive redundancy, the initial subset of environmental factors will be adjusted to form an optimized subset of environmental factors. The optimized environmental factor subset or the initial environmental factor subset is used as the environmental factor input for constructing the original feature dataset specific to the population. Based on the group-specific prediction model, a sub-group distribution probability map is generated to characterize the spatial distribution status of each specific dietary group within the target geographical area; Based on the group distribution probability map and the preset early warning triggering logic, distribution early warning information for specific dietary groups is generated and output.

2. The distribution early warning method based on the termite distribution inference model according to claim 1, characterized in that, The pre-defined termite diet group classification system is based on the termites' primary food sources, and the specific diet groups include: Wood-dwelling termite colonies primarily feed on woody structures and cellulose; Soil-dwelling termite colonies primarily feed on humus and organic matter in the soil; A type of termite colony that is both terrestrial and soil-dwelling, exhibiting feeding characteristics that combine terrestrial and soil-dwelling characteristics; In addition, there are termite colonies that cultivate fungi in their nests as a food source.

3. The distribution early warning method based on the termite distribution inference model according to claim 1, characterized in that, The environmental factor data includes: Climate data, including annual average temperature, monthly extreme temperature, annual average rainfall, and relative humidity data; Soil data, which includes soil type, soil pH, soil organic matter content, and soil moisture content; Vegetation data, which includes vegetation coverage index, vegetation type and surface litter layer thickness data; Topographic data, which includes elevation, slope and aspect data.

4. The distribution early warning method based on the termite distribution inference model according to claim 1, characterized in that, The steps for evaluating the completeness and redundancy of the initial subset of environmental factors in characterizing the living environment of the specific dietary group include: A feature selection algorithm is applied to calculate the correlation score between each environmental factor in the initial subset of environmental factors and the distribution of activity points of the specific dietary group; Calculate the mutual information score or correlation coefficient among the environmental factors in the initial subset of environmental factors; Based on preset scoring thresholds and mutual information thresholds, low-relevance factors and high-redundancy factors are identified, and the evaluation results are generated.

5. The distribution early warning method based on the termite distribution inference model according to claim 1, characterized in that, The step of generating a sub-group distribution probability map representing the spatial distribution status of each specific dietary group within the target geographical area includes: The target geographic region is gridded to generate a set of geographic cells; For each geographic cell in the set of geographic cells, extract its corresponding environmental factor data and engineering structure parameters to form a feature vector to be predicted. For each specific dietary group, the corresponding group-specific prediction model is applied to the feature vector to be predicted, and the probability value of the occurrence of the dietary group in the geographic cell is calculated. The probability values ​​of all geographic cells are combined to form a spatial distribution probability layer for this specific dietary group; The spatial distribution probability layers of all specific dietary groups are overlaid to form the subgroup distribution probability map.

6. The distribution early warning method based on the termite distribution inference model according to claim 1, characterized in that, The warning triggering logic includes: For each specific dietary group, a probability threshold and a spatial clustering threshold are pre-set; When the distribution probability of a specific dietary group in the subgroup distribution probability map is continuously higher than its corresponding probability threshold in a certain continuous area, and the area or number of cells in that area reaches the spatial clustering threshold, an early warning is triggered. The warning information includes the type of dietary group being warned, the geographical coordinates of the warning area, and the average probability of occurrence within that area.

7. The distribution early warning method based on the termite distribution inference model according to claim 1, characterized in that, The method also includes the following dynamic update steps: Receive newly added termite activity monitoring data, which includes the location of new activity points and their corresponding food group types; The newly added termite activity monitoring data and its corresponding environmental factor data are appended to the corresponding sub-group activity dataset and the original feature dataset specific to the group. According to a preset update strategy, a retraining process is triggered on one or more affected group-specific prediction models to generate an updated group-specific prediction model. The update strategies include: a trigger strategy based on data increments reaching a predetermined number, or a timed trigger strategy based on a preset time period.

8. The distribution early warning method based on the termite distribution inference model according to claim 7, characterized in that, The step of triggering the retraining process for one or more affected group-specific prediction models further includes: Before performing retraining, the performance metrics of the current population-specific prediction model are evaluated using an independent validation dataset. Perform the retraining process to generate a candidate updated prediction model; Using the independent validation dataset, evaluate the performance metrics of the candidate updated prediction models; Compare the performance metrics of the current population-specific prediction model with the performance metrics of the candidate updated prediction model. If the latter's performance metrics are better than the former's, then the candidate updated prediction model is used to replace the current population-specific prediction model. Record the operation of replacing the current population-specific prediction model with the updated prediction model of the candidate, the corresponding performance metric changes, and the timestamp information that triggered the operation.