Wild animal intelligent identification system based on protection area multi-mode perception fusion
By using a multimodal perception fusion intelligent identification system that combines acoustic and visual sensors, sentinel species can be identified in real time and predator-prey interaction events can be analyzed to construct a food chain stability index. This solves the problems of low data collection efficiency and insufficient ecosystem analysis in existing technologies, and enables efficient ecological risk assessment and strategy decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING LUXINYUAN TECHNOLOGY CO LTD
- Filing Date
- 2025-11-24
- Publication Date
- 2026-04-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing wildlife monitoring technologies suffer from low data collection efficiency, insufficient analysis depth, inability to respond to emergencies in real time, and lack of in-depth insight into the intrinsic relationships within the ecosystem, resulting in weak risk assessment and decision support capabilities.
An intelligent recognition system employing multimodal perception fusion combines a microphone array and infrared thermal imaging equipment to identify the acoustic characteristics of sentinel species in real time, generate acquisition commands, acquire multimodal datasets, perform target detection and behavior analysis, construct a food chain stability index, conduct dual risk assessments, and generate strategies.
It enables precise data collection, provides in-depth insights into interspecies interactions, offers early warnings of ecological imbalances, optimizes resource allocation and intervention strategies, and enhances monitoring efficiency and ecosystem health diagnostic capabilities.
Smart Images

Figure CN121902010A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent sensing technology, specifically to a wildlife intelligent identification system based on multimodal perception fusion in protected areas. Background Technology
[0002] Wildlife conservation is a core task in maintaining biodiversity and ecosystem balance. Traditional wildlife monitoring techniques, as an important support for conservation efforts, are increasingly revealing their inherent limitations in practice.
[0003] Current technologies generally face the dual bottlenecks of data acquisition efficiency and analytical depth. Large-scale deployments of infrared camera traps and wide-area acoustic sensors are currently the mainstream monitoring methods. However, infrared cameras employ a passive triggering mechanism, making them susceptible to interference from environmental factors (such as wind and grass), resulting in massive amounts of redundant image and video data. This not only places enormous pressure on data storage and transmission but also makes subsequent manual screening and analysis time-consuming and laborious, with long information feedback cycles, making it difficult to meet the real-time response needs for emergencies such as poaching and human-wildlife conflict. On the other hand, while acoustic monitoring alone can cover a wider area, it has inherent shortcomings in the accuracy of species identification, visual confirmation of individual behavior, and precise spatial localization of sound sources, making it impossible to construct a complete chain of event evidence.
[0004] Traditional monitoring systems are superficial in their data analysis, lacking in-depth insights into the intrinsic relationships within ecosystems. Most systems focus on species identification, counting, and the spatiotemporal distribution of their activities, outputting isolated, discrete data points. This analytical approach fails to effectively reveal complex interactions between species, such as predation, competition, and symbiosis. Consequently, these systems are highly insensitive to the health of food chains within a region. For example, abnormal behavior or population fluctuations of a key predator or prey population (i.e., sentinel species) could be an early warning sign of ecosystem imbalance. Similarly, animal foraging behavior directly reflects their response to environmental stress, and their foraging efficiency can proactively reflect potential predation threats, food shortages, or human disturbance.
[0005] In summary, existing monitoring systems are weak in risk assessment and decision support, failing to create an effective attention-guiding effect. This makes it difficult for managers to quickly identify and focus on the most urgent and important events from a sea of alerts, resulting in the dispersion and waste of management resources and the inability to achieve priority response and precise intervention for high-risk areas and high-threat events. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a multimodal perception fusion-based intelligent wildlife identification system for protected areas, thereby resolving the problems mentioned in the background section.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a multimodal perception fusion-based intelligent wildlife identification system for protected areas, comprising:
[0008] The intelligent sensing module is used to receive first modal data from the first sensing device; and analyze the first modal data in real time to identify the triggering events of the acoustic characteristics of the preset sentinel species, determine the spatial orientation information of the triggering events, generate a first acquisition command, and drive the second sensing device to execute the first acquisition command to obtain second modal data, and then align it with the first modal data to obtain a multimodal dataset.
[0009] The fusion analysis module is used to process multimodal datasets and uses target detection algorithms to identify all predator-prey interaction events in the current sensing unit within a preset time window, generate structured behavioral semantic labels, and calculate the instantaneous risk value of the event level for each predator-prey interaction event to construct a food chain stability index.
[0010] The dual risk assessment module is used to perform a weighted fusion calculation of the real-time risk value and the food chain stability index to obtain a risk level index, which is used to perform the first judgment and obtain the first assessment result.
[0011] Further analysis is performed on the second video data with structured behavioral semantic tags to calculate the foraging efficiency coefficient (FEC) of each predator-prey interaction event. Within a preset time window, the j-th perception region is aggregated, and the average value of the foraging efficiency coefficient (FEC) of all predator-prey interaction events is calculated. A second judgment is then performed to obtain the second evaluation result.
[0012] The decision-making module is used to combine the results of the first evaluation and the second evaluation to generate corresponding strategies.
[0013] Preferably, the intelligent sensing module includes a partitioning subunit and a deployment subunit;
[0014] Based on pre-set ecological and geographical data, the target wildlife reserve is divided into non-uniform grids to form several sensing areas.
[0015] The ecological and geographical data includes historical species activity heat maps, topographic data, and vegetation cover type data. The sensing areas identified by the ecological and geographical data are divided into at least two priority levels: core sensing areas and regular sensing areas. For core sensing areas identified as water sources, high-density activity hotspots, or habitats of rare species, the single-sided size of the core sensing area is set within the range of 50 to 150 meters. For regular sensing areas, the single-sided size of the regular sensing area is set within the range of 500 to 1500 meters.
[0016] A deployment subunit is used to deploy a first sensing device and a second sensing device in each sensing area, and to establish a collaborative communication link between the first sensing device and the second sensing device.
[0017] The first sensing device is a microphone array;
[0018] The second sensing device integrates infrared thermal imaging equipment and high-definition camera equipment.
[0019] Preferably, the intelligent sensing module also includes an edge analysis unit;
[0020] An edge analysis unit is used to process the first modal data collected by the first sensing device in real time and drive the second sensing device to respond in a coordinated manner, including:
[0021] The acoustic feature extraction subunit is used to perform preprocessing on the received first modal data. The preprocessing includes filtering out environmental background noise outside a preset frequency band using a bandpass filtering algorithm; and framing and windowing the preprocessed first modal data, and extracting the Mel-frequency cepstral coefficients (MFCCs) of each frame as acoustic feature vectors.
[0022] The species voiceprint recognition subunit is used to embed a pre-trained lightweight convolutional neural network model, input the acoustic feature vector into the lightweight convolutional neural network model to calculate the matching confidence score between the first modality data and one or more preset sentinel species voiceprint models; and when any matching confidence score is higher than the preset species recognition confidence threshold, the corresponding sentinel species acoustic feature is confirmed to be identified, and the category code of the corresponding sentinel species is output.
[0023] The sound source localization subunit is used to calculate the time delay of the first modal data arriving at each microphone channel based on the first modal data acquired from the microphone array of the first sensing device, using the Time-of-Arrival (TDOA) algorithm; and to calculate the azimuth and elevation angles representing the direction of the sound source based on the time delay and the geometric layout information of the microphone array, as spatial orientation information, and generate a first acquisition command; the first acquisition command includes: an angle indicating that the second sensing device needs to rotate in the horizontal direction, i.e., the azimuth angle;
[0024] Indicates the angle that the second sensing device needs to be adjusted in the vertical direction, i.e., the pitch angle;
[0025] The sentinel species is locked and continuously tracked, with at least two still photos taken and a 30-second video recorded to obtain corresponding first still image and second video data.
[0026] Preferably, the intelligent sensing module also includes a directional acquisition subunit and a spatiotemporal alignment subunit;
[0027] The directional acquisition subunit is used to receive the first acquisition command and activate the second sensing device accordingly.
[0028] The second sensing device is controlled to execute a directional acquisition program based on spatial orientation information to obtain a first static image and a second video data for the sentinel species as second modal data;
[0029] After the second sensing device completes the turning and positioning of spatial orientation information, it collects the ambient brightness. When the ambient brightness is lower than the preset brightness threshold, it automatically switches to infrared thermal imaging mode; otherwise, it maintains the high-definition camera mode.
[0030] The spatiotemporal alignment subunit is used to associate the first modality data with the second modality data to form a multimodal dataset.
[0031] Preferably, the fusion analysis module includes an image analysis unit, a behavior classification sub-unit, and an event association determination unit;
[0032] The image analysis unit is used to perform a target detection algorithm on consecutive frames in the first static image and the second video data to identify and locate the species to which one or more sentinel species belong, and associate the spatiotemporal location information of the i-th sentinel species.
[0033] The behavior classification subunit is used to perform pose analysis through image processing for each identified sentinel species to obtain the skeletal keypoint coordinate sequence of the i-th sentinel species at consecutive time points; and, based on the skeletal keypoint coordinate sequence, to calculate and generate the dynamic pose sequence of the i-th sentinel species, inputting it into a pre-built behavior recognition model to calculate the matching probability score between the dynamic pose sequence of the i-th sentinel species and multiple preset behavior types, and selecting the behavior type with the highest matching probability score as the output; the behavior types include: chirping, stalking, chasing, drinking, or alerting;
[0034] The identified species identity, spatiotemporal location information, and behavior types output by the behavior classification subunit are aggregated, and the aggregated information is formatted into a group of sentinel species.
[0035] When at least two groups of sentinel species are received, calculate their spatiotemporal proximity.
[0036] Based on two groups of sentinel species, a pre-defined ecological knowledge graph is queried to obtain the ecological relationship between the two groups of sentinel species.
[0037] By combining spatial and temporal proximity with the ecological relationships between the two groups of sentinel species, it was determined whether a predator-prey interaction event existed, including:
[0038] The rules engine is used to define the criteria for classifying a biological event as a predator-prey interaction event.
[0039] Within a preset spatiotemporal threshold, continuously identify cross-modal joint features of at least two sentinel species belonging to different species;
[0040] A predefined predator-prey relationship was confirmed between the two groups of sentinel species;
[0041] When the spatiotemporal proximity is less than or equal to a preset proximity threshold, and it is confirmed that there is a predefined predator-prey relationship between the two groups of sentinel species, a structured behavioral semantic label representing the predator-prey interaction event is generated.
[0042] Preferably, a preset ecological knowledge graph is queried to obtain the corresponding protection level, and a basic score is determined based on the protection level; based on the event type, an event weighting coefficient is determined from a preset event risk weight library; the basic score and the event weighting coefficient are multiplied by a preset operation to generate an instantaneous risk value for the event level;
[0043] Within a preset time window, aggregate multimodal datasets and calculate population activity indicators;
[0044] Based on the ecological knowledge graph, the species associated with the upstream and downstream of the food chain and the preset baseline population activity ratio are retrieved; the current population activity index is compared with the baseline population activity ratio to calculate the deviation, and the food chain stability index is obtained by correlation calculation.
[0045] Preferably, the dual risk assessment module includes a risk level fusion unit and a first assessment unit;
[0046] The risk level fusion unit is used to perform a weighted fusion calculation between the real-time risk value and the food chain stability index to obtain the risk level index.
[0047] The first assessment unit is used to preset a first risk threshold and a second risk threshold, wherein the first risk threshold is greater than the second risk threshold, and to assess the risk level index to obtain a first assessment result, including:
[0048] When the risk level index is higher than the first risk threshold, a first high-risk label is generated.
[0049] When the risk level index is between the first risk threshold and the second risk threshold, a second risk label is generated;
[0050] When the risk level index is less than the second risk threshold, a third low-risk label is generated.
[0051] Preferably, the dual risk assessment module further includes a second analysis unit, which is used to further analyze the second video data that generates structured behavioral semantic tags representing predator-prey interaction events, and to extract a second video data segment of a preset duration as the foraging cycle for the target analysis based on the timestamps recorded in the structured behavioral semantic tags.
[0052] During the current foraging cycle, the target animal entity in the second video data is continuously tracked by skeletal key points to obtain the three-dimensional spatial coordinate time series of at least one head key point, one shoulder key point, and one body centroid key point.
[0053] By traversing each frame of the three-dimensional spatial coordinate time series, and based on the relative spatial relationship and displacement velocity of key points, the animal's foraging behavior is dynamically and mutually exclusively classified into one of the following two core microstates;
[0054] The current frame is determined to belong to the ingested micro-state if and only if the vertical coordinate of the head key point is lower than the vertical coordinate of the shoulder key point, and the displacement of the body center of mass key point in a unit time is less than the first preset velocity threshold.
[0055] When the vertical coordinates of the head key points are not lower than the vertical coordinates of the shoulder key points, the current frame is determined to be in the alert and search micro-state.
[0056] Throughout the foraging cycle, extract the total intake time (Tingestion), total alert and search time (Tvigilance), search movement cost (Dsearch), and number of behavior interruptions (Ninterrupt).
[0057] The weighted total vigilance and search duration (Tvigilance), search movement cost (Dsearch), and number of behavior interruptions (Ninterrupt) are summed and added to the total intake duration (Tingestion) and placed in the denominator. The total intake duration (Tingestion) is placed in the numerator as a benefit term. The ratio of the numerator to the denominator is obtained to obtain the foraging efficiency coefficient, denoted as FEC.
[0058] Preferably, the dual risk assessment module further includes a second assessment unit, which is used to aggregate the j-th perception area within a preset time window and calculate the average value of the foraging efficiency coefficient (FEC) of several predator-prey interaction events.
[0059] An efficiency threshold is preset, and the average value of the foraging efficiency coefficient (FEC) is evaluated to obtain a second evaluation result, including: when the average value of the foraging efficiency coefficient (FEC) is less than the efficiency threshold, it indicates that there is a potential stress risk in the current sensing area, and a first stress risk label is generated; it indicates that animals in the current sensing area generally exhibit inefficient foraging behavior, and the time and energy costs far exceed the benefits, suggesting that there may be a high intensity of predation threat, severe food shortage, or frequent external disturbances in the current sensing area;
[0060] When the average value of the foraging efficiency coefficient (FEC) is greater than or equal to the efficiency threshold, a second qualified label is generated, indicating that the potential stress risk in the current sensing area is within the expected range.
[0061] Preferably, the decision-making module is used to combine the first evaluation result and the second evaluation result to generate a corresponding strategy, including:
[0062] When both the first high-risk label and the first stress risk label are identified simultaneously, a first strategy is generated, including: designating the current sensing area as the first priority recovery area, selecting at least 4 to 5 priority recovery areas on the map with a total area not less than 16% to 20% of the total area of the j-th sensing area, planting restoration plants 5m to 10m away from water sources; and establishing at least 10 to 15 predation risk buffer zones.
[0063] When the second risk label and the first stress risk label are identified at the same time, a second strategy is generated, including: taking the current sensing area as the second priority recovery area, establishing at least 2 to 3 priority recovery areas on the map with a total area not less than 10% of the total area of the j-th sensing area, and establishing 6 to 10 predation risk buffer zones between the foraging area and the woodland.
[0064] When the first high-risk label and the second qualified label are identified at the same time, a third strategy is generated, including: taking the current sensing area as the third priority recovery area, and having at least 4 to 5 priority recovery areas on the map with a total area not less than 16% to 20% of the total area of the jth sensing area.
[0065] When the second risk label and the second qualified label are identified at the same time, a fourth strategy is generated, including: taking the current sensing area as the fourth priority recovery area, and having at least 2 to 3 priority recovery areas on the map with a total area not less than 10% of the total area of the jth sensing area.
[0066] When the third low-risk label and the first stress risk label are identified at the same time, a fifth strategy is generated, including: taking the current perception area as the fifth priority recovery area, and establishing 3 to 5 predation risk buffer zones between the foraging area and the woodland, with the width of the predation risk buffer zones set to 3 to 5 meters.
[0067] When the third low-risk label and the second qualified label are identified simultaneously, no intervention is made in the current sensing area, and continuous monitoring is carried out.
[0068] This invention provides a multimodal perception fusion-based intelligent wildlife identification system for protected areas. It offers the following advantages:
[0069] (1) By using an acoustic sentinel triggering mechanism, passive recording is transformed into active and precise data collection. It can accurately locate high-value biological events from complex environmental backgrounds, instruct visual devices to capture them in a directional manner, greatly reducing data redundancy and subsequent manual analysis costs, and improving monitoring efficiency and real-time response capabilities to emergencies.
[0070] (2) By combining multimodal fusion analysis with ecological knowledge graphs, discrete species data are transformed into structured behavioral semantic tags. It deeply reveals the predation and other interactive relationships between species and constructs a food chain stability index, achieving a qualitative leap from simple species existence statistics to the diagnosis of the intrinsic health of ecosystems.
[0071] (3) The foraging efficiency coefficient (FEC) was introduced as a micro-behavioral health indicator. It can precisely quantify the environmental stress experienced by animals due to predation threats or food shortages, providing an early warning capability that traditional monitoring techniques cannot achieve, and can proactively and more sensitively reflect potential ecosystem imbalance risks.
[0072] (4) A “dual risk assessment” and “tiered strategy decision-making” system was constructed, which transformed the complex ecological situation into a clear risk level and a precise intervention strategy. This created an effective attention-guiding effect, helping managers to quickly identify and prioritize the events and areas with the highest threat from a large amount of information, thereby achieving the optimal allocation and efficient utilization of conservation resources. Attached Figure Description
[0073] Figure 1 This is a schematic diagram of the intelligent sensing module of the present invention;
[0074] Figure 2 This is a schematic diagram of the fusion analysis module of the present invention;
[0075] Figure 3 This is a flowchart illustrating the dual risk assessment module and decision-making module of the present invention.
[0076] Figure 4 This is a schematic diagram of the data example two of the present invention. Detailed Implementation
[0077] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0078] Example 1
[0079] Please see Figures 1 to 4 This invention provides a multimodal perception fusion-based intelligent wildlife identification system for protected areas, comprising:
[0080] The intelligent sensing module is used to receive first modal data from the first sensing device; and analyze the first modal data in real time to identify the triggering events of the acoustic characteristics of the preset sentinel species, determine the spatial orientation information of the triggering events, generate a first acquisition command, and drive the second sensing device to execute the first acquisition command to obtain second modal data, and then align it with the first modal data to obtain a multimodal dataset.
[0081] The fusion analysis module is used to process multimodal datasets and uses target detection algorithms to identify all predator-prey interaction events in the current sensing unit within a preset time window, generate structured behavioral semantic labels, and calculate the instantaneous risk value of the event level for each predator-prey interaction event to construct a food chain stability index.
[0082] The dual risk assessment module is used to perform a weighted fusion calculation of the real-time risk value and the food chain stability index to obtain a risk level index, which is used to perform the first judgment and obtain the first assessment result.
[0083] Further analysis is performed on the second video data with structured behavioral semantic tags to calculate the foraging efficiency coefficient (FEC) of each predator-prey interaction event. Within a preset time window, the j-th perception region is aggregated, and the average value of the foraging efficiency coefficient (FEC) of all predator-prey interaction events is calculated. A second judgment is then performed to obtain the second evaluation result.
[0084] The decision-making module is used to combine the results of the first evaluation and the second evaluation to generate corresponding strategies.
[0085] Preferably, the intelligent sensing module includes a partitioning subunit and a deployment subunit;
[0086] Based on pre-set ecological and geographical data, the wildlife reserve is divided into non-uniform grids to form several sensing areas.
[0087] The ecological and geographical data includes historical species activity heat maps, topographic data, and vegetation cover type data. The sensing areas identified by the ecological and geographical data are divided into at least two priority levels: core sensing areas and regular sensing areas. For core sensing areas identified as water sources, high-density activity hotspots, or habitats of rare species, the single-sided size of the core sensing area is set within the range of 50 to 150 meters. For regular sensing areas, the single-sided size of the regular sensing area is set within the range of 500 to 1500 meters.
[0088] A deployment subunit is used to deploy a first sensing device and a second sensing device in each sensing area, and to establish a collaborative communication link between the first sensing device and the second sensing device.
[0089] The first sensing device is a microphone array;
[0090] The second sensing device integrates infrared thermal imaging equipment and high-definition camera equipment.
[0091] Preferably, the intelligent sensing module also includes an edge analysis unit;
[0092] An edge analysis unit is used to process the first modal data collected by the first sensing device in real time and drive the second sensing device to respond in a coordinated manner, including:
[0093] The acoustic feature extraction subunit is used to perform preprocessing on the received first modal data. The preprocessing includes filtering out environmental background noise outside a preset frequency band using a bandpass filtering algorithm; and framing and windowing the preprocessed first modal data, and extracting the Mel-frequency cepstral coefficients (MFCCs) of each frame as acoustic feature vectors.
[0094] The species voiceprint recognition subunit is used to embed a pre-trained lightweight convolutional neural network model, input the acoustic feature vector into the lightweight convolutional neural network model to calculate the matching confidence score between the first modality data and one or more preset sentinel species voiceprint models; and when any matching confidence score is higher than the preset species recognition confidence threshold, the corresponding sentinel species acoustic feature is confirmed to be identified, and the category code of the corresponding sentinel species is output.
[0095] The sound source localization subunit is used to calculate the time delay of the first modal data arriving at each microphone channel based on the first modal data acquired from the microphone array of the first sensing device, using the Time-of-Arrival (TDOA) algorithm; and to calculate the azimuth and elevation angles representing the direction of the sound source based on the time delay and the geometric layout information of the microphone array, as spatial orientation information, and generate a first acquisition command; the first acquisition command includes: an angle indicating that the second sensing device needs to rotate in the horizontal direction, i.e., the azimuth angle;
[0096] Indicates the angle that the second sensing device needs to be adjusted in the vertical direction, i.e., the pitch angle;
[0097] The sentinel species is locked and continuously tracked, with at least two still photos taken and a 30-second video recorded to obtain corresponding first still image and second video data.
[0098] Preferably, the intelligent sensing module also includes a directional acquisition subunit and a spatiotemporal alignment subunit;
[0099] The directional acquisition subunit is used to receive the first acquisition command and activate the second sensing device accordingly.
[0100] The second sensing device is controlled to execute a directional acquisition program based on spatial orientation information to obtain a first static image and a second video data for the sentinel species as second modal data;
[0101] After the second sensing device completes the turning and positioning of spatial orientation information, it collects the ambient brightness. When the ambient brightness is lower than the preset brightness threshold, it automatically switches to infrared thermal imaging mode; otherwise, it maintains the high-definition camera mode.
[0102] The spatiotemporal alignment subunit is used to associate the first modality data with the second modality data to form a multimodal dataset.
[0103] Preferably, the fusion analysis module includes an image analysis unit, a behavior classification sub-unit, and an event association determination unit;
[0104] The image analysis unit is used to perform a target detection algorithm on consecutive frames in the first static image and the second video data to identify and locate the species to which one or more sentinel species belong, and associate the spatiotemporal location information of the i-th sentinel species.
[0105] The behavior classification subunit is used to perform pose analysis through image processing for each identified sentinel species to obtain the skeletal keypoint coordinate sequence of the i-th sentinel species at consecutive time points; and, based on the skeletal keypoint coordinate sequence, to calculate and generate the dynamic pose sequence of the i-th sentinel species, inputting it into a pre-built behavior recognition model to calculate the matching probability score between the dynamic pose sequence of the i-th sentinel species and multiple preset behavior types, and selecting the behavior type with the highest matching probability score as the output; the behavior types include: chirping, stalking, chasing, drinking, or alerting;
[0106] The identified species identity, spatiotemporal location information, and behavior types output by the behavior classification subunit are aggregated, and the aggregated information is formatted into a group of sentinel species.
[0107] When at least two groups of sentinel species are received, calculate their spatiotemporal proximity.
[0108] Based on two groups of sentinel species, a pre-defined ecological knowledge graph is queried to obtain the ecological relationship between the two groups of sentinel species.
[0109] By combining spatial and temporal proximity with the ecological relationships between the two groups of sentinel species, it was determined whether a predator-prey interaction event existed, including:
[0110] The rules engine is used to define the criteria for classifying a biological event as a predator-prey interaction event.
[0111] The goal is to continuously identify cross-modal joint features of at least two sentinel species belonging to different species within a preset spatiotemporal threshold. The "spatiotemporal threshold" describes the concept of a "boundary range." In the phrase "within a preset spatiotemporal threshold...", it means "within a preset temporal and spatial boundary range...". Here, the "spatiotemporal threshold" is an abstract concept.
[0112] A predefined predator-prey relationship was confirmed between the two groups of sentinel species;
[0113] When the spatiotemporal proximity is less than or equal to a preset proximity threshold, and a predefined predator-prey relationship is confirmed between two groups of sentinel species, a structured behavioral semantic label representing the predator-prey interaction event is generated. The proximity threshold is a specific constant used for comparison; preferably, a preset ecological knowledge graph is queried to obtain the corresponding protection level, and a base score is determined based on the protection level; based on the event type, an event weighting coefficient is determined from a preset event risk weight library; the base score and the event weighting coefficient are multiplied using a preset multiplication operation to generate an instantaneous risk value for the event level;
[0114] Within a preset time window, aggregate multimodal datasets and calculate population activity indicators;
[0115] Based on the ecological knowledge graph, the species associated with the upstream and downstream of the food chain and the preset baseline population activity ratio are retrieved; the current population activity index is compared with the baseline population activity ratio to calculate the deviation, and the food chain stability index is obtained by correlation calculation.
[0116] Preferably, the dual risk assessment module includes a risk level fusion unit and a first assessment unit;
[0117] The risk level fusion unit is used to perform a weighted fusion calculation of the immediate risk value and the food chain stability index to obtain the risk level index. The immediate risk value reflects short-term threats, while the food chain stability index reflects long-term health trends.
[0118] The first assessment unit is used to preset a first risk threshold and a second risk threshold, wherein the first risk threshold is greater than the second risk threshold, and to assess the risk level index to obtain a first assessment result, including:
[0119] When the risk level index is higher than the first risk threshold, the first high-risk label is generated, indicating that there is an urgent and serious ecological threat in the current sensing area;
[0120] When the risk level index is between the first risk threshold and the second risk threshold, a second risk label is generated, indicating that the current sensing area represents the existence of potential or developing ecological risks.
[0121] When the risk level index is less than the second risk threshold, a third low-risk label is generated, indicating that the current ecological status of the perceived area is stable.
[0122] The dual risk assessment module also includes a second analysis unit, which is used to further analyze the second video data that generates structured behavioral semantic tags representing predator-prey interaction events. Based on the timestamps recorded in the structured behavioral semantic tags, the second video data segments of a preset duration are extracted as the foraging cycle for the target analysis.
[0123] During the current foraging cycle, the target animal entity in the second video data is continuously tracked by skeletal key points to obtain the three-dimensional spatial coordinate time series of at least one head key point, one shoulder key point, and one body centroid key point.
[0124] By traversing each frame of the three-dimensional spatial coordinate time series, and based on the relative spatial relationship and displacement velocity of key points, the animal's foraging behavior is dynamically and mutually exclusively classified into one of the following two core microstates;
[0125] The current frame is determined to belong to the ingested micro-state if and only if the vertical coordinate of the head key point is lower than the vertical coordinate of the shoulder key point, and the displacement of the body center of mass key point in a unit time is less than the first preset velocity threshold.
[0126] When the vertical coordinates of the head key points are not lower than the vertical coordinates of the shoulder key points, the current frame is determined to be in the alert and search micro-state.
[0127] Throughout the foraging cycle, extract the total intake time (Tingestion), total alert and search time (Tvigilance), search movement cost (Dsearch), and number of behavior interruptions (Ninterrupt).
[0128] The total intake duration (Tingestion) is obtained by summing up the total duration of all frames that are identified as "intake microstates".
[0129] The total alert and search duration (Tvigilance) is obtained by summing the total duration of all frames that are identified as being in alert and search micro-states.
[0130] The search movement cost Dsearch is obtained by accumulating the total spatial displacement distance of the body's center of mass key point only during the alert and search micro-states.
[0131] The method for obtaining Ninterrupt behavior interruption is to count the total number of times the behavior microstate switches from ingestion microstate to alert and search during the foraging cycle.
[0132] The weighted total vigilance and search duration (Tvigilance), search movement cost (Dsearch), and number of behavior interruptions (Ninterrupt) are summed and added to the total intake duration (Tingestion) and placed in the denominator. The total intake duration (Tingestion) is placed in the numerator as a benefit term. The ratio of the numerator to the denominator is obtained to obtain the foraging efficiency coefficient, denoted as FEC.
[0133] Preferably, the dual risk assessment module further includes a second assessment unit, which is used to aggregate the j-th perception area within a preset time window and calculate the average value of the foraging efficiency coefficient (FEC) of several predator-prey interaction events.
[0134] An efficiency threshold is preset, and the average value of the foraging efficiency coefficient (FEC) is evaluated to obtain a second evaluation result, including: when the average value of the foraging efficiency coefficient (FEC) is less than the efficiency threshold, it indicates that there is a potential stress risk in the current sensing area, and a first stress risk label is generated; it indicates that animals in the current sensing area generally exhibit inefficient foraging behavior, and the time and energy costs far exceed the benefits, suggesting that there may be a high intensity of predation threat, severe food shortage, or frequent external disturbances in the current sensing area;
[0135] When the average value of the foraging efficiency coefficient (FEC) is greater than or equal to the efficiency threshold, a second qualified label is generated, indicating that the potential stress risk in the current sensing area is within the expected range.
[0136] Preferably, the decision-making module is used to combine the first evaluation result and the second evaluation result to generate a corresponding strategy, including:
[0137] When both the first high-risk label and the first stress risk label are identified simultaneously, a first strategy is generated, including: designating the current sensing area as the first priority recovery area, selecting at least 4 to 5 priority recovery areas on the map with a total area not less than 16% to 20% of the total area of the j-th sensing area, planting restoration plants 5m to 10m away from water sources; and establishing at least 10 to 15 predation risk buffer zones.
[0138] The most critical situation. The foundation of the ecosystem has been severely shaken (high risk), and individual animals are facing enormous and direct threats to their survival (stress risk). The strongest intervention is needed, combining emergency relief with fundamental reinforcement. Large-scale habitat restoration aims to save the collapsing ecosystem foundation; while large-scale buffer zone construction provides immediate refuge for frightened animals, reducing mortality.
[0139] When the second risk label and the first stress risk label are identified at the same time, a second strategy is generated, including: taking the current sensing area as the second priority recovery area, establishing at least 2 to 3 priority recovery areas on the map with a total area not less than 10% of the total area of the j-th sensing area, and establishing 6 to 10 predation risk buffer zones between the foraging area and the woodland.
[0140] The second strategy (medium risk + stress risk) involves preventative interventions that combine "corrective" and "stress-reducing" measures. Medium-scale habitat restoration aims to reverse the trend of system deterioration, while medium-scale buffer zone construction is used to alleviate current individual stress and prevent the problem from escalating to the critical state described in the first strategy.
[0141] When both the first high-risk label and the second qualified label are identified simultaneously, a third strategy is generated, including: designating the current sensing area as the third priority recovery zone, and establishing at least four to five priority recovery zones on the map with a total area not less than 16% to 20% of the total area of the j-th sensing area. This indicates that the problem is likely not due to "predation pressure," but rather stems from extreme food shortages, disease, or large-scale habitat degradation. Therefore, the strategy precisely addresses the root cause by undertaking large-scale habitat restoration to solve the fundamental problem, rather than investing resources in establishing predation buffer zones that are not currently urgently needed, demonstrating the precision of resource utilization.
[0142] When both the second risk label and the second qualifying label are identified simultaneously, a fourth strategy is generated, which includes: designating the current sensing area as the fourth priority restoration zone, and establishing at least two to three priority restoration zones on the map with a total area not less than 10% of the total area of the j-th sensing area; similar to the logic of the third strategy, this is an early signal of systemic problems. The strategy focuses on medium-scale habitat restoration, conducting early interventions to prevent the problem from worsening, and also avoids unnecessary buffer zone construction.
[0143] When both the third low-risk tag and the first stress risk tag are simultaneously identified, a fifth strategy is generated, which includes: designating the current sensing area as the fifth priority restoration zone, and establishing 3 to 5 predation risk buffer zones between the foraging area and the woodland, with the width of the predation risk buffer zones set at 3 to 5 meters; this reflects the "high sensitivity" of this system. Since the system's foundation is intact, large-scale habitat restoration is not necessary. The strategy establishes only a small number of key predation risk buffer zones to address this localized, nascent problem at the lowest cost and fastest speed.
[0144] When both the third low-risk label and the second qualified label are identified simultaneously, no intervention is made in the currently perceived area; continuous monitoring is conducted instead. Any intervention would be a disturbance to nature. Once the ecosystem is confirmed to be healthy, continuous monitoring is sufficient. This approach is both a scientific conservation principle and avoids unnecessary waste of resources.
[0145] Among these measures, restoring vegetation and improving water sources can directly enhance food abundance and environmental carrying capacity within the region, representing a fundamental solution to long-term, basic problems. Establishing predation risk buffer zones is an immediate solution to the problem of "individual behavioral stress." By creating densely vegetated "safe passages" or "refuges" in key areas (such as between foraging areas and woodlands), the risk of predation for animals can be reduced, alleviating their psychological stress.
[0146] The specific implementation designates a circular monitoring area with a radius ranging from 75 to 150 meters, centered on the core sensing area. This aims to ensure complete coverage of the core area while also accommodating the effective analysis range of visual sensing devices for small targets.
[0147] Data Example 1:
[0148] The scenario is set in a wetland environment (grid number: core area-07) designated as a core sensing area. This area is a typical habitat for various frog species (sentinel species, such as the marsh frog) and small to medium-sized predators (sentinel species, such as the weasel). Within this area, at least one monitoring node is deployed using a non-uniform grid partitioning scheme, consisting of a first sensing device (a high-sensitivity microphone array, number: acoustic sensor-07) and a second sensing device (a pan-tilt unit integrating infrared thermal imaging and high-definition camera equipment, number: image sensor-07).
[0149] The system's default parameters are as follows:
[0150] The lightweight convolutional neural network model built into the edge analysis unit of the Acoustic Sensor-07 has been pre-trained to include the mating calls of the swamp frog (species category code: frog-01) and the lurking sounds of the weasel (species category code: weasel-03).
[0151] Species identification confidence threshold: 0.90.
[0152] Ecological knowledge graph: stores the predator-prey relationship between "weasel-03" and "frog-01".
[0153] Risk level index thresholds: First risk threshold = 0.7, second risk threshold = 0.4.
[0154] Foraging efficiency coefficient and efficiency threshold: 0.5.
[0155] Step 1: Acoustic-triggered multimodal data acquisition;
[0156] At time T0, the intelligent sensing module of Acoustic Sensor-07 detects a continuous croaking sound, and the edge analysis unit initiates the processing flow.
[0157] Acoustic Feature Extraction and Recognition:
[0158] The acoustic feature extraction subunit performs preprocessing on the received first modal data (audio), performs framing and windowing after bandpass filtering (400 Hz-2000 Hz), and extracts the Mel frequency cepstral coefficients as acoustic feature vectors.
[0159] The species voiceprint recognition subunit inputs the feature vector into the model and calculates a matching confidence score of 0.96 for "swamp frog" (frog-01). Since this score is higher than the preset threshold of 0.90, the sentinel species "swamp frog" is confirmed and its category code is output.
[0160] Sound source localization and orientation acquisition: The sound source localization subunit uses the time difference of arrival algorithm to calculate the spatial orientation information based on the time delay of the microphone array (e.g., 0.18 ms and 0.35 ms): azimuth angle 85 degrees, pitch angle -15 degrees.
[0161] The first acquisition command is then generated, which includes the target azimuth angle of 85 degrees, the target elevation angle of -15 degrees, and parameters for continuous acquisition for 30 seconds.
[0162] The directional acquisition subunit drives the gimbal of the acoustic sensor-07 to precisely rotate to the designated position, lock onto the target, and acquire a first still image and a 30-second second video clip. At this time, the ambient light is sufficient, maintaining high-definition camera mode.
[0163] The spatiotemporal alignment subunit associates the audio and video data to form a multimodal dataset with a unified spatiotemporal label (core area -07, T0).
[0164] Step 2: Analysis of predator-prey interaction events based on multimodal fusion;
[0165] At time T1 (90 seconds after T0), acoustic sensor-07 again detected a faint low-frequency abnormal noise originating from a similar location. The system repeated step one, identified the "weasel" (Weasel-03), and acquired image data of its stalking. The fusion analysis module began processing these two spatiotemporally continuous multimodal datasets.
[0166] Image analysis and behavior classification:
[0167] The image analysis unit performs target detection on the image data of the swamp frog and the yellow weasel to identify and locate the species.
[0168] The behavior classification subunit analyzed the posture of the swamp frog and identified its behavior type as "calling"; it analyzed the posture of the weasel and identified its behavior type as "lurking".
[0169] Event correlation determination:
[0170] The event association determination unit aggregates two sets of information: one set is the croaking behavior and location of the swamp frog at time T0, and the other set is the stealthy behavior and location of the weasel at time T1.
[0171] Time Dimension: In the example, the time difference T1-T0 = 90 seconds was successfully identified as a single continuous event. This indicates that the system's time tolerance is at least 90 seconds. Considering the predator's (weasel's) stealth and ambush behavior, the complete interaction process could last several minutes. Therefore, the spatiotemporal threshold is set between 180 and 300 seconds (3 to 5 minutes). Beyond this time, the correlation between the two events decreases.
[0172] Calculations determined that the time difference was 90 seconds and the spatial distance was less than 5 meters. Based on the spatiotemporal difference, a normalized spatiotemporal proximity was calculated. First, the spatiotemporal proximity was less than a preset proximity threshold. Second, an ecological knowledge graph was consulted to confirm that there was a predator-prey relationship between the two.
[0173] Spatiotemporal proximity, denoted as Pst, is calculated using the following preferred embodiment: Pst = √[Wt × (ΔT / Tmax)² + Wd × (Δd / dmax)²]; ΔT represents the actual time difference, Tmax represents the maximum tolerable time difference, Δd represents the actual spatial distance, and dmax represents the spatial normalized baseline value; Wt and Wd are the time weighting factor and the spatial weighting factor, respectively.
[0174] Consider a boundary case: two events occur at the edge of the time tolerance (e.g., ΔT = 280 seconds) and the edge of the spatial tolerance (e.g., Δd = 45 meters), respectively.
[0175] Substituting into the formula, we get: P_st=√[0.5×(280 / 300)²+0.5×(45 / 50)²]≈√[0.5×0.87+0.5×0.81]≈0.91.
[0176] The preferred range for the proximity threshold is [0.60, 0.80], with a particularly preferred value of 0.70; 0.91 is greater than 0.70. Therefore, the system will correctly reject the classification of these two weakly correlated events as a predator-prey interaction, thereby helping to avoid false alarms.
[0177] Based on the judgment criteria of the rule engine, the system generates a structured behavioral semantic label representing the predator-prey interaction event. This label records the event type as "predator-prey interaction", the predator as "weasel-03", the prey as "frog-01", and the timestamp T1 of the event.
[0178] Step 3: Dual Risk Assessment and Decision Generation;
[0179] The dual risk assessment module is activated, and two judgments are made.
[0180] First level of assessment (macro-ecological risk):
[0181] Instant risk value calculation: The protection level of the species "Frog-01" involved in this "predator-prey interaction event" is mapped to a base score of 0.7. Combined with the event weighting coefficient of 1.2 for the event type (predator), 0.7 × 1.2 generates an instant risk value of 0.84.
[0182] The first example of the formula for calculating the food chain stability index is the weighted summation method.
[0183] This formula is used to assess the long-term, macro-level ecological health trends of a region over a period of time (e.g., the past 30 days). Calculation formula: FCSI = Σ(|D i |×W i );
[0184] The food chain stability index, the final output result, has a value range of [0,1]. Σ: Summation symbol, indicating that the calculation results for all key species within the monitoring area are accumulated. D i : Population activity deviation of the i-th sentinel species. Calculated as: (Current activity baseline activity) / Baseline activity.
[0185] The food chain stability index, denoted as FCSI, was calculated as follows: The system aggregated data from the past 30 days and found that the population activity index of "Frog-01" decreased by 30% compared to the baseline, while that of "Weasel-03" increased by 15%. The calculated food chain stability index was 0.39.
[0186] |D i|: Take the absolute value of the deviation, indicating that the system only cares about the magnitude of the deviation, not the direction (increase or decrease).
[0187] W i : The ecological weight factor of the i-th sentinel species, preset based on its importance in the ecosystem.
[0188] Specific calculations in this case:
[0189] |D (蛙-01) |=|30%|=0.3;|D (鼬-03) |=|+15%|=0.15;W (蛙-01) =1.0; W (鼬-03) =0.6;
[0190] FCSI=(|D( 蛙) |×W (蛙) )+(|D (鼬) |×W (鼬) ); FCSI=(0.30×1.0)+(0.15×0.6); FCSI=0.30+0.09; FCSI=0.390;
[0191] Risk level index calculation: The risk level fusion unit performs weighted fusion (with weights of 0.4 and 0.6 respectively), calculates the weighted value of the immediate risk part: 0.84 × 0.4 = 0.336; calculates the weighted value of the long-term trend part: 0.390 × 0.6 = 0.234; add the two together to calculate the risk level index of 0.570.
[0192] First assessment result: Since the risk level index of 0.570 is between the second risk threshold of 0.4 and the first risk threshold of 0.7, the first assessment unit generates a second medium risk label.
[0193] Second level of judgment (micro-behavioral coercion):
[0194] The second analysis unit extracts and analyzes multiple video data segments containing foraging behaviors of swamp frogs.
[0195] Example of foraging efficiency coefficient calculation: Analyzing a 60-second foraging cycle, by tracking key skeletal points, the following statistics were obtained: total intake time was 25 seconds, total alert and search time was 35 seconds, search movement cost was 5 meters, and the number of behavior interruptions was 8.
[0196] The total cost is calculated to be 51.5 based on the preset weights (0.8, 1.5, 2.0 in this embodiment).
[0197] The second analysis unit extracts and analyzes multiple video data segments containing foraging behaviors of swamp frogs.
[0198] Example of foraging efficiency coefficient (FEC) calculation: Analyzing one foraging cycle (60 seconds), by tracking key skeletal points, the following results were obtained:
[0199] Total intake time Tingestion = 25 seconds (head lower than shoulders and minimal displacement).
[0200] Total alert and search duration: Tvigilance = 35 seconds.
[0201] Search movement cost Dsearch = 5 meters.
[0202] The number of behavioral interruptions (Ninterrupt) is 8 (the number of times the behavior switches from intake to alert).
[0203] Let the weights be w1=0.8, w2=1.5, w3=2.0;
[0204] The total cost is calculated as follows: (w1×35)+(w2×5)+(w3×8)=28+7.5+16=51.5.
[0205] Calculate the foraging efficiency coefficient (FEC):
[0206] FEC = Tingestion / (Tingestion + Total Cost);
[0207] FEC=25 / (25+51.5)=25 / 76.5≈0.3268;
[0208] Second assessment result: The average FEC of all predator-prey interaction events (e.g., 50 samples) within the aggregation area is 0.41. Since 0.41 < 0.5 (efficiency threshold), the second assessment unit generates the first stress risk label.
[0209] The decision-making module received the "second medium risk label" and the "first coercion risk label".
[0210] Based on the corresponding strategy matrix, the system matches and generates a second strategy.
[0211] Strategy details: Core area-07 is designated as the second priority restoration area. Two sub-areas (preferably near water sources) with a total area of no less than 10% of the core area are automatically marked on the system map for subsequent vegetation restoration. Simultaneously, six recommended predation risk buffer zones are planned between the foraging area and the woodland. This strategy is sent to the reserve management platform to guide subsequent ecological interventions.
[0212] Data Example 2:
[0213] Eagle-rabbit predation interactions in alpine meadow areas;
[0214] Scene and preset parameters:
[0215] Scenario Setting: At the boundary between alpine meadow and woodland (grid number: Alpine 03), which is designated as the core sensing area, this area is a typical activity zone for pikas (sentinel species) and their medium-sized predator, the sparrowhawk (sentinel species). Acoustic sensor 03 and image sensor 03 monitoring nodes are deployed within this area.
[0216] System preset parameters:
[0217] The model pre-training included the alarm screech of a "pika" (species category code: pika-01) and the swooping sound of a "sparrowhawk" (species category code: sparrowhawk-02).
[0218] Species identification confidence threshold: 0.95 (high-altitude environments have lower background noise, so a higher threshold can be set).
[0219] Ecological knowledge graph: Stores the predator-prey relationship between "Sparrowhawk-02" and "Pika-01".
[0220] Risk level index thresholds: First risk threshold = 0.7, second risk threshold = 0.4.
[0221] Foraging efficiency coefficient and efficiency threshold: 0.5.
[0222] Step 1: Acoustic-triggered multimodal data acquisition;
[0223] At time T0, the intelligent sensing module of acoustic sensor 03 detects a short, high-frequency alarm scream from a rabbit, and the edge analysis unit is immediately activated.
[0224] Acoustic Feature Extraction and Recognition: The system processes the received first modal data and extracts acoustic feature vectors after high-pass filtering (above 2500 Hz). The species voiceprint recognition subunit inputs the feature vectors into the model and calculates a matching confidence score of 0.98 for "pika-01", which is higher than the threshold of 0.95, thus confirming the identification of the sentinel species "pika".
[0225] Sound source localization and orientation acquisition: The sound source localization subunit calculates the sound source azimuth as: azimuth angle 155 degrees, elevation angle 25 degrees. Then, the first acquisition command is generated, driving the pan-tilt unit of image sensor 03 to precisely rotate to the specified position, and acquiring the first static image of a rabbit standing upright at the entrance of its burrow and 30 seconds of second video data, forming a multimodal dataset.
[0226] Step 2: Analysis of predator-prey interaction events based on multimodal fusion;
[0227] At time T1 (15 seconds after T0), acoustic sensor 03 detected a low, rumbling sound of wind breaking, while at the same time, a fast-moving shadow appeared in the field of view of image sensor 03.
[0228] Secondary acquisition and analysis of multimodal data: The system immediately identifies the new sound source and confirms it as the swooping sound of "Sparrowhawk-02". Simultaneously, the image analysis unit detects and identifies the sparrowhawk entity in the video data. The behavior classification subunit identifies the pika's behavior as "alert" and the sparrowhawk's behavior as "swooping pursuit".
[0229] Event association determination: The event association determination unit aggregates information and determines that within a very short spatiotemporal proximity (15-second time difference, overlapping spatial locations), two species with a predator-prey relationship exist, and their behaviors (vigilance, pursuit) are highly correlated. If the system determines that the rules are met, it generates structured behavioral semantic tags representing the predator-prey interaction event.
[0230] Step 3: Dual Risk Assessment and Decision Generation;
[0231] The dual risk assessment module is activated, and two judgments are made.
[0232] First level of assessment (macro-ecological risk):
[0233] Real-time risk value calculation: As a key species in the region, the pika has a base score of 0.8 for its protection level mapping. Combined with the weighting coefficient of 1.2 for event type (predation), the real-time risk value is 0.96.
[0234] The food chain stability index is calculated as shown in the second example (predator-prey balance ratio method): The system aggregates data from the past 60 days and finds that the population activity index of "pika-01" has decreased significantly by 50% compared to the baseline, while the activity frequency of "sparrowhawk-02" has only increased slightly by 5%. The calculated food chain stability index is (0.5 / 1.0) / (1.05 / 1.0) = 0.476, which is severely underestimated. The predator-prey balance ratio model pays close attention to the causal relationship between species, namely the core fact that "sparrowhawk" depends on "pika" for survival. It is concerned not only with the magnitude but also with the direction of change. A decrease in prey (smaller numerator) and an increase in predators (larger denominator) is the most dangerous signal, leading to a sharp drop in the index.
[0235] To avoid misjudgment: If only the weighted total deviation model of Example 1 is used, when one species (such as pikas) decreases while another unrelated species (such as plants) increases, the total deviation may not change significantly, thus masking the serious problem of pika population collapse. To avoid one-sidedness: If only the predator-prey balance ratio model of Example 2 is used, it may only focus on a few key food chains, ignoring the widespread decline of multiple species caused by environmental changes (such as water pollution). The results of the two models can corroborate each other. The simultaneous increase of macro-indices and decrease of micro-indices enhances the confidence of risk assessment.
[0236] Risk level index calculation: Perform weighted fusion (weights 0.4 and 0.6), and the risk level index is calculated as (0.4×0.96)+(0.6×(1-0.476))=0.384+0.3144=0.6984.
[0237] First assessment result: Since the risk level index is less than the first risk threshold of 0.7, the first assessment unit generates a second medium risk label.
[0238] Second level of judgment (micro-behavioral coercion):
[0239] Example of foraging efficiency coefficient calculation: The second analysis unit analyzed recently collected pika foraging videos. In a 60-second foraging cycle, it was found that the total intake time was only 15 seconds, while the total alert and search time was as high as 45 seconds, with 12 interruptions in behavior during this period.
[0240] Calculate the total cost (weights same as in Data Example 1): (w1×Tvigilance)+(w2×Dsearch)+(w3×Ninterrupt);
[0241] Total cost = (0.8 × 45) + (1.5 × 3) + (2.0 × 12) = 64.5.
[0242] Calculate the foraging efficiency coefficient (FEC):
[0243] FEC = Tingestion / (Tingestion + Total Cost) FEC = 15 / (15 + 64.5) FEC = 15 / 79.5 FEC ≈ 0.19;
[0244] The calculated foraging efficiency coefficient is much lower than 0.2, for example, 0.19.
[0245] Second assessment result: The average foraging efficiency coefficient of all pika foraging events within the aggregation area is 0.25. Since this average value is much lower than the efficiency threshold of 0.5, the second assessment unit generates a first stress risk label.
[0246] Decision generation: The decision module receives the "second risk label" and the "first coercion risk label". This combination corresponds to the highest level of urgency, and the system matches and generates the second strategy.
[0247] The second strategy is as follows: mark "High Mountain-03" as the second priority recovery zone, automatically mark 2 to 3 priority recovery zones with a total area of not less than 10% of the total area of the sensing area on the system map, and plan and establish 6 to 10 predation risk buffer zones between the foraging area and the woodland.
[0248] This embodiment aims to fundamentally revolutionize the current situation in traditional wildlife conservation work, which suffers from single monitoring dimensions, insufficient analytical depth, and lagging decision-making response. Its core technical principle is built upon an intelligent closed loop of "acoustic sentinel triggering - multi-modal homomorphic fusion semanticization - dual risk quantification assessment - graded strategy precise response". First, this invention abandons the passive and inefficient continuous recording mode of traditional equipment. It innovatively utilizes a low-power, wide-coverage first sensing device (microphone array) as a "stethoscope" for the ecological environment. By capturing high-value acoustic signals emitted by sentinel species in real time, it accurately triggers event recording and immediately guides a high-resolution second sensing device (HD and infrared camera equipment) for directional collaborative acquisition. This cross-modal collaborative mechanism improves the capture rate of key predator-prey interaction events while reducing system energy consumption and data redundancy. Secondly, after acquiring the data, the system transforms the raw, fragmented multimodal data into structured behavioral semantic tags that can be deeply understood by machines through a fusion analysis module. This "semanticization" process is the key to this invention. It combines target detection, skeletal keypoint tracking, posture sequence analysis, and ecological knowledge graph association, enabling the system to accurately "understand" the complex interactions between species, laying a solid foundation for subsequent quantitative risk assessment. The essence of this invention lies in its unique dual risk assessment module, which provides a three-dimensional diagnosis of ecological health from both macro and micro dimensions: Macroscopically, by fusing the "immediate risk value" reflecting short-term threats with the "food chain stability index" revealing long-term trends, a comprehensive "risk level index" is generated to fully assess the immediate impact and long-term health status of the ecosystem; Microscopically, it innovatively proposes the quantitative indicator of "foraging efficiency coefficient (FEC)," which, through detailed analysis of animals' micro-behaviors such as intake, vigilance, searching, and interruption during foraging, constructs a mathematical model that can quantify animals' "psychological stress" or "environmental stress," thereby capturing implicit, early warning signals in the early stages of population decline. Based on this technical principle, this invention realizes a shift in monitoring mode from passive waiting to proactive early warning, improving monitoring efficiency and timeliness; through behavioral semantics and dual risk assessment, the analysis depth goes from the superficial appearance of species counting to the connotation of ecological relationship health; at the same time, it transforms complex ecological risks into objective and comparable quantitative data, providing managers with unprecedentedly accurate decision-making basis, making ecological intervention more scientific, targeted and economical.This ability to make precise decisions stems from a rigorous strategy-making logic that deeply understands the different combinations of two risk labels: when "high risk" and "stress risk" coexist (as in the "eagle-rabbit" example), it indicates a macro-level imbalance in the ecosystem and that individual species are already under immense pressure. This is the most urgent systemic crisis signal, thus requiring the activation of the highest-level first strategy, namely, large-scale habitat restoration and the construction of numerous predation buffer zones, addressing both the symptoms and the root cause; when "medium risk" and "stress risk" coexist (as in the "weasel-frog" example), it represents a "warning period" in which the system is sliding towards imbalance, requiring a medium-level second strategy to "nip problems in the bud"; when risk and stress are not synchronized, such as "high risk" but "no stress," it may suggest that the risk originates from non-predation factors, and the strategy focuses on habitat restoration; conversely, "low risk" but "stress" may only indicate the emergence of new predation pressure in a localized area, which can be precisely resolved by establishing buffer zones on a small scale and at low cost.
[0249] The threshold is set to facilitate comparison. The size of the threshold depends on the amount of sample data and the number of bases set by those skilled in the art for each set of sample data; as long as it does not affect the ratio between the parameter and the quantized value, it is acceptable.
[0250] The above formulas are all derived from software simulation using a large amount of data and are selected to be close to the actual values. The coefficients in the formulas are set by those skilled in the art according to the actual situation. The above description is only a preferred embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any equivalent substitutions or changes made by those skilled in the art within the technical scope disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the protection scope of the present invention.
Claims
1. A wildlife intelligent identification system based on multimodal perception fusion in a protected area, characterized in that: include: The intelligent sensing module is used to receive first-mode data from the first sensing device; It analyzes the first modal data in real time to identify the triggering events of the acoustic characteristics of the preset sentinel species, determines the spatial orientation information of the triggering events, generates a first acquisition command, and drives the second sensing device to execute the first acquisition command to acquire the second modal data. After aligning the second modal data with the first modal data, it summarizes the data to obtain a multimodal dataset. The first sensing device is a microphone array. The second sensing device integrates an infrared thermal imaging device and a high-definition camera device. After the second sensing device completes the turning and positioning of spatial orientation information, it collects the ambient brightness. When the ambient brightness is lower than the preset brightness threshold, it automatically switches to the infrared thermal imaging mode; otherwise, it maintains the high-definition camera mode. The fusion analysis module is used to process the multimodal dataset and use a target detection algorithm to identify all predator-prey interaction events in the current sensing unit within a preset time window, generate structured behavioral semantic labels, and calculate the instantaneous risk value of the event level of each predator-prey interaction event to construct a food chain stability index. The dual risk assessment module is used to perform a weighted fusion calculation on the real-time risk value and the food chain stability index to obtain a risk level index, which is used to perform the first judgment and obtain the first assessment result. Further analysis is performed on the second video data with structured behavioral semantic tags to calculate the foraging efficiency coefficient (FEC) of each predator-prey interaction event. Within a preset time window, the j-th perception region is aggregated, and the average value of the foraging efficiency coefficient (FEC) of all predator-prey interaction events is calculated. A second judgment is then performed to obtain the second evaluation result. The decision-making module is used to combine the results of the first evaluation and the second evaluation to generate corresponding strategies.
2. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The intelligent sensing module includes a partitioning subunit and a deployment subunit; Based on preset ecological and geographical data, the zoning sub-units divide the target wildlife reserve into several sensing areas through non-uniform grid division. Among them, the ecological and geographical data includes historical species activity heat maps, topographic data and vegetation cover type data. The perception area marked by the ecological and geographical data is divided into at least two priority areas: core perception area and regular perception area. Among them, the core perception area is marked as a water source, a high-density activity hotspot or a habitat of rare species. The deployment subunit is used to deploy a first sensing device and a second sensing device in each sensing area, and to establish a collaborative communication link between the first sensing device and the second sensing device.
3. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The intelligent sensing module also includes an edge analysis unit; The edge analysis unit is used to process the first modal data collected by the first sensing device in real time and drive the second sensing device to perform a coordinated response, including: An acoustic feature extraction subunit is used to perform preprocessing on the received first modal data. The preprocessing includes filtering out environmental background noise outside a preset frequency band using a bandpass filtering algorithm; and framing and windowing the preprocessed first modal data, and extracting the Mel-frequency cepstral coefficients (MFCCs) of each frame as an acoustic feature vector. The species voiceprint recognition subunit is used to embed a pre-trained lightweight convolutional neural network model, input the acoustic feature vector into the lightweight convolutional neural network model to calculate the matching confidence score between the first modality data and one or more preset sentinel species voiceprint models; and when any matching confidence score is higher than the preset species recognition confidence threshold, the corresponding sentinel species acoustic feature is confirmed to be identified, and the category code of the corresponding sentinel species is output. The sound source localization subunit is used to calculate the time delay of the first modal data arriving at each microphone channel based on the first modal data acquired from the microphone array of the first sensing device, using the Time of Arrival (TDOA) algorithm; and to calculate the azimuth and pitch angles representing the direction of the sound source based on the time delay and the geometric layout information of the microphone array, as spatial orientation information, and generate a first acquisition command; the first acquisition command includes: an angle indicating that the second sensing device needs to rotate in the horizontal direction, i.e., the azimuth angle; Indicates the angle that the second sensing device needs to be adjusted in the vertical direction, i.e., the pitch angle; The sentinel species is locked and continuously tracked, with at least two still photos taken and a 30-second video recorded to obtain corresponding first still image and second video data.
4. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The intelligent sensing module also includes a directional acquisition subunit and a spatiotemporal alignment subunit; The directional acquisition subunit is used to receive the first acquisition command and activate the second sensing device accordingly. The second sensing device is controlled to execute a directional acquisition program based on the spatial orientation information to obtain a first static image and a second video data for the sentinel species as second modal data; The spatiotemporal alignment subunit is used to associate the first modal data with the second modal data to form a multimodal dataset.
5. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The fusion analysis module includes an image analysis unit, a behavior classification subunit, and an event association determination unit; The image analysis unit is used to perform a target detection algorithm on consecutive frames in the first static image and the second video data to identify and locate the species to which one or more sentinel species belong, and associate the spatiotemporal location information of the i-th sentinel species. The behavior classification subunit is used to perform pose analysis through image processing calculation for each identified sentinel species to obtain the skeletal key point coordinate sequence of the i-th sentinel species at consecutive time points. Furthermore, based on the skeletal key point coordinate sequence, the dynamic posture sequence of the i-th sentinel species is calculated and generated, and input into the pre-built behavior recognition model to calculate the matching probability score between the dynamic posture sequence of the i-th sentinel species and multiple preset behavior types, and select the behavior type with the highest matching probability score as the output. Behavioral types include: chirping, stealth, chasing, drinking, or alertness; The identified species identity, spatiotemporal location information, and behavior types output by the behavior classification subunit are aggregated, and the aggregated information is formatted into a group of sentinel species. When at least two groups of sentinel species are received, calculate their spatiotemporal proximity. Based on two groups of sentinel species, a pre-defined ecological knowledge graph is queried to obtain the ecological relationship between the two groups of sentinel species. By combining spatial and temporal proximity with the ecological relationships between the two groups of sentinel species, it was determined whether a predator-prey interaction event existed, including: The rules engine is used to define the criteria for classifying a biological event as a predator-prey interaction event. Within a preset spatiotemporal threshold, continuously identify cross-modal joint features of at least two sentinel species belonging to different species; It was confirmed that a predefined predator-prey relationship exists between the two groups of sentinel species; When the spatiotemporal proximity is less than or equal to a preset proximity threshold, and it is confirmed that there is a predefined predator-prey correspondence between the two groups of sentinel species, a structured behavioral semantic label representing the predator-prey interaction event is generated.
6. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The event type and the species involved are parsed from the structured behavioral semantic tags. The corresponding protection level is obtained by querying the preset ecological knowledge graph and the basic score is determined according to the protection level. Based on the event type, the event weighting coefficient is determined from the preset event risk weight library. Perform a preset multiplication operation between the base score and the event weighting coefficient to generate the instantaneous risk value of the event level; Within a preset time window, aggregate multimodal datasets and calculate population activity indicators; Based on the ecological knowledge graph, the species associated with the upstream and downstream of the food chain and the preset baseline population activity ratio are retrieved. The current population activity index is compared with the baseline population activity ratio to calculate the deviation, and the food chain stability index is obtained by correlation calculation.
7. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The dual risk assessment module includes a risk level fusion unit and a first assessment unit; The risk level fusion unit is used to perform a weighted fusion calculation between the instantaneous risk value and the food chain stability index to obtain the risk level index. The first assessment unit is used to preset a first risk threshold and a second risk threshold, wherein the first risk threshold is greater than the second risk threshold, and to assess the risk level index to obtain a first assessment result, including: When the risk level index is higher than the first risk threshold, a first high-risk label is generated. When the risk level index is between the first risk threshold and the second risk threshold, a second risk label is generated; When the risk level index is less than the second risk threshold, a third low-risk label is generated.
8. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The dual risk assessment module also includes a second analysis unit, which is used to further analyze the second video data that generates structured behavioral semantic tags representing predator-prey interaction events. Based on the timestamps recorded in the structured behavioral semantic tags, a second video data segment of a preset duration is extracted as the foraging cycle for the target analysis. During the current foraging cycle, the target animal entity in the second video data is continuously tracked by skeletal key points to obtain the three-dimensional spatial coordinate time series of at least one head key point, one shoulder key point, and one body centroid key point. By traversing each frame of the three-dimensional spatial coordinate time series, and based on the relative spatial relationship and displacement velocity of key points, the animal's foraging behavior is dynamically and mutually exclusively classified into one of the following two core microstates; The current frame is determined to belong to the ingested micro state if and only if the vertical coordinate of the head key point is lower than the vertical coordinate of the shoulder key point, and the displacement of the body center of mass key point in a unit time is less than the first preset velocity threshold. When the vertical coordinate of the head key point is not lower than the vertical coordinate of the shoulder key point, the current frame is determined to be in the alert and search micro-state. Throughout the foraging cycle, extract the total intake time (Tingestion), total alert and search time (Tvigilance), search movement cost (Dsearch), and number of behavior interruptions (Ninterrupt). The weighted total vigilance and search duration (Tvigilance), search movement cost (Dsearch), and number of behavior interruptions (Ninterrupt) are summed and added to the total intake duration (Tingestion) in the denominator. The total intake duration (Tingestion) is then placed in the numerator as a benefit term. The ratio of the numerator to the denominator is used to obtain the foraging efficiency coefficient, denoted as FEC.
9. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The dual risk assessment module also includes a second assessment unit, which is used to aggregate the j-th perception area within a preset time window and calculate the average value of the foraging efficiency coefficient (FEC) of several predator-prey interaction events. A preset efficiency threshold is set, and the average value of the foraging efficiency coefficient (FEC) is evaluated to obtain a second evaluation result, including: when the average value of the foraging efficiency coefficient (FEC) is less than the efficiency threshold, it indicates that there is a potential stress risk in the current sensing area, and a first stress risk label is generated. When the average value of the foraging efficiency coefficient (FEC) is greater than or equal to the efficiency threshold, a second qualified label is generated, indicating that the potential stress risk in the current sensing area is within the expected range.
10. The wildlife intelligent identification system based on multimodal perception fusion in protected areas according to claim 1, characterized in that, The decision-making module is used to combine the first evaluation result and the second evaluation result to generate a corresponding strategy, including: When both the first high-risk label and the first stress risk label are identified simultaneously, a first strategy is generated, including: designating the current sensing area as the first priority recovery area, selecting at least 4 to 5 priority recovery areas on the map with a total area not less than 16% to 20% of the total area of the j-th sensing area, planting restoration plants 5m to 10m away from water sources; and establishing at least 10 to 15 predation risk buffer zones. When the second risk label and the first stress risk label are identified at the same time, a second strategy is generated, including: taking the current sensing area as the second priority recovery area, establishing at least 2 to 3 priority recovery areas on the map with a total area not less than 10% of the total area of the j-th sensing area, and establishing 6 to 10 predation risk buffer zones between the foraging area and the woodland. When the first high-risk label and the second qualified label are identified at the same time, a third strategy is generated, including: taking the current sensing area as the third priority recovery area, and having at least 4 to 5 priority recovery areas on the map with a total area not less than 16% to 20% of the total area of the jth sensing area. When the second risk label and the second qualified label are identified at the same time, a fourth strategy is generated, including: taking the current sensing area as the fourth priority recovery area, and having at least 2 to 3 priority recovery areas on the map with a total area not less than 10% of the total area of the jth sensing area. When the third low-risk label and the first stress risk label are identified at the same time, a fifth strategy is generated, including: taking the current perception area as the fifth priority recovery area, and establishing 3 to 5 predation risk buffer zones between the foraging area and the woodland, with the width of the predation risk buffer zones set to 3 to 5 meters. When the third low-risk label and the second qualified label are identified simultaneously, no intervention is made in the current sensing area, and continuous monitoring is carried out.