Animal population non-contact monitoring method and system based on AI vision
Patent Information
- Application Number
- CN202611024667.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-25
AI Technical Summary
常规做法中,通常依赖单一的表型特征进行个体识别,特征提取方式较简单,对光照变化、姿态差异、部分遮挡等环境噪声敏感
[0051]基于自然生物标记特征的复合表达,避免了传统物理标记对动物的伤害与应激反应,实现完全无接触式监测。结合度量学习构建的个体特征嵌入空间,使跨时空场景下的个体再识别精度显著提升,解决了野外环境中因光照、角度差异导致的识别困难问题。该方法大幅降低了人工实地观测的劳动强度,支持对大规模种群中每个个体的长期持续追踪。
Smart Images

Figure CN122821592A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and ecological monitoring technology, and in particular to a non-contact monitoring method and system for animal populations based on AI vision. Background Technology
[0002] In existing technologies, vision-based animal population monitoring typically relies on manual observation, physical tagging, or GPS positioning devices. Traditional manual observation methods are inefficient, struggle to sustain large-scale habitat coverage, and disrupt animal behavior. Physical tagging methods require capturing individual animals, posing invasive risks that can lead to stress or injury, while the tags are also expensive and prone to detachment. While GPS tracking can provide precise location information, its limited battery life, device weight restricts monitoring of small species, and hardware costs limit deployment for large populations.
[0003] With the development of camera traps and computer vision technologies, some solutions attempt to achieve individual identification through image recognition. Conventional methods typically rely on single phenotypic features for individual identification, with relatively simple feature extraction methods that are sensitive to environmental noise such as changes in lighting, pose differences, and partial occlusion. The appearance of the same animal individual can vary significantly under different spatiotemporal conditions, and existing methods often suffer from unstable feature representation, leading to significant fluctuations in re-identification accuracy.
[0004] Furthermore, existing monitoring systems primarily focus on individual counts or identification, lacking the ability to automatically infer social relationships within a population. Kinship analysis typically relies on manual analysis of behavioral videos or molecular biology methods, which is not only costly and time-consuming but also difficult to update in real time. The quantification of behavioral interaction patterns mainly depends on manual annotation, which is highly subjective and cannot handle large-scale spatiotemporal data. The construction of population social network structures mostly remains at the theoretical model level, lacking an end-to-end system that automatically interfaces with monitoring data. Summary of the Invention
[0005] This invention provides a non-contact animal population monitoring method and system based on AI vision, which can solve the problems in the prior art.
[0006] A first aspect of this invention provides a non-contact animal population monitoring method based on AI vision, comprising:
[0007] The process involves acquiring animal image sequences from multiple spatiotemporal nodes, extracting natural biomarker features of individuals from the animal image sequences, and including composite feature expressions of phenotypic texture features and morphological contour features.
[0008] An individual feature embedding space is constructed based on metric learning. The composite feature expression is mapped to the embedding space to form an individual feature vector. Cross-temporal individual re-identification is achieved by minimizing the feature vector distance of the same individual under different spatiotemporal conditions, resulting in a spatiotemporal trajectory sequence with individual identity.
[0009] Based on the spatial co-occurrence patterns and temporal interaction patterns among individuals in the spatiotemporal trajectory sequence, the social association strength between individual pairs is calculated, and the population social network topology is constructed.
[0010] In the aforementioned population social network topology, by analyzing the frequency of interactions, spatial proximity persistence, and behavioral synchronicity among individuals, and combining the temporal constraints of individual life cycle events, the kinship level among individuals is inferred, and a population phylogenetic tree structure is generated.
[0011] The feature vectors of the newly monitored individuals are backfilled into the individual feature embedding space for incremental updates, and the calculation rules for the strength of social associations are adaptively adjusted using the newly added kinship labeling.
[0012] An individual feature embedding space is constructed based on metric learning. The composite feature representation is mapped to the embedding space to form an individual feature vector. Cross-spatial-temporal individual re-identification is achieved by minimizing the feature vector distance of the same individual under different spatiotemporal conditions, resulting in a spatiotemporal trajectory sequence with individual identity identifiers, including:
[0013] Construct a training sample set of triplets, where each triplet contains anchor features, positive sample features, and negative sample features;
[0014] The metric learning network is trained based on the triplet training sample set. By minimizing the feature vector distance between the anchor feature and the positive sample feature in the embedding space, and simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space, the metric learning network learns to map the composite feature representation into individual feature vectors that are invariant to changes in illumination, viewpoint, and pose.
[0015] For an animal image to be identified, its composite feature representation is input into the trained metric learning network to obtain a query feature vector. The distance metric between the query feature vector and all individual feature vectors stored in the feature database is calculated. When the minimum distance metric satisfies the similarity judgment condition, the query feature vector is associated with the corresponding individual and the feature vector set of the corresponding individual is updated.
[0016] Spatial location information with the same individual association identifier is linked together in chronological order to form a spatiotemporal trajectory sequence.
[0017] The metric learning network is trained based on the triplet training sample set. This is achieved by minimizing the feature vector distance between the anchor feature and the positive sample feature in the embedding space, while simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space.
[0018] Construct a triplet loss function, which includes a positive pair distance term and a negative pair distance term. Quantize the feature vector distance between anchor features and positive sample features in the embedding space based on the positive pair distance term, and quantize the feature vector distance between anchor features and negative sample features in the embedding space based on the negative pair distance term.
[0019] Triple samples are extracted from the triple training sample set. The anchor features, positive sample features, and negative sample features of the triple samples are input into the metric learning network. The input features are mapped into feature vectors in the embedding space according to the metric learning network. The distance between the feature vectors is calculated and substituted into the triple loss function to obtain the values of the positive pair distance term and the negative pair distance term.
[0020] The gradient of the triplet loss function with respect to the network parameters of the metric learning network is calculated through backpropagation. The network parameters are then updated using the gradient, thereby decreasing the value of the positive pair distance term and increasing the value of the negative pair distance term.
[0021] The extraction, mapping, calculation, and update processes are repeatedly executed to minimize the feature vector distance between the anchor feature and the positive sample feature in the embedding space, while simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space.
[0022] Based on the spatial co-occurrence patterns and temporal interaction patterns among individuals in the spatiotemporal trajectory sequence, the social association strength between individual pairs is calculated, and the population social network topology is constructed, including:
[0023] Spatiotemporal trajectory sequences are divided into spatiotemporal grids. Within each grid cell, the co-occurrence frequency and co-occurrence duration of individual pairs are counted. Stable spatial co-occurrence relationships are identified by analyzing the co-occurrence distribution patterns of individual pairs in multiple grid cells. These stable spatial co-occurrence relationships characterize that individual pairs have a continuous spatial clustering preference.
[0024] Based on the changes in the motion trajectory of individuals in the spatiotemporal trajectory sequence, proximity events and following events between individual pairs are extracted. The proximity events are identified by detecting the decreasing trend of the distance between individuals, and the following events are identified by detecting the temporal consistency of the individual's motion direction. The frequency and persistence of the proximity events and the following events are used as quantitative indicators of the temporal interaction pattern.
[0025] The strength value of the stable spatial co-occurrence relationship and the quantitative index of the temporal interaction pattern are weighted in multiple dimensions to generate a value of the social association strength between individual pairs.
[0026] Using all monitored individuals as network nodes, and the social association strength value as the edge weight between nodes, a weighted graph structure is constructed. In the weighted graph structure, connections with edge weights exceeding the association threshold are retained to form a population social network topology.
[0027] Using all monitored individuals as network nodes, and the social association strength values as edge weights between nodes, a weighted graph structure is constructed. In this weighted graph structure, connections with edge weights exceeding an association threshold are retained to form the population social network topology, including:
[0028] Each monitored individual is mapped to a network node in a weighted graph structure, and a node feature vector is constructed for each network node.
[0029] Calculate the social association strength between any two network nodes, and use the social association strength as the edge weight connecting the two network nodes to construct a fully connected weighted graph structure.
[0030] Statistical analysis is performed on the edge weight distribution in the fully connected weighted graph structure, and the dynamic association threshold is determined by calculating the mean and standard deviation of the edge weights;
[0031] Remove edge connections in the fully connected weighted graph structure whose edge weights are lower than the dynamic association threshold, and retain edge connections whose edge weights are higher than the dynamic association threshold to form a sparse population social network topology.
[0032] In the aforementioned population social network topology, by analyzing the frequency of interactions, spatial proximity persistence, and behavioral synchronicity among individuals, and combining this with the temporal constraints of individual life cycle events, the kinship level among individuals is inferred, generating a population phylogenetic tree structure including:
[0033] Extract individual pairs with edge connections from the topology of the population social network, calculate the cumulative value of interaction frequency, duration of spatial proximity and similarity of behavioral patterns of the individual pairs, and fuse the three into a kinship indicator value through weighted coefficients;
[0034] The first observation time and age stage marker of an individual are extracted as life cycle event records, and intergenerational time constraints are established based on the life cycle event records;
[0035] Individual pairs whose kinship indicator values exceed the judgment threshold and satisfy the generational time constraints are marked as kinship links. Kinship levels are determined based on the age differences between individuals. A population phylogenetic tree structure is constructed with individuals as nodes and kinship links as directed edges.
[0036] Incremental updates are performed by backfilling the feature vectors of newly monitored individuals into the individual feature embedding space, and adaptive adjustments are made to the calculation rules for the strength of social associations using newly added kinship annotations, including:
[0037] When a new monitored individual is identified, the composite feature expression of the new monitored individual is extracted and mapped to a new feature vector through a metric learning network. The distance metric matrix between the new feature vector and the feature vectors already stored in the individual feature embedding space is calculated. Based on the distance metric matrix, the influence of the new feature vector on the topology of the individual feature embedding space is analyzed. The new feature vector is then backfilled into the individual feature embedding space to complete the incremental update.
[0038] New kinship labels are extracted from the phylogenetic tree structure of the population. The new kinship labels include pairs of individuals with verified kinship. Observational data of the pairs of individuals are extracted in three dimensions: interaction frequency, spatial proximity persistence, and behavioral synchronicity.
[0039] The distribution characteristics of the observed data in three dimensions are statistically analyzed, and the differences between the distribution characteristics and those of unrelated individuals in the three dimensions are analyzed to quantify the contribution of each dimension to kinship identification. Based on the contribution, the weight ratios of the interaction frequency dimension, spatial proximity persistence dimension, and behavioral synchronicity dimension in the calculation rule of social association strength are adjusted, and the calculation rule of social association strength is adaptively adjusted using the adjusted weight ratios.
[0040] A second aspect of the present invention provides a non-contact animal population monitoring system based on AI vision, comprising:
[0041] The feature extraction unit is used to acquire animal image sequences collected from multiple spatiotemporal nodes, and extract natural biological marker features of individuals from the animal image sequences. The natural biological marker features include a composite feature expression of phenotypic texture features and morphological contour features.
[0042] The individual identification unit is used to construct an individual feature embedding space based on metric learning, map the composite feature expression to the embedding space to form an individual feature vector, and achieve cross-temporal individual re-identification by minimizing the feature vector distance of the same individual under different spatiotemporal conditions, thereby obtaining a spatiotemporal trajectory sequence with individual identity identifier;
[0043] The network construction unit is used to calculate the social association strength between individual pairs based on the spatial co-occurrence pattern and temporal interaction pattern between individuals in the spatiotemporal trajectory sequence, and to construct the population social network topology.
[0044] The phylogenetic inference unit is used to infer the kinship level between individuals and generate a population phylogenetic tree structure by analyzing the frequency of interaction, spatial proximity persistence and behavioral synchronicity between individuals, and combining the temporal constraints of individual life cycle events in the topology of the population social network.
[0045] The incremental update unit is used to backfill the feature vector of the newly monitored individual into the individual feature embedding space for incremental update, and to adaptively adjust the calculation rules of the social association strength using the newly added kinship label.
[0046] A third aspect of the present invention provides an electronic device, comprising:
[0047] processor;
[0048] Memory used to store processor-executable instructions;
[0049] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0050] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0051] This method utilizes composite expressions based on natural biomarker features, avoiding the harm and stress caused to animals by traditional physical markers, and achieving completely contactless monitoring. The integration of an individual feature embedding space constructed using metric learning significantly improves the accuracy of individual re-identification across spatiotemporal scenarios, solving the identification difficulties caused by differences in lighting and angles in the wild. This method greatly reduces the labor intensity of manual field observation and supports long-term continuous tracking of each individual in large-scale populations.
[0052] Spatiotemporal trajectories automatically extracted from animal image sequences reveal quantitative patterns of spatial co-occurrence and behavioral interactions among individuals. By analyzing interaction frequency, spatial proximity persistence, and behavioral synchronicity, combined with temporal constraints of life-cycle events, kinship levels can be inferred without gene sampling. The generated population phylogenetic tree structure accurately reflects the real kinship network, providing a reliable data foundation for studying reproductive strategies and family structure.
[0053] The constructed social network topology dynamically presents the strength of social connections within the population, enabling the intuitive quantification of social characteristics such as dominance hierarchy and alliance relationships. Joint analysis based on spatiotemporal co-occurrence patterns and temporal interaction patterns effectively distinguishes between random encounters and stable social relationships, avoiding the errors of subjective judgment inherent in traditional methods. This topology can support subsequent in-depth studies such as disease transmission simulation and foraging cooperation analysis. Attached Figure Description
[0054] Figure 1 This is a flowchart illustrating the non-contact animal population monitoring method based on AI vision, as described in an embodiment of the present invention.
[0055] Figure 2 This is a flowchart illustrating the population kinship inference based on generational time constraints in an embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0058] Figure 1 This is a flowchart illustrating the non-contact animal population monitoring method based on AI vision, as described in an embodiment of the present invention.
[0059] AI-based vision-based non-contact animal population monitoring methods include:
[0060] The process involves acquiring animal image sequences from multiple spatiotemporal nodes, extracting natural biomarker features of individuals from the animal image sequences, and including composite feature expressions of phenotypic texture features and morphological contour features.
[0061] An individual feature embedding space is constructed based on metric learning. The composite feature expression is mapped to the embedding space to form an individual feature vector. Cross-temporal individual re-identification is achieved by minimizing the feature vector distance of the same individual under different spatiotemporal conditions, resulting in a spatiotemporal trajectory sequence with individual identity.
[0062] Based on the spatial co-occurrence patterns and temporal interaction patterns among individuals in the spatiotemporal trajectory sequence, the strength of social associations between individual pairs is calculated, and the topology of the population social network is constructed.
[0063] In the aforementioned population social network topology, by analyzing the frequency of interactions, spatial proximity persistence, and behavioral synchronicity among individuals, and combining the temporal constraints of individual life cycle events, the kinship level among individuals is inferred, and a population phylogenetic tree structure is generated.
[0064] The feature vectors of the newly monitored individuals are backfilled into the individual feature embedding space for incremental updates, and the calculation rules for the strength of social associations are adaptively adjusted using the newly added kinship labeling.
[0065] In one optional implementation, an individual feature embedding space is constructed based on metric learning. The composite feature representation is mapped into the embedding space to form an individual feature vector. Cross-spatial individual re-identification is achieved by minimizing the feature vector distance of the same individual under different spatiotemporal conditions, resulting in a spatiotemporal trajectory sequence with individual identification, including:
[0066] Construct a training sample set of triplets, where each triplet contains anchor features, positive sample features, and negative sample features;
[0067] The metric learning network is trained based on the triplet training sample set. By minimizing the feature vector distance between the anchor feature and the positive sample feature in the embedding space, and simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space, the metric learning network learns to map the composite feature representation into individual feature vectors that are invariant to changes in illumination, viewpoint, and pose.
[0068] For an animal image to be identified, its composite feature representation is input into the trained metric learning network to obtain a query feature vector. The distance metric between the query feature vector and all individual feature vectors stored in the feature database is calculated. When the minimum distance metric satisfies the similarity judgment condition, the query feature vector is associated with the corresponding individual and the feature vector set of the corresponding individual is updated.
[0069] Spatial location information with the same individual association identifier is linked together in chronological order to form a spatiotemporal trajectory sequence.
[0070] In constructing the individual feature embedding space, the core driving force of metric learning comes from the rational organization of the triplet training sample set. Each triplet in the training sample set consists of three parts: anchor features, positive sample features, and negative sample features. Anchor features correspond to the composite feature expression extracted from a specific individual in a particular observation; positive sample features are composite feature expressions extracted from the same individual under different spatiotemporal conditions; and negative sample features are composite feature expressions from another different individual. The quality of triplet construction directly determines the generalization ability of the metric learning network. To improve training efficiency, a semi-difficult negative sample mining strategy is adopted: when selecting negative samples, samples that are close to the anchor features but still belong to different individuals are prioritized, rather than randomly selecting negative samples that are extremely far away. This forces the network to form a more refined discrimination boundary at the difficult boundary. For wildlife population monitoring scenarios, since the appearance of the same individual varies greatly in different seasons and behavioral states, the coverage of positive samples needs to encompass as many variations in lighting, viewing angle, and posture as possible to ensure sufficient intra-class compactness of the embedding space.
[0071] Based on the aforementioned triplet training sample set, the metric learning network optimizes its parameters by minimizing the feature vector distance between the anchor features and positive sample features in the embedding space, while simultaneously maximizing the feature vector distance between the anchor features and negative sample features in the embedding space. Specifically, let the embedding vector obtained after mapping the anchor features through the network be... The embedding vector corresponding to the positive sample feature is The embedding vector corresponding to the negative sample features is The training objective function then adopts the triplet loss form:
[0072] ;
[0073] in This represents the Euclidean distance metric function. The margin hyperparameter controls the minimum separation margin between positive and negative sample pairs in the embedding space. This is achieved by minimizing... The network is forced to map different observations of the same individual to neighboring regions in the embedding space, while mapping observations of different individuals to regions more than a certain distance from each other. The network backbone employs a deep convolutional neural network to extract high-level semantic features. A fully connected mapping layer is then added to the output layer to compress the features into fixed-dimensional embedding vectors. L2 normalization is applied to these embedding vectors to ensure that all individual feature vectors are distributed across a unit hypersphere, thus making the distance metrics numerically comparable. Data augmentation techniques, including random brightness perturbation, random cropping, and horizontal flipping, are introduced during training to enhance the network's invariance to changes in illumination, viewpoint, and pose.
[0074] After training, for the animal image to be identified, its composite feature representation is input into the trained metric learning network to obtain the corresponding query feature vector. The feature library pre-stores a set of feature vectors for each confirmed individual. Corresponding to a set of feature vectors ,in Represents an individual The number of accumulated feature vectors. Calculate the query feature vector. The distance metric between each individual and all stored feature vectors in the feature library is used, and the minimum value of the corresponding distance set for each individual is taken as the representative distance between that individual and the query sample. :
[0075] ;
[0076] Traverse all known individuals in the feature library and find those that make Individual index for obtaining the global minimum value ,Right now:
[0077] ;
[0078] When the minimum distance metric value Below the preset similarity threshold At that time, it was determined whether the animal or individual in the query image was real or artificial. For the same individual, the query feature vector will be used. Associated with individuals And add it to the individual The feature library is updated online using the feature vector set. If... Exceeding the threshold If the query image corresponds to a new individual that has not yet been recorded, a new individual identity identifier is assigned to it, and it is then identified as such. Establish a new set of feature vectors based on the initial feature vectors. Threshold The selection of the operating point needs to be determined in conjunction with the degree of appearance difference between individuals of a specific animal species. It is usually determined by plotting the receiver operating characteristic curve on the validation set and selecting the operating point that minimizes the sum of the false positive rate and the false negative rate.
[0079] Based on individual identity association, spatial location information with the same individual association identifier is concatenated in chronological order to form a spatiotemporal trajectory sequence. Spatial location information originates from geographic coordinates or camera deployment locations recorded during image acquisition. Combined with image acquisition timestamps, each valid identification event can be assigned both temporal and spatial annotations. For two identification records of the same individual at adjacent time points, if the time interval and displacement are within a biologically reasonable range, the two records are included in the same continuous trajectory segment; if the time interval is too long or the displacement exceeds a reasonable range, discontinuity markers are inserted into the trajectory sequence to indicate that there are observation gaps in that segment. The final spatiotemporal trajectory sequence uses individual identity identifiers as indexes to record the spatial location of the individual at each moment within the monitoring period, forming the foundational data for subsequent population social network analysis and kinship inference.
[0080] In practical deployments, the size of the feature database continues to grow with the extension of the monitoring cycle, posing a challenge to query efficiency. To maintain real-time recognition capabilities, the feature vector set of each individual in the feature database can be periodically clustered and compressed. A small number of cluster center vectors that can represent the range of appearance variations of that individual are retained to replace the original full set of feature vectors in distance calculations, thereby significantly reducing retrieval overhead while maintaining recognition accuracy. The number of cluster centers can be dynamically adjusted according to the number of historical observations of an individual. For individuals with fewer observations, all original feature vectors are retained, while for individuals with more than a certain number of observations, the cluster compression mechanism is activated to ensure that the storage and query efficiency of the feature database is always kept within an acceptable range.
[0081] In one optional implementation, the metric learning network is trained based on the triplet training sample set. Minimizing the feature vector distance between the anchor feature and the positive sample feature in the embedding space, while simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space, includes:
[0082] Construct a triplet loss function, which includes a positive pair distance term and a negative pair distance term. Quantize the feature vector distance between anchor features and positive sample features in the embedding space based on the positive pair distance term, and quantize the feature vector distance between anchor features and negative sample features in the embedding space based on the negative pair distance term.
[0083] Triple samples are extracted from the triple training sample set. The anchor features, positive sample features, and negative sample features of the triple samples are input into the metric learning network. The input features are mapped into feature vectors in the embedding space according to the metric learning network. The distance between the feature vectors is calculated and substituted into the triple loss function to obtain the values of the positive pair distance term and the negative pair distance term.
[0084] The gradient of the triplet loss function with respect to the network parameters of the metric learning network is calculated through backpropagation. The network parameters are then updated using the gradient, thereby decreasing the value of the positive pair distance term and increasing the value of the negative pair distance term.
[0085] The extraction, mapping, calculation, and update processes are repeatedly executed to minimize the feature vector distance between the anchor feature and the positive sample feature in the embedding space, while simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space.
[0086] The construction of the triplet loss function is a core step in training a metric learning network. In the embedding space, any triplet sample consists of three parts: an anchor sample, a positive sample, and a negative sample. The anchor sample and the positive sample come from images of the same individual acquired under different spatiotemporal conditions, while the negative sample comes from images of different individuals. By simultaneously constraining the positive and negative pair distances, the triplet loss function drives the metric learning network to cluster feature vectors of the same individual and push feature vectors of different individuals apart in the embedding space, thereby forming a discriminative distribution structure of individual features.
[0087] The triplet loss function consists of a positive pair distance term and a negative pair distance term. The positive pair distance term quantifies the Euclidean distance between anchor features and positive sample features in the embedding space, reflecting the similarity between feature vectors of the same individual. The negative pair distance term quantifies the Euclidean distance between anchor features and negative sample features in the embedding space, reflecting the difference between feature vectors of different individuals. The goal of the triplet loss function is to minimize the value of the positive pair distance term while maximizing the value of the negative pair distance term, and the two terms must satisfy a condition determined by the margin hyperparameter. The controlled safety margin constraint requires that the negative pair distance be at least greater than the positive pair distance. This results in a clear intra-class aggregation and inter-class separation structure within the embedding space. Specifically, the triplet loss function can be expressed as:
[0088] ;
[0089] in, The embedding vector is obtained after mapping the anchor features through a metric learning network. The embedding vector corresponding to the positive sample features. This is the embedding vector corresponding to the features of the negative samples. The Euclidean distance metric function is used. is the interval hyperparameter in the triplet loss. At this point, the triplet has already satisfied the interval constraint, and its contribution to the loss is zero; only when the above difference is greater than zero does the triplet generate gradient update-driven network parameters.
[0090] During the training data preparation phase, triplet samples are drawn in batches from the triplet training sample set. To improve training efficiency and convergence quality, a hard sample mining strategy is typically used for triplet extraction, that is, prioritizing the selection of triplet samples that meet the following criteria. Triples that are close to the margin boundary contribute most significantly to improving the network's discriminative ability and can effectively avoid the gradient vanishing problem caused by a large number of simple samples during training. The extracted triple samples contain three sets of original feature data: anchor point features, positive sample features, and negative sample features. These feature data come from the composite feature expression of phenotypic texture features and morphological contour features extracted in the previous step.
[0091] The anchor features, positive sample features, and negative sample features from the extracted triplet samples are input into a metric learning network. The metric learning network uses a deep convolutional neural network as its backbone, connecting a fully connected embedding layer after the feature extraction layer to compress and map the high-dimensional input features into a fixed-dimensional embedding space vector. The anchor features, positive sample features, and negative sample features are then forward-propagated through the same network structure with shared weights, each outputting its corresponding embedding vector. , and The embedding vector output by the network is usually processed... Normalization processes project all embedded vectors onto the unit hypersphere, thereby eliminating the interference of vector magnitudes during Euclidean distance calculations and improving the stability of the distance metric.
[0092] After completing the forward computation of the embedding vector, based on Calculate separately and The direct distance between them and and The negative pair distance between the positive and negative pairs is calculated, and the calculated positive and negative pair distance terms are substituted into the triplet loss function to obtain the loss value for the current batch. This loss value comprehensively reflects the degree of clustering of feature vectors of the same individual and the degree of separation of feature vectors of different individuals under the current network parameters. If there are multiple triples in a batch, the average loss value of all triples is taken as the overall training loss for that batch.
[0093] In obtaining batch loss Then, the gradient of the loss function with respect to all learnable network parameters in the metric learning network is calculated using the backpropagation mechanism. Backpropagation propagates gradient information layer by layer along the computation graph, from the embedding layer to the convolutional feature extraction layer, ensuring that the parameters of each network layer receive effective optimization signals. After calculating the gradient, stochastic gradient descent or its variants (such as the Adam optimizer) are used to update the network parameters. The parameter update direction decreases the value of the positive distance term, meaning that the embedding vectors of the same individual are closer to each other in space; simultaneously, it increases the value of the negative distance term, meaning that the embedding vectors of different individuals are farther apart in space. Each parameter update strictly follows the gradient direction, ensuring that the network's ability to discriminate individual features is improved after each training step.
[0094] The entire process of extracting triplet samples, forward mapping to obtain embedding vectors, calculating the distance between positive and negative pairs and substituting it into the loss function, and updating network parameters through backpropagation is repeatedly executed, constituting the iterative training loop of the metric learning network. As training rounds accumulate, the network parameters are continuously adjusted, and the feature vectors of the same individual in the embedding space gradually converge towards a dense cluster structure, while the spacing between feature vector clusters of different individuals continuously expands. During training, the convergence status of the network is dynamically monitored by periodically evaluating the ratio of the accuracy of individual re-identification to the intra-class and inter-class distances in the embedding space on the validation set. When the recognition performance on the validation set stabilizes and the loss value no longer decreases significantly, training is stopped and the network parameters are saved, completing the training of the metric learning network. A fully trained metric learning network possesses the ability to map image features of the same individual collected under different spatiotemporal conditions such as lighting, angle, and season to adjacent positions in the embedding space, providing a stable and reliable feature embedding foundation for subsequent cross-spatial individual re-identification and the construction of spatiotemporal trajectory sequences.
[0095] In one optional implementation, the social association strength between individual pairs is calculated based on the spatial co-occurrence pattern and temporal interaction pattern among individuals in the spatiotemporal trajectory sequence, and the population social network topology is constructed by including:
[0096] Spatiotemporal trajectory sequences are divided into spatiotemporal grids. Within each grid cell, the co-occurrence frequency and co-occurrence duration of individual pairs are counted. Stable spatial co-occurrence relationships are identified by analyzing the co-occurrence distribution patterns of individual pairs in multiple grid cells. These stable spatial co-occurrence relationships characterize that individual pairs have a continuous spatial clustering preference.
[0097] Based on the changes in the motion trajectory of individuals in the spatiotemporal trajectory sequence, proximity events and following events between individual pairs are extracted. The proximity events are identified by detecting the decreasing trend of the distance between individuals, and the following events are identified by detecting the temporal consistency of the individual's motion direction. The frequency and persistence of the proximity events and the following events are used as quantitative indicators of the temporal interaction pattern.
[0098] The strength value of the stable spatial co-occurrence relationship and the quantitative index of the temporal interaction pattern are weighted in multiple dimensions to generate a value of the social association strength between individual pairs.
[0099] Using all monitored individuals as network nodes, and the social association strength value as the edge weight between nodes, a weighted graph structure is constructed. In the weighted graph structure, connections with edge weights exceeding the association threshold are retained to form a population social network topology.
[0100] Using all monitored individuals as network nodes, and the social association strength values as edge weights between nodes, a weighted graph structure is constructed. In this weighted graph structure, connections with edge weights exceeding an association threshold are retained to form the population social network topology, including:
[0101] Each monitored individual is mapped to a network node in a weighted graph structure, and a node feature vector is constructed for each network node.
[0102] Calculate the social association strength between any two network nodes, and use the social association strength as the edge weight connecting the two network nodes to construct a fully connected weighted graph structure.
[0103] Statistical analysis is performed on the edge weight distribution in the fully connected weighted graph structure, and the dynamic association threshold is determined by calculating the mean and standard deviation of the edge weights;
[0104] Remove edge connections in the fully connected weighted graph structure whose edge weights are lower than the dynamic association threshold, and retain edge connections whose edge weights are higher than the dynamic association threshold to form a sparse population social network topology.
[0105] When performing spatiotemporal gridding on spatiotemporal trajectory sequences, the monitoring area is divided into several rectangular grid units according to a fixed spatial resolution, and the time axis is segmented according to fixed time windows, thus forming a three-dimensional spatiotemporal grid system. Within each grid unit, the occurrence records of all individuals are statistically analyzed, extracting the number of times any two individuals appear simultaneously in the same grid unit and the cumulative duration of their simultaneous residence within that grid unit. Co-occurrence information within a single grid unit only reflects local spatial clustering; therefore, further analysis of the co-occurrence distribution patterns of individual pairs across multiple grid units is required. If an individual pair exhibits a high frequency of co-occurrence and a long duration of co-occurrence in multiple geographically adjacent or functionally related grid units, then that individual pair is considered to have a stable spatial clustering preference, and this is identified as a stable spatial co-occurrence relationship. This stable spatial co-occurrence relationship can eliminate the interference of accidental encounters and truly reflect the spatial associations that individuals maintain continuously in their daily activities.
[0106] When extracting temporal interaction patterns, the motion trajectory of each individual is analyzed frame by frame, and the spatial distance between any two individuals is calculated over time. When a monotonically decreasing trend in the distance between individuals is detected over several consecutive time steps, and the distance reduction exceeds a preset proximity threshold, a proximity event is determined to have occurred within that time period. The frequency of proximity events and the duration of each proximity event together constitute a quantitative description of the proximity event. The identification of following events relies on the temporal consistency analysis of motion direction: the motion direction angle sequence of two individuals within a continuous time window is extracted, and the statistics of the difference between their direction angles are calculated. When the mean of the direction angle difference is consistently lower than the direction consistency threshold and the duration of this state exceeds the shortest following duration, a following event is determined to have occurred. Following events are further divided into two types: "active following" and "passive following." The former requires the movement start time of the subsequent individual to be later than that of the preceding individual, while the latter does not impose temporal constraints. The values of the four dimensions of proximity event frequency, average duration of proximity events, following event frequency, and average duration of following events are integrated into a quantitative index vector for temporal interaction patterns.
[0107] The calculation of social association strength involves a multi-dimensional weighted fusion of the strength value of stable spatial co-occurrence relationships and quantitative indicators of temporal interaction patterns. Let the individual's... The stable spatial co-occurrence intensity between them is The normalized value of the frequency of proximity events is The normalized value of the average duration of the near event is The event frequency normalization value is followed by the event frequency normalization value. The normalized value of the average duration of the following event is The strength of social connections Calculated using a weighted summation method, specifically expressed as follows: ,in , , , , These are the weight coefficients for the corresponding dimensions, and the sum of all weight coefficients is 1. The initial values of the weight coefficients can be set based on prior knowledge of the species' behavioral ecology; for example, for species with strong territoriality, the spatial co-occurrence weight... The weighting can be appropriately increased; for social species with obvious following behavior, the relevance weight of following events can be increased. , It can be appropriately increased. The normalization process uses the min-max normalization method to map each dimension's indicators to... Intervals are used to eliminate the influence of dimensional differences on the weighted results.
[0108] All monitored individuals are mapped to network nodes in a weighted graph structure. Each node is accompanied by a node feature vector, which contains the coordinates of the center of the individual's average activity range, a histogram of activity time distribution, and a representative feature vector corresponding to the individual's feature embedding space. The social association strength is calculated between any two network nodes. ,by As a connection node With nodes We construct a fully connected weighted graph structure using edge weights. The number of edges in the fully connected weighted graph is... ,in The graph structure is relatively dense when the population size is large, and directly using a fully connected graph for subsequent analysis will introduce a large number of noisy edges. Therefore, it is necessary to sparsify the graph structure.
[0109] Statistical analysis is performed on the distribution of all edge weights in a fully connected weighted graph structure, and the mean of the edge weights is calculated. with standard deviation This is used to determine the dynamic association threshold. The specific calculation method is as follows: ,in This is the threshold adjustment coefficient, whose value is adaptively determined based on population size and monitoring data density. When the population size is small or the monitoring duration is short, resulting in insufficient data, [further adjustments may be necessary]. Smaller values are chosen to retain more weakly correlated edges, preventing the network structure from becoming too sparse; when the population size is large and monitoring data is sufficient. A larger value is selected to filter out noisy edges and retain ecologically significant strong connections. The dynamic threshold mechanism is more adaptable than a fixed threshold, and can maintain reasonable network connectivity under different species and monitoring conditions.
[0110] Remove edge weights from a fully connected weighted graph that are below the dynamic association threshold. After connecting the edges, retain the edge weights that are higher than 1. Edge connections are used to form a sparsed population social network topology. In the sparsified network, the degree centrality, betweenness centrality, and clustering coefficient of each node are calculated to determine its role and status within the population's social structure. Individuals with high degree centrality are typically core members of the population and have strong associations with multiple individuals; local subgraphs with high clustering coefficients correspond to stable small groups or family units within the population. These topological features provide important structural constraints for subsequent kinship inference. The sparsified population social network topology is stored in both adjacency matrix and edge list formats for easy graph computation and visualization analysis.
[0111] In one optional implementation, within the population social network topology, by analyzing the frequency of interactions, spatial proximity persistence, and behavioral synchronicity among individuals, and combining this with the temporal constraints of individual lifecycle events, the kinship level among individuals is inferred, and a population phylogenetic tree structure is generated, including:
[0112] Extract individual pairs with edge connections from the topology of the population social network, calculate the cumulative value of interaction frequency, duration of spatial proximity and similarity of behavioral patterns of the individual pairs, and fuse the three into a kinship indicator value through weighted coefficients;
[0113] The first observation time and age stage marker of an individual are extracted as life cycle event records, and intergenerational time constraints are established based on the life cycle event records;
[0114] Individual pairs whose kinship indicator values exceed the judgment threshold and satisfy the generational time constraints are marked as kinship links. Kinship levels are determined based on the age differences between individuals. A population phylogenetic tree structure is constructed with individuals as nodes and kinship links as directed edges.
[0115] like Figure 2 As shown, the method includes:
[0116] After obtaining the topology of the population's social network, it is necessary to further mine information on kinship among individuals and ultimately construct a phylogenetic tree structure reflecting the population's reproductive history and bloodline inheritance. This process begins by extracting pairs of individuals with edge connections from the established social network—that is, combinations of individuals who have had social interaction records during the monitoring period. For each valid pair of individuals… Three types of behavioral indicators were statistically analyzed: cumulative interaction frequency, duration of spatial proximity, and similarity of behavioral patterns.
[0117] The cumulative interaction frequency reflects the total number of direct interactions between two individuals within the entire monitoring time window, including close-range interactions detectable by visual recognition systems, such as contact, following, and synchronized foraging. Spatial proximity duration measures the cumulative time two individuals remain in a spatially proximate state, recorded in minutes or hours, and normalized for differences in sampling density across different monitoring periods to avoid statistical bias caused by uneven sampling frequency. Behavioral pattern similarity is calculated by extracting behavioral sequences of two individuals within the same time period and using sequence alignment methods to calculate the correlation coefficient of their behavioral state transition patterns, with values normalized to a range of... The higher the value in the interval, the closer the daily behavioral rhythms of the two individuals are.
[0118] The above three types of indicators are combined using weighted coefficients to form a kinship indicator value. The specific calculation method is as follows:
[0119] ;
[0120] in, For individuals The score is the normalized cumulative value of interaction frequency. This is the normalized value of the duration of spatial proximity. This represents the behavioral pattern similarity value. , , These are the corresponding weighting coefficients, and the sum of the three is 1. The weighting coefficients can be set a priori based on the species' ecological habits; for example, for mammals with strong maternal attachment behavior, The weight of this should be appropriately increased to more fully capture the persistent spatial accompaniment characteristics between parents and offspring.
[0121] While extracting kinship indicator values, it is also necessary to establish an individual's life cycle event record. For each identified individual, the timestamp of its first appearance within the monitoring range is extracted from the spatiotemporal trajectory sequence and recorded as the first observation time. Based on observable phenotypic features such as body size, coat maturity, and tooth wear in the images, individuals were categorized into four age stages: juvenile, sub-adult, adult, and old age, each labeled with an integer. This indicates that, for some species, the accuracy of age group determination can be further refined by combining behavioral characteristics during the breeding season.
[0122] Based on the aforementioned life cycle event records, intergenerational time-series constraints are established. The core logic of these constraints is that the first observation of a parent individual must be earlier than that of an offspring individual, and the age difference between the two must satisfy the biologically feasible reproductive interval range for the species. Specifically, if an individual... Inferred to be an individual For a parent to meet the following two conditions: ,and ,in The minimum intergenerational age stage difference, typically one to two stages, is set based on the species' sexual maturity age. For individual pairs that do not meet the above temporal constraints, even their kinship indicator value... A high level of kinship cannot be labeled as a parent-offspring kinship connection, but should be classified as a strong social connection between individuals of the same generation.
[0123] After calculating the kinship indicator value and establishing intergenerational time constraints, kinship connections are screened and determined. A kinship determination threshold is set. Only when an individual... Kinship indicator value If the intergenerational time sequence constraint is also met, the individual pair is marked as a valid kinship connection. Threshold The value can be adaptively adjusted according to the population size and monitoring data density: when monitoring data is sufficient and individual behavior records are complete, it can be appropriately increased. To reduce the false positive rate; when data is sparse in the early monitoring stage, the error rate can be appropriately reduced. It is also verified in conjunction with a manual review mechanism.
[0124] For pairs of individuals marked as kinship-linked, the kinship level is further determined based on their age differences. Pairs with an age difference of 2 or more are inferred to be direct parents, corresponding to a kinship level of 1. Pairs with an age difference of 1 and whose first observation time difference falls within the species' reproductive cycle are inferred to be parents or grandparents / grandchildren, requiring further differentiation based on specific time-series information. Pairs of individuals with the same or similar age but higher kinship indices are inferred to be siblings, corresponding to a kinship level of 2. Pairs with kinship indices in the intermediate range and small age differences are inferred to be cousins or more distant relatives, corresponding to a kinship level of 3. This refinement of kinship levels helps distinguish the impact of different strengths of genetic association on population behavioral structure in subsequent population dynamics analysis.
[0125] After labeling all valid kinship connections, a phylogenetic tree structure is constructed, using each identified individual in the population as a node and valid kinship connections as directed edges, with the edges pointing from parents to offspring. For offspring individuals with multiple parental candidates, the candidate with the highest kinship indicator value and satisfying all constraints is selected as the primary parent node and connected. The remaining candidate parents are retained in the phylogenetic tree as secondary edges for subsequent analysis. Individuals whose parental origin cannot be determined are treated as the root node of the phylogenetic tree and marked as the population's foundational individual.
[0126] The final generated phylogenetic tree structure is stored in the form of a directed acyclic graph. Node attributes include individual identification, age group marker, first observation time, and sex information. Edge attributes include kinship level, kinship indicator value, and confidence level for determining effective kinship connections. This phylogenetic tree structure can be directly used for population genetic diversity analysis, reproductive success rate statistics, and parameter estimation of population growth models, providing a quantitative basis for wildlife conservation and management decisions. With the continuous accumulation of monitoring data, the phylogenetic tree structure supports incremental updates. Newly identified individuals can be dynamically incorporated into the phylogenetic structure based on their feature vectors and social association patterns with existing individuals, enabling continuous tracking and recording of the population's reproductive history.
[0127] In one optional implementation, incrementally updating the feature vector of the newly monitored individual by backfilling it into the individual feature embedding space, and adaptively adjusting the calculation rules for the strength of social association using the newly added kinship annotations, includes:
[0128] When a new monitored individual is identified, the composite feature expression of the new monitored individual is extracted and mapped to a new feature vector through a metric learning network. The distance metric matrix between the new feature vector and the feature vectors already stored in the individual feature embedding space is calculated. Based on the distance metric matrix, the influence of the new feature vector on the topology of the individual feature embedding space is analyzed. The new feature vector is then backfilled into the individual feature embedding space to complete the incremental update.
[0129] New kinship labels are extracted from the phylogenetic tree structure of the population. The new kinship labels include pairs of individuals with verified kinship. Observational data of the pairs of individuals are extracted in three dimensions: interaction frequency, spatial proximity persistence, and behavioral synchronicity.
[0130] The distribution characteristics of the observed data in three dimensions are statistically analyzed, and the differences between the distribution characteristics and those of unrelated individuals in the three dimensions are analyzed to quantify the contribution of each dimension to kinship identification. Based on the contribution, the weight ratios of the interaction frequency dimension, spatial proximity persistence dimension, and behavioral synchronicity dimension in the calculation rule of social association strength are adjusted, and the calculation rule of social association strength is adaptively adjusted using the adjusted weight ratios.
[0131] When a new monitored individual is identified, a composite feature representation is extracted from image frames of that individual collected at different spatiotemporal nodes. This composite feature representation integrates phenotypic texture features and morphological contour features. The composite feature representation is then input into a trained metric learning network to obtain a corresponding new feature vector. After completing the mapping, calculate... Construct a distance metric matrix by considering the distances between the vectors and all stored feature vectors in the individual feature embedding space. ,in This represents the total number of feature vectors stored in the current embedding space. Each element in the matrix corresponds to... The Euclidean distance value between the vector and a stored vector.
[0132] Based on distance metric matrix Analyze the impact of the new eigenvectors on the topological structure of the embedding space, specifically, with Centered on the target area, statistics are collected within a preset neighborhood radius. Number of stored feature vectors within the range And the set of individual identity identifiers to which these neighborhood vectors belong. Smaller values and the fact that the vectors in the neighborhood belong to multiple different individuals indicate that... Falling into the sparse boundary region of the embedded space causes minimal disturbance to the existing topology and can be directly... Insert into the embedded space and assign a new individual identity identifier. If If a new monitored individual has a large vector and its neighborhood vectors are highly concentrated around a known individual, cross-validation with the cross-temporal individual re-identification process is triggered to determine whether the new monitored individual is a recurrence of the known individual at a new spatiotemporal node. After topological impact analysis, The data is then backfilled into the embedding space, completing the incremental update. The incremental update uses an append-only approach, without modifying existing feature vectors, thus ensuring the consistency of historical trajectory data.
[0133] After incrementally updating the embedding space, new phylogenetic labels are extracted from the population phylogenetic tree structure. These new phylogenetic labels consist of two parts: first, individual pairs that have been inferred through lifecycle event time-series constraints and meet the confidence threshold; and second, individual pairs that have been validated through subsequent manual verification or external biological sample analysis. For each validated phylogenetic pair... Three dimensions of observational data were extracted from spatiotemporal trajectory sequences with individual identifiers: cumulative observations of interaction frequency, cumulative observations of spatial proximity persistence, and similarity observations of behavioral synchronicity. These three dimensions of observational data are used in kinship inference. , , The indicators have corresponding meanings, but in this context they are used as observation records of verified samples for statistical analysis, rather than for calculating kinship.
[0134] The observational data of all verified kinship pairs across the three dimensions were aggregated into a positive kinship sample set. Simultaneously, pairs of individuals confirmed to be unrelated were extracted from the population phylogenetic tree to form a negative non-kinship sample set. The distribution characteristics across the three dimensions—mean, variance, and skewness coefficient—were calculated for both the positive and negative sample sets. Taking interaction frequency as an example, if the mean of the positive sample set in this dimension is significantly higher than that of the negative sample set, and the overlap between their distributions is small, it indicates that interaction frequency has a high discriminative contribution to kinship identification.
[0135] When quantifying the contribution of each dimension of features to kinship identification, a discriminative power evaluation method based on distribution differences is adopted. For the first... Dimensions ( The discriminant score is obtained by calculating the ratio of the mean difference between the positive and negative sample sets to the pooled standard deviation in this dimension. Its expression is:
[0136] ;
[0137] in, and The positive sample set is respectively in the th The mean and standard deviation of the dimension. and The negative sample set is in the th Mean and standard deviation of the dimension. The larger the value, the stronger the ability of that dimension to distinguish between related and unrelated individuals.
[0138] Based on the discriminative power scores of each dimension, the weights of the interaction frequency dimension, spatial proximity persistence dimension, and behavioral synchronicity dimension in the social association strength calculation rule are adaptively adjusted. The adjustment method involves normalizing the discriminative power scores of the three dimensions to obtain the updated weight allocation. :
[0139] ;
[0140] in , , The updated weights correspond to the three dimensions of interaction frequency, spatial proximity persistence, and behavioral synchronicity, respectively. The updated weight ratios replace the original weight coefficients in the social association strength calculation rules, enabling the calculation of social association strength to more accurately reflect the impact of kinship on inter-individual interaction patterns in subsequent monitoring periods.
[0141] To avoid excessive oscillations in weight adjustments due to insufficient sample size for single kinship labeling, a smoothing mechanism is introduced during weight updates. Let the number of currently accumulated positive kinship pairs be... ,when Below the preset confidence sample size threshold At that time, a weighted fusion method is used to combine the newly calculated Compared with the previous round of weighting Mix:
[0142] ;
[0143] when At that time, directly adopt As an update weight, there is no need for smooth fusion. This mechanism ensures that the weight ratio can transition smoothly when the accumulation of kinship samples is insufficient in the early stage of monitoring, and avoids drastic fluctuations in the calculation rules caused by a small number of abnormal samples.
[0144] After completing the adaptive adjustment of the weight allocation, the updated weights will be... The calculation rules for social association strength are synchronously written into the configuration, and the weights of affected edges in the constructed population social network topology are recalculated. The recalculation only applies to edges involving individuals related to newly added kinship labels; the weights of edges not involving individuals remain unchanged, thus achieving localized refinement updates to the population social network while maintaining computational efficiency. Each adaptive adjustment triggered by a new kinship label records version information, including the adjustment timestamp, the sample size involved in the statistics, and the discriminative power scores for each dimension, facilitating subsequent backtracking analysis and model iteration optimization.
[0145] The method further includes:
[0146] Animal image sequences are acquired from multiple visual acquisition devices deployed at different locations and altitudes within the monitoring area. These devices include ground-mounted high-resolution cameras, adjustable-angle gimbal cameras, infrared trigger cameras, and aerial cameras mounted on drones. The resolution of each acquisition device is no less than 1920×1080 pixels, with a frame rate set at 25 to 30 frames per second under visible light conditions. Under low-light conditions, the device switches to thermal infrared imaging mode and adjusts the frame rate to 10 to 15 frames per second to reduce power consumption. Each acquisition device simultaneously records metadata information such as shooting timestamp, geographic coordinates, altitude, device orientation angle, and ambient light intensity, embedding this information into the EXIF field of the image file. When the image sequence is transmitted to the edge computing node, a quality assessment is first performed. The Laplacian variance of each frame is calculated as a sharpness indicator. When this indicator is below a preset threshold, the frame is marked as blurry and its subsequent processing priority is reduced. Simultaneously, image contrast and color saturation are calculated. When the contrast is below a set lower limit or the saturation is close to zero, an exposure anomaly is identified, triggering an automatic exposure parameter adjustment mechanism on the device.
[0147] Animal target detection is performed on qualified image sequences. A deep learning-based target detection network is used to extract animal regions from the images. The input layer of the detection network receives a normalized RGB three-channel image. After multiple convolution and pooling operations, candidate regions are generated on the feature map. A region proposal network is used to filter candidate boxes with a confidence score higher than 0.6. Each candidate box is classified and bounding box regression is performed to obtain the animal target's category label and precise location coordinates. Natural biomarker features are extracted from the detected animal target regions. These natural biomarker features include phenotypic texture features and morphological contour features. Phenotypic texture features are extracted by applying the local binary pattern operator and the histogram of oriented gradients operator to the target region. The target region is divided into several 16×16 pixel grid cells. The gradient magnitude distribution in eight directions is statistically analyzed in each grid cell and concatenated to form a texture description vector. Morphological contour features are extracted by first performing edge detection on the target region to obtain a binary contour image. Then, the Fourier descriptor of the contour is calculated as a shape representation. The first 32 low-frequency coefficients of the Fourier descriptor are taken as the morphological feature vector to ensure scale and rotation invariance. The phenotypic texture feature vector and the morphological contour feature vector are weighted and concatenated with a weight ratio of 3 to 2 to form a composite feature representation. The composite feature representation has a dimension of 512, and the numerical range of each dimension is mapped to 0 to 1 through max-min normalization.
[0148] An individual feature embedding space is constructed to map composite feature representations into discriminative individual feature vectors. This embedding space is obtained through training a metric learning network, which consists of four fully connected layers with 512, 256, 128, and 64 neurons respectively. The activation function is a modified linear unit (MLU). During training, triplet training samples are prepared. Each triplet contains an anchor sample, a positive sample, and a negative sample. The anchor sample and positive sample come from images of the same individual observed at different times or from different perspectives, while the negative sample comes from images of different individuals. The triplet loss function includes a positive pair distance term and a negative pair distance term. The positive pair distance term calculates the Euclidean distance between the anchor feature vector and the positive sample feature vector in the embedding space, and the negative pair distance term calculates the Euclidean distance between the anchor feature vector and the negative sample feature vector in the embedding space. A margin constraint parameter of 0.5 is set, requiring the difference between the positive and negative pair distance terms to be greater than this margin constraint parameter. The weight parameters of the metric learning network are iteratively updated using a gradient descent algorithm until the triplet loss function converges. After training, the newly observed animal composite feature representations are input into the metric learning network. After forward propagation, a 64-dimensional individual feature vector is obtained. This feature vector is compared with the stored individual feature vectors in the embedding space. The cosine similarity between the new feature vector and each known individual feature vector in the database is calculated. When the maximum similarity exceeds 0.85, it is determined to be a known individual and the corresponding individual identity is output. When the maximum similarity is less than 0.85, it is determined to be a new individual and a new identity is assigned to it. At the same time, the feature vector is stored in the individual feature database.
[0149] Based on individual identity identifiers, observation records of the same individual at different times and spatial locations are linked. Each observation record includes a timestamp, geographic coordinates, individual identity identifier, and behavioral status annotation. The observation records of the same individual are arranged in chronological order to form a spatiotemporal trajectory sequence. The social association strength between individual pairs is calculated based on this spatiotemporal trajectory sequence. This social association strength comprehensively considers three dimensions: interaction frequency, spatial proximity persistence, and behavioral synchronicity. Interaction frequency is obtained by counting the number of times two individuals co-occur within a time window (set to 7 consecutive days). A co-occurrence is defined as when the spatial distance between two individuals within the same time period is less than 50 meters. Spatial proximity persistence is obtained by calculating the cumulative duration of two individuals maintaining spatial proximity during the observation period. Maintaining proximity is considered when the distance between two individuals is less than a threshold at multiple consecutive time points. Behavioral synchronicity is obtained by comparing the behavioral status annotations of two individuals within the same time period. A behavioral synchronization event is defined as when the behavioral status annotations of two individuals are identical. The values of the three dimensions of interaction frequency, spatial proximity persistence, and behavioral synchronicity are normalized to between 0 and 1, and then weighted and summed according to the weight coefficients of 0.4, 0.35, and 0.25 to obtain the value of social association strength.
[0150] A population social network topology is constructed, with each identified animal individual as a network node and the strength of the social association between individual pairs as the weight of the edges. When the social association strength exceeds 0.3, a connection edge is established between the corresponding two nodes, forming a weighted undirected graph representation of the social network. Further, the kinship level between individuals is inferred from the social network topology. First, confirmed kinship pairs are extracted from the population phylogenetic tree structure as reference samples. Feature values of these reference samples are extracted in three dimensions: interaction frequency, spatial proximity persistence, and behavioral synchronicity. The distribution range of feature values corresponding to different kinship levels is statistically analyzed. For example, the median interaction frequency of mother-child pairs is 18 times per week, the median spatial proximity persistence is 0.82, and the median behavioral synchronicity is 0.76, while the corresponding medians for non-related pairs are 3 times per week, 0.21, and 0.33, respectively. For the individual pair to be determined, its observed values in three dimensions are extracted and matched with the feature distribution of each kinship level. The probability of the individual pair belonging to each kinship level is calculated. When the probability of belonging to a certain kinship level exceeds 0.7, the individual pair is determined to have the corresponding kinship.
[0151] The kinship inference results are validated by incorporating the temporal constraints of individual life cycle events. The first observation time of an individual is extracted from monitoring records as an approximation of the life cycle starting point, and records from the individual's juvenile stage are extracted as the basis for age determination. When a mother-child relationship is determined between two individuals, the difference between their first observation times is checked to see if it conforms to the reproductive cycle constraints of the species. For example, if the gestation period of a species is 8 months, if the difference between the first observation time of a candidate mother-child pair in juvenile and adult individuals is less than 8 months, the kinship inference is deemed invalid and corrected to a non-related relationship. If the time difference is between 8 and 24 months, the kinship determination result is retained. Based on the validated kinship inference results, a population phylogenetic tree structure is constructed. The phylogenetic tree is stored in a tree graph data structure, where each node represents an individual and records the individual's identity, sex, estimated age, and number of observation records. The edges connecting nodes represent kinship relationships and are labeled with the kinship type and confidence level.
[0152] When a new individual is identified in the monitoring area, its composite feature representation is extracted and mapped to a new feature vector through a metric learning network. The distance metric matrix between the new feature vector and the feature vectors already stored in the individual feature embedding space is calculated. The rows of this matrix correspond to the new feature vector, the columns correspond to the known individual feature vectors in the database, and the matrix elements are the Euclidean distances between corresponding feature vector pairs. The minimum distance value and the variance of the distance distribution in the distance metric matrix are analyzed. When the minimum distance value is greater than a set threshold of 1.2, the new individual is determined to be significantly different from all known individuals. The new feature vector is then backfilled into the individual feature embedding space, and the index structure of the embedding space is updated to support subsequent fast retrieval.
[0153] Simultaneously, newly added kinship labels are extracted from the population phylogenetic tree structure. These newly added kinship labels are derived from kinship confirmed through genetic sampling or long-term behavioral observation. Historical observation data of these newly confirmed kinship pairs are extracted in three dimensions: interaction frequency, spatial proximity persistence, and behavioral synchronicity. The new data is merged with existing reference samples, and the distinguishing ability of each dimension feature for kinship identification is recalculated. The distinguishing ability is quantified by calculating the ratio of inter-class variance to intra-class variance of each dimension feature between kinship groups and non-kinship groups. The weight coefficients of the three dimensions in the social association strength calculation rule are adjusted according to the updated distinguishing ability values. If the distinguishing ability of the interaction frequency dimension is improved, its weight coefficient is increased, and the weight coefficients of other dimensions are correspondingly decreased to ensure that the sum of the three weight coefficients is 1. The social association strength of all individual pairs in the population is recalculated using the adjusted weight coefficients, and the social network topology and phylogenetic tree structure are updated.
[0154] A second aspect of the present invention provides a non-contact animal population monitoring system based on AI vision, comprising:
[0155] The feature extraction unit is used to acquire animal image sequences collected from multiple spatiotemporal nodes, and extract natural biological marker features of individuals from the animal image sequences. The natural biological marker features include a composite feature expression of phenotypic texture features and morphological contour features.
[0156] The individual identification unit is used to construct an individual feature embedding space based on metric learning, and to map the composite feature expression into the embedding space to form an individual feature vector. By minimizing the feature vector distance of the same individual under different spatiotemporal conditions, cross-spatiotemporal individual re-identification is achieved, resulting in a spatiotemporal trajectory sequence with individual identity identifier.
[0157] The network construction unit is used to calculate the social association strength between individual pairs based on the spatial co-occurrence pattern and temporal interaction pattern among individuals in the spatiotemporal trajectory sequence, and to construct the population social network topology.
[0158] The phylogenetic inference unit is used to infer the kinship level between individuals and generate a population phylogenetic tree structure by analyzing the frequency of interaction, spatial proximity persistence and behavioral synchronicity between individuals, and combining the temporal constraints of individual life cycle events in the topology of the population social network.
[0159] The incremental update unit is used to backfill the feature vector of the newly monitored individual into the individual feature embedding space for incremental update, and to adaptively adjust the calculation rules of the social association strength using the newly added kinship label.
[0160] A third aspect of the present invention provides an electronic device, comprising:
[0161] processor;
[0162] Memory used to store processor-executable instructions;
[0163] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0164] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0165] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A non-contact animal population monitoring method based on AI vision, characterized in that, include: The process involves acquiring animal image sequences from multiple spatiotemporal nodes, extracting natural biomarker features of individuals from the animal image sequences, and including composite feature expressions of phenotypic texture features and morphological contour features. An individual feature embedding space is constructed based on metric learning. The composite feature expression is mapped to the embedding space to form an individual feature vector. Cross-temporal individual re-identification is achieved by minimizing the feature vector distance of the same individual under different spatiotemporal conditions, resulting in a spatiotemporal trajectory sequence with individual identity. Based on the spatial co-occurrence patterns and temporal interaction patterns among individuals in the spatiotemporal trajectory sequence, the strength of social associations between individual pairs is calculated, and the topology of the population social network is constructed. In the aforementioned population social network topology, by analyzing the frequency of interactions, spatial proximity persistence, and behavioral synchronicity among individuals, and combining the temporal constraints of individual life cycle events, the kinship level among individuals is inferred, and a population phylogenetic tree structure is generated. The feature vectors of the newly monitored individuals are backfilled into the individual feature embedding space for incremental updates, and the calculation rules for the strength of social associations are adaptively adjusted using the newly added kinship labeling.
2. The method according to claim 1, characterized in that, An individual feature embedding space is constructed based on metric learning. The composite feature representation is mapped to the embedding space to form an individual feature vector. Cross-spatial-temporal individual re-identification is achieved by minimizing the feature vector distance of the same individual under different spatiotemporal conditions, resulting in a spatiotemporal trajectory sequence with individual identity identifiers, including: Construct a training sample set of triplets, where each triplet contains anchor features, positive sample features, and negative sample features; The metric learning network is trained based on the triplet training sample set. By minimizing the feature vector distance between the anchor feature and the positive sample feature in the embedding space, and simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space, the metric learning network learns to map the composite feature representation into individual feature vectors that are invariant to changes in illumination, viewpoint, and pose. For an animal image to be identified, its composite feature representation is input into the trained metric learning network to obtain a query feature vector. The distance metric between the query feature vector and all individual feature vectors stored in the feature database is calculated. When the minimum distance metric satisfies the similarity judgment condition, the query feature vector is associated with the corresponding individual and the feature vector set of the corresponding individual is updated. Spatial location information with the same individual association identifier is linked together in chronological order to form a spatiotemporal trajectory sequence.
3. The method according to claim 2, characterized in that, The metric learning network is trained based on the triplet training sample set. This is achieved by minimizing the feature vector distance between the anchor feature and the positive sample feature in the embedding space, while simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space. Construct a triplet loss function, which includes a positive pair distance term and a negative pair distance term. Quantize the feature vector distance between anchor features and positive sample features in the embedding space based on the positive pair distance term, and quantize the feature vector distance between anchor features and negative sample features in the embedding space based on the negative pair distance term. Triple samples are extracted from the triple training sample set. The anchor features, positive sample features, and negative sample features of the triple samples are input into the metric learning network. The input features are mapped into feature vectors in the embedding space according to the metric learning network. The distance between the feature vectors is calculated and substituted into the triple loss function to obtain the values of the positive pair distance term and the negative pair distance term. The gradient of the triplet loss function with respect to the network parameters of the metric learning network is calculated through backpropagation. The network parameters are then updated using the gradient, thereby decreasing the value of the positive pair distance term and increasing the value of the negative pair distance term. The extraction, mapping, calculation, and update processes are repeatedly executed to minimize the feature vector distance between the anchor feature and the positive sample feature in the embedding space, while simultaneously maximizing the feature vector distance between the anchor feature and the negative sample feature in the embedding space.
4. The method according to claim 1, characterized in that, Based on the spatial co-occurrence patterns and temporal interaction patterns among individuals in the spatiotemporal trajectory sequence, the social association strength between individual pairs is calculated, and the population social network topology is constructed, including: Spatiotemporal trajectory sequences are divided into spatiotemporal grids. Within each grid cell, the co-occurrence frequency and co-occurrence duration of individual pairs are counted. Stable spatial co-occurrence relationships are identified by analyzing the co-occurrence distribution patterns of individual pairs in multiple grid cells. These stable spatial co-occurrence relationships characterize that individual pairs have a continuous spatial clustering preference. Based on the changes in the motion trajectory of individuals in the spatiotemporal trajectory sequence, proximity events and following events between individual pairs are extracted. The proximity events are identified by detecting the decreasing trend of the distance between individuals, and the following events are identified by detecting the temporal consistency of the individual's motion direction. The frequency and persistence of the proximity events and the following events are used as quantitative indicators of the temporal interaction pattern. The strength value of the stable spatial co-occurrence relationship and the quantitative index of the temporal interaction pattern are weighted in multiple dimensions to generate a value of the social association strength between individual pairs. Using all monitored individuals as network nodes, and the social association strength value as the edge weight between nodes, a weighted graph structure is constructed. In the weighted graph structure, connections with edge weights exceeding the association threshold are retained to form a population social network topology.
5. The method according to claim 4, characterized in that, Using all monitored individuals as network nodes, and the social association strength values as edge weights between nodes, a weighted graph structure is constructed. In this weighted graph structure, connections with edge weights exceeding an association threshold are retained to form the population social network topology, including: Each monitored individual is mapped to a network node in a weighted graph structure, and a node feature vector is constructed for each network node. Calculate the social association strength between any two network nodes, and use the social association strength as the edge weight connecting the two network nodes to construct a fully connected weighted graph structure. Statistical analysis is performed on the edge weight distribution in the fully connected weighted graph structure, and the dynamic association threshold is determined by calculating the mean and standard deviation of the edge weights; Remove edge connections in the fully connected weighted graph structure whose edge weights are lower than the dynamic association threshold, and retain edge connections whose edge weights are higher than the dynamic association threshold to form a sparse population social network topology.
6. The method according to claim 1, characterized in that, In the aforementioned population social network topology, by analyzing the frequency of interactions, spatial proximity persistence, and behavioral synchronicity among individuals, and combining this with the temporal constraints of individual life cycle events, the kinship level among individuals is inferred, generating a population phylogenetic tree structure including: Extract individual pairs with edge connections from the topology of the population social network, calculate the cumulative value of interaction frequency, duration of spatial proximity and similarity of behavioral patterns of the individual pairs, and fuse the three into a kinship indicator value through weighted coefficients; The first observation time and age stage marker of an individual are extracted as life cycle event records, and intergenerational time constraints are established based on the life cycle event records; Individual pairs whose kinship indicator values exceed the judgment threshold and satisfy the generational time constraints are marked as kinship links. Kinship levels are determined based on the age differences between individuals. A population phylogenetic tree structure is constructed with individuals as nodes and kinship links as directed edges.
7. The method according to claim 1, characterized in that, Incremental updates are performed by backfilling the feature vectors of newly monitored individuals into the individual feature embedding space, and adaptive adjustments are made to the calculation rules for the strength of social associations using newly added kinship annotations, including: When a new monitored individual is identified, the composite feature expression of the new monitored individual is extracted and mapped to a new feature vector through a metric learning network. The distance metric matrix between the new feature vector and the feature vectors already stored in the individual feature embedding space is calculated. Based on the distance metric matrix, the influence of the new feature vector on the topology of the individual feature embedding space is analyzed. The new feature vector is then backfilled into the individual feature embedding space to complete the incremental update. New kinship labels are extracted from the phylogenetic tree structure of the population. The new kinship labels include pairs of individuals with verified kinship. Observational data of the pairs of individuals are extracted in three dimensions: interaction frequency, spatial proximity persistence, and behavioral synchronicity. The distribution characteristics of the observed data in three dimensions are statistically analyzed, and the differences between the distribution characteristics and those of unrelated individuals in the three dimensions are analyzed to quantify the contribution of each dimension to kinship identification. Based on the contribution, the weight ratios of the interaction frequency dimension, spatial proximity persistence dimension, and behavioral synchronicity dimension in the calculation rule of social association strength are adjusted, and the calculation rule of social association strength is adaptively adjusted using the adjusted weight ratios.
8. A non-contact animal population monitoring system based on AI vision, used to implement the method as described in any one of claims 1-7, characterized in that, include: The feature extraction unit is used to acquire animal image sequences collected from multiple spatiotemporal nodes, and extract natural biological marker features of individuals from the animal image sequences. The natural biological marker features include a composite feature expression of phenotypic texture features and morphological contour features. The individual identification unit is used to construct an individual feature embedding space based on metric learning, map the composite feature expression to the embedding space to form an individual feature vector, and achieve cross-temporal individual re-identification by minimizing the feature vector distance of the same individual under different spatiotemporal conditions, thereby obtaining a spatiotemporal trajectory sequence with individual identity identifier; The network construction unit is used to calculate the social association strength between individual pairs based on the spatial co-occurrence pattern and temporal interaction pattern between individuals in the spatiotemporal trajectory sequence, and to construct the population social network topology. The phylogenetic inference unit is used to infer the kinship level between individuals and generate a population phylogenetic tree structure by analyzing the frequency of interaction, spatial proximity persistence and behavioral synchronicity between individuals, and combining the temporal constraints of individual life cycle events in the topology of the population social network. The incremental update unit is used to backfill the feature vector of the newly monitored individual into the individual feature embedding space for incremental update, and to adaptively adjust the calculation rules of the social association strength using the newly added kinship label.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.