Wild animal ecological insight and behavior analysis method, electronic equipment and storage medium
By employing dynamic primitive decomposition and occlusion repair techniques, combined with feature space construction and 3D reconstruction, the problems of data fragmentation and occlusion in wildlife monitoring have been solved. This enables the precise quantification of the 3D physiological parameters and subtle behaviors of individual wild animals, supporting intelligent assessment of population health risks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN EFERCRO ELECTRONIC TECHNOLOGY CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies cannot effectively link observation data of the same animal at different times and from different angles in wildlife monitoring, resulting in data fragmentation and severe obstruction, which makes it impossible to achieve refined management of individual health status and early disease warning.
By using dynamic primitive decomposition, occlusion extrapolation and repair, a feature space is constructed for individual identity re-identification and cross-temporal association. Combined with 3D reconstruction and kinematic models, the quantification of 3D physiological parameters and detection of subtle behavioral anomalies of wild animal individuals are realized.
It enables continuous tracking of individual wild animal identities, precise quantification of three-dimensional physiological parameters, and accurate detection of subtle behaviors, allowing for intelligent assessment of population health risks.
Smart Images

Figure CN121862384A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wildlife monitoring and ecological protection technology, specifically to a method for wildlife ecological insight and behavior analysis, electronic equipment, and storage medium. Background Technology
[0002] Passive monitoring based on infrared-triggered cameras is currently the primary method for wildlife surveillance. However, the complexities of the wild environment and the concealment of animal activities result in highly fragmented, discontinuous, and severely obscured monitoring data. Existing technologies are largely limited to species identification and population statistics, failing to effectively correlate observation data of the same animal at different times and angles. Furthermore, it is difficult to recover continuous and accurate individual three-dimensional physiological parameters and subtle movement characteristics from fragmented images. This leads to managers obtaining only scattered snapshots, unable to construct complete individual growth profiles, severely limiting the ability to conduct refined management of individual wildlife health and provide early disease warnings.
[0003] Therefore, there is an urgent need for a technical solution that can utilize fragmented and heavily occluded monocular video data to achieve continuous tracking of individual wildlife identities, precise quantification of three-dimensional physiological parameters, detection of subtle behavioral anomalies, and intelligent assessment of population health risks. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides a method, electronic device, and storage medium for wildlife ecological insight and behavioral analysis, which enables automated, quantitative, and three-dimensional analysis of the health status of individual wild animals under complex field conditions.
[0005] According to a first aspect of this application, a method for wildlife ecological insight and behavioral analysis is provided, comprising the following steps: S1. Process the acquired wildlife video stream, and obtain a denoised dynamic 3D primitive sequence and target contour image through dynamic primitive decomposition, occlusion extrapolation and repair. S2. Based on the target contour image, extract and encode biological features, and construct a feature space through triplet constraints to achieve individual identity re-identification and cross-temporal association; S3. Based on the dynamic three-dimensional primitive sequence and identity identifier, through feature decoupling, fusion and three-dimensional reconstruction, drive the parametric mesh model to fit the body shape, and perform virtual measurement to quantify body size parameters and body condition score; S4. Based on the three-dimensional skeleton sequence, a kinematic model is constructed. After the sparse motion data is completed and calibrated, motion anomalies are identified and located through micro-motion slice analysis to obtain motion anomaly discrimination results. S5. Based on the body size parameters, body condition scores and motion abnormality discrimination results, continuous physiological state is deduced, long-term physiological trends are extracted through spectral domain filtering, interactive behaviors are determined based on three-dimensional spatial relationships, and multiple body analysis results are aggregated to assess and reveal population health risks.
[0006] According to a second aspect of this application, an electronic device is provided, including a processor and a memory, the memory storing a computer program, wherein the processor executes the program to implement any of the above-described method steps.
[0007] According to a third aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements any of the above-described method steps.
[0008] The beneficial effects of this application are as follows: Through the method provided in this application, the three-dimensional analysis of fragmented field videos is achieved by dynamic primitive decomposition and anti-occlusion reconstruction. By using feature learning and continuous temporal flow modeling, the precise quantitative tracking of individual wildlife identity, three-dimensional body shape, and subtle behavior is completed, and finally, a systematic non-contact assessment from individual physiological parameters to population health risks is achieved. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0010] Figure 1 A flowchart illustrating a method for wildlife ecological insight and behavior analysis provided in an embodiment of this application; Figure 2 A schematic diagram of the dynamic primitive decomposition and anti-occlusion processing provided in an embodiment of this application; Figure 3 A schematic diagram illustrating the construction of an individual identity feature space according to an embodiment of this application; Figure 4 A schematic diagram of a three-dimensional body reconstruction process provided in an embodiment of this application; Figure 5 This is a schematic flowchart illustrating the detection of subtle motion anomalies provided in an embodiment of this application. Figure 6 A flowchart illustrating a population health risk assessment is provided for one embodiment of this application; Figure 7 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0012] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0013] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0014] In biodiversity conservation and wildlife management, field monitoring networks based on monocular cameras face fundamental challenges posed by unstructured environments: animal activity is highly uncertain and concealed, resulting in highly fragmented, discontinuous video data with severe occlusion. Traditional methods treat these discrete "snapshots" as independent events, failing to correlate observations of the same animal at different times, and making it even more difficult to recover complete spatiotemporal motion information and precise morphological structure from localized, blurry images.
[0015] To overcome this bottleneck, this invention proposes a method for wildlife ecological insight and behavioral analysis.
[0016] See Figure 1 The method includes the following steps: S1. Process the acquired wildlife video stream, and obtain a denoised dynamic 3D primitive sequence and target contour image through dynamic primitive decomposition, occlusion extrapolation and repair.
[0017] This step aims to address the issues of image information loss and fragmentation caused by target occlusion and motion blur in outdoor video streams. Its core lies in converting the continuous video stream into two key outputs: a denoised dynamic 3D primitive sequence and a target contour image. (See [link to relevant documentation]). Figure 2 The specific implementation consists of three closely connected sub-processes: S11. Perform inter-frame dense correspondence calculation on the input video stream, segment out rigid geometric primitives and solve their six-degree-of-freedom motion parameters; S12. When a primitive is occluded, inertial extrapolation and virtual geometric completion are performed based on its state before occlusion, and the motion trajectory is smoothed and filtered to obtain the denoised dynamic three-dimensional primitive sequence. S13. Based on the primitive sequence, determine the estimated pose of the target in the video frame, extract the image region accordingly, and combine it with historical features for alignment and interpolation repair to generate the target contour image.
[0018] This embodiment first performs dense inter-frame correspondence calculation on the input video stream. This calculation can be implemented using conventional computer vision algorithms such as optical flow or structure from motion (SfM) matching. For example, the Farneback dense optical flow algorithm can be used to establish a dense inter-frame motion field by calculating the motion vector of each pixel between consecutive video frames; or feature extraction and matching algorithms such as ORB and SIFT can be used, combined with epipolar geometric constraints and fundamental matrix estimation, to recover the motion parameters of rigid primitives in the scene. Based on this, the video scene is segmented into multiple rigid geometric primitives based on pixel motion consistency. This process not only distinguishes between the foreground (animal) and background but also further refines the animal target into multiple component units (such as the head and limb segments) that can be considered as rigid body motion within a short time. Subsequently, the six-degree-of-freedom motion parameters of each rigid geometric primitive are solved, forming a preliminary motion data stream describing the translation and rotation of each component in three-dimensional space.
[0019] Real-time detection is performed to address occlusion of primitives. When occlusion is detected, inertial extrapolation is performed based on the primitive's motion state before occlusion. Specifically, based on its velocity, acceleration, and other motion states before occlusion, a kinematic model is used to predict the trajectory during the occlusion period. Simultaneously, virtual geometric completion is performed to generate a virtual model of the occluded component in 3D space, ensuring the continuity of the animal's overall geometric structure. Furthermore, the primitive motion trajectory is smoothed and filtered to eliminate abnormal data points caused by image noise, matching errors, or momentary jitter, thereby obtaining a denoised dynamic 3D primitive sequence.
[0020] To obtain a high-quality target contour, the spatial information obtained in the preceding steps is utilized. The estimated spatial location and pose of the target within the video frame are determined based on the primitive sequence. Based on this estimate, the target image region to be repaired is extracted from the corresponding frame of the input video stream. Next, combining high-confidence historical features and occlusion buffer features, texture and edge information lost due to occlusion or blurring is recovered through feature alignment and matching. Finally, the region is aligned and interpolated for repair, generating a complete target contour image—a binarized mask with clear edges and removed environmental interference.
[0021] As can be seen, this step successfully transforms fragmented, low-quality video streams acquired in the field into a continuous and denoised dynamic 3D primitive sequence and high-quality target contour images that can be directly used in subsequent steps. This is achieved by decoupling rigid motion units from the video stream, performing physically driven extrapolation and completion during occlusion, and restoring the visual appearance using historical information. This processing lays a crucial data foundation for achieving stable individual tracking and accurate 3D analysis in unstructured environments.
[0022] S2. Based on the target contour image, extract and encode biological features, and construct a feature space through triple constraints to achieve individual identity re-identification and cross-temporal association.
[0023] This step aims to address the challenges of "high similarity in appearance among individuals of the same species" and "significant differences in appearance among the same individual due to environmental and seasonal changes" in wildlife monitoring. Its core lies in transforming the visual information contained in the target's outline image into a unique and spatiotemporally stable digital identity, thereby enabling precise tracking of individuals. (See also...) Figure 3 The specific implementation includes the following four sub-processes: S21. Based on the target contour image, divide the region of interest into anatomical functional regions, extract multi-scale biological features and encode them into feature vectors; S22. Optimize the feature space using triplet data containing anchor samples, positive samples, and negative samples to cluster features of the same individual and separate features of different individuals; S23. For a specific target individual, calculate and generate a feature centroid vector representing the identity of the individual based on its multiple feature vectors; S24. Generate feature centroid vectors based on individual multi-frame feature vectors, and perform identity re-identification or new individual registration through similarity retrieval.
[0024] First, based on the target contour image, regions of interest (ROIs) are divided according to anatomical functional areas. This division is automatic based on the biological characteristics of the target species; for example, for felines, the focus is on facial markings and lateral stripes, and for ungulates, the focus is on angular features and ear notches. Within each region, multi-scale biological features are extracted: using a deep convolutional neural network with residual connections, low-level texture details (such as hair color distribution, scar shape, and speckle geometry) and high-level topological structures (such as limb proportions, spinal curvature, and body contour) are extracted in parallel. Subsequently, these multi-scale biological features are compressed and encoded into a fixed-length one-dimensional feature vector, i.e., a "transient feature descriptor" of the observed target at that moment is generated through a global average pooling layer.
[0025] Then, to construct a discriminative space that enables feature aggregation of all variants of the same target (different angles, lighting, partial occlusion) and feature separation of different targets, this step optimizes the feature space using triplet data containing anchor point samples, positive samples, and negative samples. Specifically: a target image is selected as the anchor point, another image confirmed as the same target is selected as the positive sample, and an image of the same species known as a different target is selected as the negative sample. By defining a boundary constraint loss function, the feature extraction network is iteratively optimized during training, forcibly reducing the Euclidean or cosine distance between the anchor point and the positive sample feature vector in the feature space, while increasing the distance between the anchor point and the negative sample feature vector. Through this process, the system learns and constructs a high-dimensional manifold feature space, in which all observation data of the same individual are mapped into a compact cluster, and the center of the cluster constitutes the individual's "identity anchor point".
[0026] Subsequently, addressing the scarcity of field data, a few-shot learning strategy was employed to generate individual identifiers. For a specific target individual, a feature centroid vector representing that individual's identity was calculated based on its multiple feature vectors. In practice, only 3-5 frames of high-quality images of the individual need to be collected, and their centroids in the feature space are calculated through weighted averaging. This feature centroid vector is defined as the unique "digital biometric fingerprint" of the biological individual and stored in a dynamic identity database. For newly emerging targets that cannot match any fingerprint in the database, the system automatically registers a new identity ID and initializes its profile.
[0027] Finally, during continuous monitoring, online retrieval is continuously performed to achieve tracking. Similarity retrieval is used for identity re-identification or new individual registration: whenever a new target image is captured, the system calculates in real time the distance between its instantaneous feature vector and all existing feature centroid vectors in the database. If the minimum distance is below a preset threshold, it is determined to be identity re-identification, and the system automatically appends the observation data (such as time, location, and physiological parameters) to the corresponding individual's historical file, achieving trajectory connection across several days or even months. Simultaneously, the system employs a sliding window mechanism, using the latest high-confidence observation features to fine-tune the individual's feature centroid vector, enabling it to adaptively follow morphological changes caused by growth, molting, etc., preventing identification failure due to feature drift.
[0028] As can be seen, this step, by transforming unstructured image data into a structured identity ID stream, successfully solves the challenge of continuously and accurately tracking individual wild animals amidst complex appearance changes. It provides crucial object consistency guarantees for all subsequent individual-based longitudinal analyses (such as growth curve fitting and health trend assessment), enabling the integration of discrete observational data into a biologically meaningful individual lifecycle profile.
[0029] S3. Based on the dynamic three-dimensional recognition primitive sequence and identity marker feature decoupling, fusion and three-dimensional reconstruction, drive the parametric mesh model to fit the body shape, and perform virtual measurement to quantify body size parameters and body condition score.
[0030] This step aims to overcome the limitations of lacking depth sensors in field monitoring and solve the problem of how to recover the true three-dimensional spatial structure of wild animals from monocular two-dimensional images and perform non-contact precision body size measurements. Its core is to employ a structural decoupling and interactive correction method to reconstruct accurate, measurable three-dimensional body proportions from blurred monocular images. (See also...) Figure 4 The specific implementation includes the following four sub-processes: S31. Extract the two-dimensional pixel coordinates of the key anatomical points of the target, and deduce the relative depth distribution characteristics of each part of the target from the image texture gradient and perspective relationship; S32. The depth distribution features are calibrated using the two-dimensional pixel coordinate constraints, and a three-dimensional skeleton key point sequence is synthesized by back projection; S33. Use the skeleton sequence to drive the parameterized biological mesh template to optimize posture and body shape deformation; S34. Perform virtual measurements on the optimized 3D model, calculate body size parameters, and quantify body condition scores based on surface geometry information.
[0031] In this embodiment, a parallel processing architecture is designed to address the uncertainty of depth information in monocular vision. It extracts the two-dimensional pixel coordinates of key anatomical points of the target and infers the relative depth distribution features of different parts of the target from image texture gradients and perspective relationships. Specifically, this includes: Two-dimensional planar structural feature extraction: Through a first processing branch focused on high-confidence structural information, the preprocessed target image is input and the pixel positions (X, Y coordinates) of anatomical key points such as the scapula vertex, hip joint center, and nose tip in the two-dimensional image coordinate system of the camera imaging are directly extracted using depth convolution operation. This process does not involve depth estimation and aims to ensure the geometric accuracy of the planar projection structure.
[0032] Spatial Depth Feature Probabilistic Inference: Through a second processing branch focused on inferring ambiguous information, the potential depth distribution is inferred from the texture gradient, shadow occlusion relationship, and perspective distortion of the same image. Due to the inherent uncertainty of monocular depth, this branch introduces a set of pre-defined statistical distribution vectors as priors to assist the model in searching for the most reasonable depth hypothesis in the feature space, and outputs the probability distribution features of the relative distance (Z coordinate) of each part of the target relative to the camera optical center.
[0033] Then, the determinism of the two-dimensional structure is used to constrain the uncertainty of the depth information. The depth distribution features are calibrated using the two-dimensional pixel coordinate constraints: a bidirectional correlation matrix is established between the planar structural features and the depth features, and the correlation weights between keypoint positions and the depth distribution are calculated. Using a high-confidence planar skeleton structure as a rigid constraint, and based on prior knowledge such as the constant proportion of limb lengths in biology, outliers caused by perspective in depth estimation are geometrically calibrated. Subsequently, a three-dimensional skeleton keypoint sequence is synthesized through backprojection: the calibrated depth information (Z-axis) and the precise planar pixel coordinates (X, Y-axis) are backprojected to synthesize a sparse three-dimensional skeleton keypoint sequence of the target in the camera coordinate system.
[0034] Based on this, in order to reconstruct a continuous body surface model reflecting muscle fullness and body contour from the sparse 3D skeleton key sequence, this embodiment performs parametric model fitting. The sparse 3D skeleton key sequence drives a parametric biological mesh template for posture and body shape deformation optimization: First, a standard parametric statistical model of the target species is retrieved as a general mesh template. Then, the 3D skeleton key point cloud generated in the previous step is used as control anchor points, and the standard template is aligned in posture through rigid rotation and translation. Virtual measurement is performed on the optimized 3D model: While maintaining the skeleton posture, the shape parameters of the mesh model are iteratively optimized based on the target silhouette edge contour in the original image. By minimizing the projection error of the 3D mesh on the 2D plane, its contour perfectly matches the original image, thereby accurately reconstructing the 3D body contour details of the target individual, such as abdominal sagging, back width, and muscle lines.
[0035] Finally, this embodiment performs automated and standardized virtual measurements on the reconstructed, precise 3D mesh model. Body size parameters are calculated: standard anatomical measurement points are automatically identified and selected, and the Euclidean distance between these point pairs is calculated in 3D space, directly converting them into actual physical dimensions (such as shoulder height, body length, and chest circumference), completely eliminating perspective shortening errors caused by tilted shooting angles. Body condition scoring is quantified based on surface geometry information: by calculating the curvature to volume ratio of the mesh at specific locations, a surface smoothness index and volume / skeleton ratio reflecting fat reserve levels are generated, serving as a digital and objective basis for assessing an individual's nutritional status (lean, normal, obese).
[0036] It can be seen that this step, through an innovative decoupling-fusion reconstruction strategy, successfully achieved millimeter-level precision reconstruction and automated measurement of the three-dimensional anatomy of individual wild animals using only a monocular camera. It elevates traditional two-dimensional image analysis to a quantifiable three-dimensional spatial analysis level, providing a direct and accurate foundation of three-dimensional physiological parameter data for subsequent growth trend fitting and health assessment.
[0037] S4. Based on the three-dimensional skeleton sequence, a kinematic model is constructed. After the sparse motion data is completed and calibrated, motion anomalies are identified and located through micro-motion slice analysis to obtain motion anomaly discrimination results.
[0038] This step aims to address the technical challenges of low frame rate shooting in field monitoring equipment, which is limited by storage and bandwidth. This results in discontinuous capture of rapid animal movements and missed detection of early pathological micro-movements that are difficult to discern with the naked eye. Through computational reconstruction and feature amplification techniques, it achieves motion signal recovery from sparse to dense at the data level and abnormal behavior screening from macro to micro at the semantic level. (See also...) Figure 5 Specifically, it includes the following four sub-processes: S41. Construct a species-specific skeletal topology model, perform bidirectional temporal interpolation on low frame rate skeleton data, and generate continuous motion trajectories. S42. Apply angular velocity constraints in the rotating manifold space to perform biomechanical smoothing on the motion trajectory; S43. Divide the motion trajectory into micro-time slices, extract multi-scale motion features, and identify abnormal patterns by analyzing the feature distances between symmetrical slices; S44. Based on the anomaly analysis results of the feature distance, locate the specific time of the anomaly and the associated skeletal location, and generate the motion anomaly discrimination result.
[0039] First, a biokinematic topological model of the target species is constructed as the basis for physical constraints. This model digitally defines rigid constraints such as standard skeletal connections, joint degrees of freedom restrictions, and limb length ratios for the species. For example, for deer, it is explicitly defined that the knee joint can only flex and extend forward and backward, but not laterally twist. Next, the three-dimensional skeleton sequence generated in step S3 (a set of three-dimensional coordinates of key points sampled at discrete time points) is used as observations and input into a linear recursive state estimator. This estimator does not treat each moment as an isolated event, but assumes that the animal's motion state (position, velocity, acceleration) evolves continuously over time. A bidirectional temporal scanning strategy is used to simultaneously perform forward prediction and backward smoothing: by performing kinematically guided interpolation calculations between two sparse measured skeleton poses, not only is the motion state of the current frame used to predict the next frame, but information from subsequent frames is also used to correct the estimation of the previous frame. By aggregating this bidirectional information flow, a series of high-temporal-resolution virtual skeleton poses are automatically generated between two actual observation points based on biological motion inertia. This numerically completes the low-frame-rate video into a high-frame-rate continuous motion trajectory, effectively filling the observation blind spot caused by insufficient sampling rate.
[0040] Subsequently, to prevent the motion trajectories generated by the interpolation from violating biomechanical principles, this embodiment introduces a geometric angular velocity consistency constraint for physical calibration. Specifically, the joint motion data is transformed from the traditional Euclidean coordinate space to a Lie algebraic manifold space, which is more suitable for describing rotation, for calculation. In this space, the relative rotational angular velocities of joints between adjacent interpolated postures are monitored and calculated in real time. A geometric angular velocity loss function is defined. If the calculated rate of change of angular velocity exceeds the biological limits that the muscles and ligaments of the species can drive (e.g., unnatural instantaneous joint retraction or ultra-high-speed rotation), a smoothing penalty is automatically applied to the trajectory segment, forcing it to revert to a smooth curve that conforms to the characteristics of gravity and muscle drive. This step eliminates false motion anomalies caused by data noise or interpolation algorithm defects, ensuring that subsequently detected abnormal signals truly reflect the physiological state of the organism.
[0041] After obtaining a physically plausible continuous motion flow, fine motion feature analysis and anomaly detection are performed. First, the calibrated motion trajectory is divided into micro-time slices according to extremely short time windows, such as a non-overlapping segment every 0.5 seconds. Each segment represents an instantaneous, basic "micro-motion unit." For each micro-motion slice, a U-shaped multi-scale feature fusion network is used for encoding: the bottom layer of the network extracts micro-displacement details within the slice (such as "the subtle tremor amplitude when the foot lands"), and the top layer extracts macro-motion semantics (such as "stepping with the left foreleg"). The high and low layer features are fused through jump connections to generate a high-dimensional feature description vector for the slice. Subsequently, in the feature space, the feature distance between periodic or symmetrical slices is calculated for analysis: for walking gait, the feature distance between the left forelimb stepping slice and the right forelimb stepping slice is calculated. If the distance increases significantly, it indicates bilateral movement asymmetry, possibly indicating unilateral limping; for stationary or slow-moving states, the feature distance between consecutive adjacent slices is calculated. If the distance shows irregular high-frequency fluctuations, it indicates involuntary muscle tremors or control impairment.
[0042] Finally, anomaly localization and discrimination results are generated. Based on the aforementioned feature distance analysis results, an "anomaly confidence curve" is generated that changes over time. Peak periods exceeding a preset threshold in the curve are automatically marked, precisely locating the specific video frame number and duration of the anomaly. Simultaneously, the attention weight map in the feature fusion network is used for reverse attribution to locate key skeletal nodes causing abrupt changes in feature distance. For example, if attention is highly focused on the "left posterior ankle joint," this area is identified as the source of the anomaly. Combining environmental data (e.g., excluding shivering caused by hypothermia) with individual health records, structured motion anomaly discrimination results are generated, such as: "Target ID-X exhibits 'shortened left hind limb support phase' at [timestamp], accompanied by grade II limping characteristics, suspected to be a recurrence of a previous joint injury." This achieves earlier and more objective early warning of pathological behavior than manual observation.
[0043] S5. Based on the body size parameters, body condition scores and motion abnormality discrimination results, continuous physiological state is deduced, long-term physiological trends are extracted through spectral domain filtering, interactive behaviors are determined based on three-dimensional spatial relationships, and multiple body analysis results are aggregated to assess and reveal population health risks.
[0044] This step aims to address the core challenges of uneven observation intervals and significant environmental noise interference in field monitoring. By establishing a continuous dynamic model that supports arbitrary timestamp input, it enables the derivation of continuous and pure physiological state evolution curves from discrete, noisy individual observation data, ultimately achieving a quantitative assessment of risks from microscopic individual behavior to macroscopic population ecosystem risks. (See also...) Figure 6 Specifically, it includes the following five sub-processes: S51. Bind body size parameters, body condition scores, and motion abnormality discrimination results to timestamps to construct a multidimensional physiological time series; S52. Fit the spatiotemporal velocity field function of the sequence, and deduce the physiological state and three-dimensional position at any consecutive moment through integration; S53. Perform frequency domain transformation on physiological time series data, and use spectral differentiation to extract low-frequency trends and suppress high-frequency noise; S54. Based on individual identity and deduced 3D position and motion state, calculate the relative relationship between individuals and match interaction semantic rules to determine interaction behavior events; S55. Divide the spatial distribution range of multiple observed individuals into analysis units, aggregate the density, health index and interaction frequency of individuals in each unit to construct an ecological feature vector, and identify population health risks through system state evolution trajectory after dimensionality reduction analysis.
[0045] First, data integration and structured encapsulation are performed. The time-series body size parameters (such as shoulder height and body length) and body condition scores (BCS) quantified in step S3 are correlated and integrated with the motion anomaly discrimination results (such as lameness level and tremor index) obtained in step S4. This integration is not a simple stacking, but rather aligns data from different sources and dimensions onto a unified timeline based on individual identification. Subsequently, each integrated data unit (state tensor) is precisely bound to its corresponding absolute physical timestamp, forming a set of "timestamp-state tensor" pairs. This process solves the inherent problem of highly irregular time points in field observations, constructing a unified and time-aligned multidimensional physiological time series data, providing standardized input for continuous extrapolation models that process non-uniform time series.
[0046] Then, continuous physiological state inference based on spatiotemporal velocity field is performed. This application abandons the traditional discrete recursive prediction architecture based on fixed time steps (such as RNN, LSTM), and instead adopts a more advanced continuous flow matching or neural ordinary differential equation solver as the core prediction engine. The core idea of this model is to learn a spatiotemporal velocity field function. This function does not directly predict the state value at the next moment, but learns the mapping from the current state to the rate of change of state (i.e., "velocity"). Fitting the spatiotemporal velocity field function of this sequence: Through training, the model learns what the derivative (trend of change) of any given multidimensional physiological state is over time. In the inference stage, by numerically integrating this velocity field function, the system can calculate the theoretical physiological state and three-dimensional spatial position of an animal at any continuous future or past time point, starting from a known observed state. For example, between two actual observations 10 days apart, the model can generate intermediate states at a virtual resolution once per hour, thereby depicting a smooth weight change curve, clearly showing whether the animal is steadily gaining weight or suddenly losing weight.
[0047] Next, the system extracts long-term trends reflecting biological essence from noisy time-series data. First, it performs a frequency domain transformation on the physiological time-series data, mapping it from the time domain to the frequency domain for analysis using a discrete Fourier transform. In the frequency domain, short-term random disturbances (such as single measurement errors or temporary soil adhesion to the target) manifest as high-frequency components, while actual individual growth, emaciation, or chronic health deterioration manifest as low-frequency components. The system uses spectral differential operators to extract low-frequency trends. Spectral differential can directly calculate the global derivative in the frequency domain, making it more robust to local noise spikes than the time-domain difference method, and obtaining a more stable long-term trend signal. Simultaneously, the system implements a dynamic weighting mechanism to suppress high-frequency noise. This mechanism calculates the "innovation" value between the model's predicted value and the new observation value and analyzes its spectral characteristics. If the "innovation" value is small and its energy is concentrated in the low frequency, it is considered valid information, and the observation's weight is increased to update the state estimate; if the "innovation" value is large and exhibits high-frequency characteristics, it is determined to be sudden noise, and its weight is automatically reduced to prevent it from contaminating the long-term health trend line.
[0048] Next, automated semantic determination of ecological interaction events based on three-dimensional spatial logic is performed, transforming quantified spatial information into ecological behavioral semantics. Its input depends on the three-dimensional spatial positions and orientations of different individuals at the same moment, derived from S5.2. The system calculates the Euclidean distance vector between individuals, the angle between head / torso orientation vectors, and their respective velocity vectors. These calculated spatial relationship parameters are then fed into a pre-defined interaction semantic rule base for matching. This rule base is predefined based on animal behavioral ecology knowledge, for example: Standoff determination rule: If the distance between two different species is less than D1 meters, and the angle between their heads is close to 180 degrees (face to face), and this state lasts for more than T1 seconds.
[0049] Predation / chasing determination rule: If the velocity vector of predator individual B continuously points towards prey individual A, and B's velocity is greater than A's, and the distance between the two continuously shortens to near zero within a time interval Δt.
[0050] Rules for determining social behavior: If the centroids of multiple individuals of the same species are distributed in three-dimensional space within a sphere of radius R, and their trajectories show a high correlation through covariance analysis.
[0051] Through rule matching, the original 3D coordinate data stream is automatically transformed into a structured interactive behavior event log, such as "Time T, Individual ID-1 and Individual ID-2 are in a standoff, lasting X seconds".
[0052] Finally, macroscopic ecological dynamics maps and population health risk identification are performed. First, the monitoring area is divided into spatial grids or topological nodes as the basic spatial units for analysis. For each unit, the system aggregates three key types of information from all observed individuals: population density (number of individuals / km²), average health index (a combined value based on body condition score and anomaly index), and frequency of interaction events (such as the number of confrontations or chases per unit time), thus constructing a multidimensional vector representing the ecological state of that unit, i.e., the ecological feature vector. The vectors from all units together constitute the ecological vector field at that moment. Subsequently, manifold learning techniques such as multidimensional scaling analysis (MDS) are used to perform dimensionality reduction analysis on the ecological feature vector, mapping the high-dimensional ecological vector field to a two-dimensional or three-dimensional visualization space. In this low-dimensional space, each point represents the state of the entire ecosystem at a certain moment. By observing the evolutionary trajectories of these state points over time, the system can perform advanced pattern recognition: if a trajectory point undergoes periodic or quasi-static movement within a small area for a long period, forming a stable attractor, it indicates that the ecosystem is in a healthy dynamic equilibrium; if a trajectory point suddenly deviates from its original trajectory and rapidly disperses in an unknown direction, it strongly suggests that a structural mutation has occurred in the ecosystem, posing a risk to population health (such as disease outbreaks or a decline in overall physical condition due to food shortages). This data-driven macroscopic map provides managers with a panoramic ecological health situation perception and early warning capability that goes beyond individual case reports.
[0053] This step overcomes the challenge of uneven observation by constructing a continuous-time dynamic model, enabling "blind-spot-free" extrapolation of individual physiological states. Through spectral domain analysis and three-dimensional spatial logic, it achieves robust extraction of long-term health trends and automated understanding of complex ecological interactions, respectively. Finally, through the construction and dimensionality reduction analysis of the ecological feature field, it aggregates massive amounts of individual data into an intuitive macro-risk map, completing the entire closed-loop chain from precise micro-level monitoring to macro-level scientific decision-making.
[0054] In addition, this application also provides an electronic device 100, including a processor 101 and a memory 102, wherein the memory stores a computer program, and the processor executes the program to implement the above-mentioned wildlife ecological insight and behavior analysis method.
[0055] Finally, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for wildlife ecological insight and behavior analysis.
[0056] Accordingly, the electronic device and computer-readable storage medium have the same technical effects as the above-described method, which will not be elaborated further here.
[0057] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0058] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0059] The above are merely optional embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made based on the inventive concept of this application and the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included within the patent protection scope of this application.
Claims
1. A method for wildlife ecological insight and behavioral analysis, characterized in that, The method includes the following steps: S1. Process the acquired wildlife video stream, and obtain a denoised dynamic 3D primitive sequence and target contour image through dynamic primitive decomposition, occlusion extrapolation and repair. S2. Based on the target contour image, extract and encode biological features, and construct a feature space through triplet constraints to achieve individual identity re-identification and cross-temporal association; S3. Based on the dynamic three-dimensional primitive sequence and identity identifier, through feature decoupling, fusion and three-dimensional reconstruction, drive the parametric mesh model to fit the body shape, and perform virtual measurement to quantify body size parameters and body condition score; S4. Based on the three-dimensional skeleton sequence, a kinematic model is constructed. After the sparse motion data is completed and calibrated, motion anomalies are identified and located through micro-motion slice analysis to obtain motion anomaly discrimination results. S5. Based on the body size parameters, body condition scores and motion abnormality discrimination results, continuous physiological state is deduced, long-term physiological trends are extracted through spectral domain filtering, interactive behaviors are determined based on three-dimensional spatial relationships, and multiple body analysis results are aggregated to assess and reveal population health risks.
2. The method according to claim 1, characterized in that, Step S1 specifically includes: S11. Perform inter-frame dense correspondence calculation on the input video stream, segment out rigid geometric primitives and solve their six-degree-of-freedom motion parameters; S12. When a primitive is occluded, inertial extrapolation and virtual geometric completion are performed based on its state before occlusion, and the motion trajectory is smoothed and filtered to obtain the denoised dynamic three-dimensional primitive sequence. S13. Based on the primitive sequence, determine the estimated pose of the target in the video frame, extract the image region accordingly, and combine it with historical features for alignment and interpolation repair to generate the target contour image.
3. The method according to claim 1, characterized in that, Step S2 specifically includes: S21. Based on the target contour image, divide the region of interest into anatomical functional regions, extract multi-scale biological features and encode them into feature vectors; S22. Optimize the feature space using triplet data containing anchor samples, positive samples, and negative samples to cluster features of the same individual and separate features of different individuals; S23. For a specific target individual, calculate and generate a feature centroid vector representing the identity of the individual based on its multiple feature vectors; S24. Generate feature centroid vectors based on individual multi-frame feature vectors, and perform identity re-identification or new individual registration through similarity retrieval.
4. The method according to claim 1, characterized in that, Step S3 specifically includes: S31. Extract the two-dimensional pixel coordinates of the key anatomical points of the target, and deduce the relative depth distribution characteristics of each part of the target from the image texture gradient and perspective relationship; S32. The depth distribution features are calibrated using the two-dimensional pixel coordinate constraints, and a three-dimensional skeleton key point sequence is synthesized by back projection; S33. Use the skeleton sequence to drive the parameterized biological mesh template to optimize posture and body shape deformation; S34. Perform virtual measurements on the optimized 3D model, calculate body size parameters, and quantify body condition scores based on surface geometry information.
5. The method according to claim 1, characterized in that, Step S4 specifically includes: S41. Construct a species-specific skeletal topology model, perform bidirectional temporal interpolation on low frame rate skeleton data, and generate continuous motion trajectories. S42. Apply angular velocity constraints in the rotating manifold space to perform biomechanical smoothing on the motion trajectory; S43. Divide the motion trajectory into micro-time slices, extract multi-scale motion features, and identify abnormal patterns by analyzing the feature distances between symmetrical slices; S44. Based on the anomaly analysis results of the feature distance, locate the specific time of the anomaly and the associated skeletal location, and generate the motion anomaly discrimination result.
6. The method according to claim 1, characterized in that, Step S5 specifically includes: S51. Bind body size parameters, body condition scores, and motion abnormality discrimination results to timestamps to construct a multidimensional physiological time series; S52. Fit the spatiotemporal velocity field function of the sequence, and deduce the physiological state and three-dimensional position at any consecutive moment through integration; S53. Perform frequency domain transformation on physiological time series data, and use spectral differentiation to extract low-frequency trends and suppress high-frequency noise; S54. Based on individual identity and deduced 3D position and motion state, calculate the relative relationship between individuals and match interaction semantic rules to determine interaction behavior events; S55. Divide the spatial distribution range of multiple observed individuals into analysis units, aggregate the density, health index and interaction frequency of individuals in each unit to construct an ecological feature vector, and identify population health risks through system state evolution trajectory after dimensionality reduction analysis.
7. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program, and the processor executing the program to implement the method of any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method of any one of claims 1 to 6.