Wild animal and plant species identification method

By collecting and preprocessing biological and environmental data, multimodal features are extracted and time alignment and dynamic weight adjustment are performed, comprehensive feature vectors are generated, and large models are used for identification, which solves the problem of degradation of recognition accuracy caused by mutations in environmental parameters in the prior art, and achieves more efficient species recognition.

CN120470544AActive Publication Date: 2025-08-12ZHEJIANG NONGCHAOER SMART TECH CO LTD

Patent Information

Application Number
CN202510968474.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-08-12
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

The prior art does not consider the temporal dynamics of environmental data in wildlife species identification, resulting in a decrease in the accuracy of species identification when environmental parameters are mutations and fails to effectively fuse multimodal data.

Method used

By collecting biological data and environmental data, multimodal features are extracted after preprocessing, and feature fusion is performed based on time alignment and dynamic adjustment of weights, comprehensive feature vectors are generated, and large models are used for identification.

Benefits of technology

It improves the accuracy and robustness of wildlife species identification, and can maintain efficient identification in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470544A_ABST
    Figure CN120470544A_ABST
Patent Text Reader

Abstract

The invention provides a wild animal and plant species identification method, which comprises the following steps: preprocessing collected biological data and environmental data associated with a target species to obtain to-be-analyzed data, the biological data comprising an image, an audio and positioning data associated with the target species, the environment data comprises meteorological and physical environment data associated with the living environment of the target species, and the target species are wild animals or wild plants; performing feature extraction on the to-be-analyzed data, and aligning the extracted multi-modal features based on time; based on the dynamically adjusted weight of each modal feature, performing feature fusion on the aligned multi-modal features to generate a comprehensive feature vector of the target species; and processing the comprehensive feature vector based on the target large model, obtaining at least one species name of the target species and the confidence of each species name, and determining a final name based on the confidence. According to the method, environment feature processing and dynamic weight adjustment mechanisms are introduced, so that species identification can be intelligently and accurately carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of ecological technology, and in particular to a method for identifying wild animal and plant species. Background Art

[0002] With the improvement of global awareness of ecological protection and biodiversity conservation, the demand for monitoring and research on wild animals and plants is growing.

[0003] Currently, when monitoring and researching wildlife, researchers rely on artificial intelligence (AI) to conduct intelligent analysis of wildlife-related images, videos, audio, and other data for species research. Processing environmental data related to species survival is relatively straightforward. Existing technologies (such as purely visual classification models) typically fail to consider the temporal dynamics of environmental data. This inability to identify species in specific environments is due to a lack of consideration for the impact of environmental factors on species identification. Furthermore, existing technologies typically perform static weight fusion on collected species images and audio, which can affect species identification accuracy when sudden changes in environmental parameters lead to a step-down decrease in the quality of a specific modality (e.g., a sharp drop in audio signal-to-noise ratio during heavy rain). Summary of the Invention

[0004] In view of the above problems, embodiments of the present application provide a method for identifying wild animal and plant species that overcomes the above problems or at least partially solves the above problems.

[0005] In a first aspect, embodiments of the present application provide a method for identifying wild animal and plant species, comprising: Preprocessing the collected biological data and environmental data associated with the target species to obtain data to be analyzed corresponding to the target species, the biological data including image data, audio data, and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environment data associated with the living environment of the target species, where the target species is a wild animal or a wild plant; Extracting features from the data to be analyzed and aligning the extracted multimodal features based on time, the multimodal features including biometric features and environmental features, the biometric features including image features extracted from the image data, audio features extracted from the audio data, and positioning features extracted from the positioning data; Based on the dynamically adjusted weights of each modal feature, the aligned multimodal features are subjected to feature fusion to generate a comprehensive feature vector indicating comprehensive feature information of the target species, wherein the weights of each modal feature are dynamically adjusted based on the real-time contribution of the multimodal feature to species identification; Based on the target large model, the comprehensive feature vector is subjected to inference analysis to obtain at least one species name corresponding to the target species and the confidence of each species name, and the species name of the target species is determined based on the confidence of each species name.

[0006] In a second aspect, an embodiment of the present application provides a system for identifying wild animal and plant species, including: a collection device and a cloud platform; The collection device collects biological data and environmental data associated with the target species, the biological data including image data, audio data, and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environment data associated with the living environment of the target species, wherein the target species is a wild animal or a wild plant; The cloud platform preprocesses the data provided by the acquisition device, obtains the data to be analyzed corresponding to the target species, extracts features from the data to be analyzed, and aligns the extracted multimodal features based on time, wherein the multimodal features include biological features and environmental features, and the biological features include image features extracted from the image data, audio features extracted from the audio data, and positioning features extracted from the positioning data; The cloud platform performs feature fusion on the aligned multimodal features based on the dynamically adjusted weights of each modal feature to generate a comprehensive feature vector indicating comprehensive feature information of the target species. The weights of each modal feature are dynamically adjusted based on the real-time contribution of the multimodal feature to species identification. The cloud platform performs reasoning analysis on the comprehensive feature vector based on the deployed target large model, obtains at least one species name corresponding to the target species and the confidence of each species name, and determines the species name of the target species based on the confidence of each species name.

[0007] The technical solution of the embodiment of the present application collects biological data and environmental data associated with the target species. After obtaining the data to be analyzed corresponding to the target species based on data preprocessing, feature extraction is performed on the data to be analyzed, the extracted multimodal features are aligned based on time, and based on the weights of each modal feature after dynamic adjustment, the aligned multimodal features are feature fused to generate a comprehensive feature vector. The comprehensive feature vector is inferred and analyzed based on the target large model, and the species name of the target species and the confidence level of each species name are predicted. The large model technology can be used to realize the intelligent identification of wild animals and plants, and by introducing environmental feature processing and dynamic weight adjustment mechanism, a more comprehensive multimodal data fusion solution is provided to optimize the generation of comprehensive feature vectors, and then the accuracy and robustness of wild animal and plant identification can be significantly improved by analyzing the optimized comprehensive feature vectors. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A schematic diagram showing a method for identifying wild animal and plant species provided in an embodiment of the present application; Figure 2 An example diagram showing the method of dynamically adjusting the weights of each modal feature in a baseline feature vector to generate a comprehensive feature vector according to an embodiment of the present application; Figure 3 A schematic diagram of a wild animal and plant species identification system provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0009] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0010] It should be understood that references throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, the appearance of "in one embodiment" or "in an embodiment" throughout this specification does not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. The term "a plurality" in the embodiments of the present application may include two or more.

[0011] In the various embodiments of the present application, it should be understood that the size of the serial numbers of the following processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0012] The present invention provides a method for identifying wild animal and plant species. Figure 1 As shown, the method is applied to the cloud platform and includes the following steps: Step 101: Preprocess the collected biological data and environmental data associated with the target species to obtain data to be analyzed corresponding to the target species. The biological data includes image data, audio data, and positioning data associated with the target species. The environmental data includes at least meteorological data and physical environment data associated with the living environment of the target species. The target species is a wild animal or a wild plant.

[0013] After using the acquisition equipment to collect biological and environmental data associated with the target species, the collected data is preprocessed to obtain the data to be analyzed corresponding to the target species. The collected biological data includes at least image data, audio data, and positioning data, and the collected environmental data includes at least meteorological data and physical environment data. Specifically, environmental data includes but is not limited to temperature, humidity, soil conditions, light intensity, water resources, air pressure, wind speed, season, and other data related to the species' living environment. By collecting environmental data associated with the target species, contextual information on the species' survival and activities can be provided, allowing for a more comprehensive description of species characteristics, thereby helping to improve the accuracy of species identification.

[0014] In one specific embodiment, data collection is performed using unmanned devices (such as drones and unmanned boats) equipped with various sensors. The sensors on the unmanned devices are controlled to collect data at a set frequency and add attributes such as timestamps and location information. This enables unified sensor control and synchronized data collection, thereby obtaining multi-dimensional information related to species. The unmanned devices can travel along a preset or autonomously planned path, acquiring image data, audio data, location data, and environmental data during travel. For unmanned devices traveling along a preset path, a GIS (Geographic Information System) tool or path planning software is used to pre-plan the path (e.g., a path represented by a series of longitude and latitude coordinates). The navigation system of the unmanned device automatically adjusts its direction and speed based on these coordinates. For autonomously planned paths, a path planning algorithm is used in conjunction with real-time environmental data (such as obstacle locations and water flow speed) to dynamically generate a path. The navigation system of the unmanned device automatically adjusts its direction and speed based on the real-time location information and path planning results.

[0015] After collecting multi-dimensional data associated with the target species, the collected data is preprocessed, such as denoising, enhancement, format conversion, etc. This can provide relatively rich data while improving data quality and laying the foundation for subsequent data analysis.

[0016] Step 102: extract features from the data to be analyzed and align the extracted multimodal features based on time. The multimodal features include biometric features and environmental features. The biometric features include image features extracted from the image data, audio features extracted from the audio data, and positioning features extracted from the positioning data.

[0017] After obtaining the target species' data for analysis through data preprocessing, multimodal features are extracted from the data. These features include biological and environmental features. Biological features include image features, audio features, and positioning features. For example, convolutional neural networks are used to extract wildlife and plant features from images, such as texture, shape, and color. Audio features such as Mel-spectrum spectra and Mel-frequency cepstral coefficients (MFCCs) are extracted, reflecting the frequency and temporal characteristics of the audio. Two-dimensional coordinates of positioning data are extracted as features. Environmental features include, for example, climate, soil, hydrological (water quality, depth, temperature, flow velocity, etc.), topographical features, and vegetation characteristics.

[0018] After extracting multimodal features, these features are aligned over time to ensure consistency across the temporal dimension of the multimodal data. For example, asynchronous biological data such as images, audio, and positioning data can be aligned with dynamic environmental data such as meteorological and physical environment data over time to address the issue of inconsistent timing across multiple sources in natural scenarios.

[0019] Step 103: Based on the dynamically adjusted weights of each modal feature, the aligned multimodal features are subjected to feature fusion to generate a comprehensive feature vector indicating the comprehensive feature information of the target species. The weights of each modal feature are dynamically adjusted based on the real-time contribution of the multimodal feature to species identification.

[0020] After aligning the multimodal features, the weights of each modal feature are dynamically adjusted, and feature fusion is performed based on the dynamically adjusted weights and the aligned multimodal features to generate a comprehensive feature vector indicating the comprehensive feature information of the target species.

[0021] Dynamic weight adjustment is based on the contribution of biological and environmental features to species identification. Specifically, the weight of each modal feature is dynamically adjusted based on its real-time contribution to species identification. The essence of dynamic weight adjustment is to quantify the reliability of each modal feature in real time and achieve adaptive fusion of features through weight assignment. Dynamic weight adjustment optimizes the generation of comprehensive feature vectors, leveraging them to improve the accuracy and robustness of species identification. Specifically, for example, dynamic weight adjustment based on the real-time contribution of each modal feature to species identification (e.g., higher weight for images on sunny days and higher weight for audio at night) can enhance recognition robustness in complex environments.

[0022] The process of fusing features from different modalities to generate a comprehensive feature vector involves combining image, audio, positioning, and environmental features with corresponding weights to create a high-dimensional feature vector. This process, through the introduction of environmental features and dynamic weight adjustment, provides a more comprehensive multimodal data fusion solution. This multimodal data fusion approach to generating a comprehensive feature vector leverages the strengths of data from different modalities, improving the accuracy and robustness of species identification.

[0023] Existing technologies typically fuse traditional biological image and audio features, but this fusion method cannot identify species in specific environments. This solution models the association of traditional biological image features (such as bird feather texture) and audio features (such as bird song spectra) with environmental characteristics (such as temperature, humidity, and altitude). This approach uses environmental constraints to narrow the species identification range (e.g., only species that are likely to occur at a specific humidity).

[0024] Step 104: perform inference analysis on the comprehensive feature vector based on the target large model to obtain at least one species name corresponding to the target species and the confidence level of each species name, and determine the species name of the target species based on the confidence level of each species name.

[0025] After determining the comprehensive feature vector through feature fusion, the comprehensive feature vector is inferred and analyzed based on the target large model used for species identification to obtain at least one species name corresponding to the target species output by the target large model and the confidence of each species name. Based on the at least one species name and the corresponding confidence, the final species name of the target species is determined to use the large model for wild animal and plant species identification.

[0026] The confidence of the species name indicates the degree of certainty of the model in the prediction result of the species name, which is usually expressed as a probability value. Based on the confidence of each species name output by the target large model for the target species, the target species can be effectively identified.

[0027] The above implementation scheme of the present application collects biological data and environmental data associated with the target species. After obtaining the data to be analyzed corresponding to the target species based on data preprocessing, feature extraction is performed on the data to be analyzed, the extracted multimodal features are aligned based on time, and based on the dynamically adjusted weights of each modal feature, the aligned multimodal features are fused to generate a comprehensive feature vector. The comprehensive feature vector is inferred and analyzed based on the target large model, and the species name of the target species and the confidence level of each species name are predicted. The large model technology can be used to realize the intelligent identification of wild animals and plants, and by introducing environmental feature processing and dynamic weight adjustment mechanism, a more comprehensive multimodal data fusion solution is provided to optimize the generation of comprehensive feature vectors, and then the accuracy and robustness of wild animal and plant species identification can be significantly improved by analyzing the optimized comprehensive feature vectors.

[0028] The following describes the process of obtaining the data to be analyzed based on data preprocessing. When preprocessing the collected biological data and environmental data associated with the target species to obtain the data to be analyzed corresponding to the target species, it includes: Perform image denoising, image enhancement, and image standardization on image data associated with the target species to obtain images to be analyzed; Removing background noise from audio data associated with the target species to obtain audio to be analyzed; Convert the latitude and longitude coordinates of the target species into plane coordinates, and combine them with the data measured by the inertial measurement unit to obtain the positioning data to be analyzed; Standardize the environmental data associated with the target species to a uniform range to obtain the environmental data to be analyzed.

[0029] During image data preprocessing, image denoising algorithms (such as Gaussian filtering and median filtering) are used to remove noise from the image, and image enhancement and standardization are performed. Image enhancement enhances image details through methods such as histogram equalization and contrast adjustment. Image standardization normalizes image data to a uniform size and format. Image preprocessing, including denoising, enhancement, and standardization, can improve image quality.

[0030] When preprocessing audio data, use audio noise reduction algorithms (such as spectral subtraction and wavelet transform) to remove background noise. Normalization and pre-emphasis can also be used to further process the audio. Pre-emphasis enhances high-frequency components in the audio signal and highlights important features. Normalization scales the amplitude of the audio signal to a specific range to improve numerical stability. Through these audio processing methods, the quality of the audio data can be gradually improved.

[0031] When preprocessing positioning data, the longitude and latitude coordinates are converted into two-dimensional plane coordinates. Combined with IMU (Inertial Measurement Unit) data, a weighted fusion algorithm is used to obtain more accurate positioning information. When preprocessing environmental data, the environmental data is normalized to a uniform range.

[0032] Among them, the preprocessed image data can be stored in JPEG or PNG format, with timestamp and positioning information added as file name or metadata; the preprocessed audio data can be stored in WAV or MP3 format, with timestamp and positioning information added as file name or metadata; the preprocessed positioning data and environmental data can be stored in CSV or JSON format, containing timestamp, longitude and latitude, environmental parameters and other information.

[0033] Optionally, after obtaining the data to be analyzed corresponding to the target species based on data preprocessing, feature extraction is performed on the data to be analyzed, and the extracted multimodal features are aligned based on time, including: Image features, audio features, positioning features, and environmental features with timestamps are extracted from the data to be analyzed; the extracted features of different modalities are aligned according to the timestamps to obtain multimodal features that are consistent in the time dimension.

[0034] When extracting features from the data to be analyzed, for example, convolutional neural networks can be used to extract image features, extract features such as the Mel-spectrogram and MFCC of audio, extract positioning features from positioning data, and extract environmental features from environmental data. Environmental features are obtained by converting environmental data into feature vectors that reflect the living environment of the target species. For example, environmental data such as temperature and humidity can be standardized to serve as environmental features.

[0035] After feature extraction, image features, audio features, positioning features, and environmental features with precise timestamps are obtained, and feature alignment is performed based on the timestamps so that data at the same time point can be matched, ensuring data consistency in the time dimension.

[0036] Through the above-mentioned data preprocessing process, the data quality can be improved, providing a better data foundation and more powerful data support for the subsequent identification of wild animal and plant species; by extracting features from the preprocessed data and aligning the extracted features based on time, multimodal features that are consistent in the time dimension can be obtained.

[0037] In an optional embodiment of the present application, when the aligned multimodal features are subjected to feature fusion based on the dynamically adjusted weights of each modal feature to generate a comprehensive feature vector indicating comprehensive feature information of the target species, the method includes: Perform feature splicing on the aligned image features, audio features, positioning features, and environmental features to generate a high-dimensional feature vector; Assign initial weights to each modal feature in the multimodal features, and fuse the assigned initial weights into a high-dimensional feature vector to generate a baseline feature vector. The initial weights of environmental features are determined based on the degree of impact of environmental data on the survival and identification of the target species. The initial weights of each modal feature in the biological features are assigned based on the estimated importance of each modal feature in species identification. Based on the target parameters, the weights of each modal feature in the baseline feature vector are dynamically adjusted, and the weight change amplitude is controlled based on the weight smoothing biological constraint mechanism. A comprehensive feature vector is generated according to the baseline feature vector with dynamically adjusted weights. The target parameters include at least one of the following: the full modal weight corresponding to the system time period, the confidence of the target large model for each modal feature, the user's feedback information on the recognition results output by the target large model in the previous recognition cycle, the newly introduced sample data in the current recognition cycle, and the scenario in which the current recognition cycle is located.

[0038] After the extracted features are aligned based on the timestamps, the aligned image features, audio features, positioning features, and environmental features are concatenated to generate a high-dimensional feature vector. For example, F high-dimensional feature association = [F image, F audio, F positioning, F environment], where F image represents image features, F audio represents audio features, F positioning represents positioning features, and F environment represents environmental features.

[0039] Based on the importance and relevance of the data from different modalities, initial weights are assigned to each modality and fused into a high-dimensional feature vector to generate a baseline feature vector. The initial weights for environmental features are determined based on the impact of environmental data on the survival and identification of the target species. The initial weights for each modality in the biological signature are assigned based on the estimated importance of each modality in species identification. For example, the baseline feature vector F = αF image + βF audio + γF positioning + δF environment, where α, β, γ, and δ are initial weight coefficients, satisfying α + β + γ + δ = 1.

[0040] When generating a baseline feature vector, a comprehensive feature vector can be generated by dynamically adjusting weights. This dynamic weight adjustment process is based on the contribution of biological and environmental features to species identification. When dynamically adjusting weights, the weights of each modal feature can be adjusted based on the attention mechanism.

[0041] For example, a multimodal model based on the Transformer architecture learns the correlations between different modal data and adaptively adjusts weights. For example, a multi-head attention mechanism is used to perform a weighted summation of features from different modalities: F comprehensive feature vector = Attention(F image, F audio, F location, F environment). The resulting comprehensive feature vector incorporates information from image, audio, location, and environmental data, enabling a more comprehensive description of species characteristics. Furthermore, the attention mechanism dynamically adjusts the weights of each modal feature based on the importance and relevance of the data, performing a weighted fusion to generate a comprehensive feature vector. This provides a highly reliable comprehensive feature vector, facilitating subsequent species analysis.

[0042] As a method of dynamically adjusting weights, the attention mechanism adaptively allocates weights by calculating the similarity or correlation between input features. In the case of generating a baseline feature vector, the embodiment of the present application can also dynamically adjust the weights of each modal feature in the baseline feature vector based on the target parameter, and control the amplitude of the weight change based on the weight smoothing biological constraint mechanism. After adjusting the weights of each modal feature in the baseline feature vector, the comprehensive feature vector is determined. The target parameters include at least one of the following: the full modal weight corresponding to the system time period, the confidence of the target large model for each modal feature, the user's feedback information on the recognition results output by the target large model in the previous recognition cycle, the newly introduced sample data in the current recognition cycle, and the scene in which the current recognition cycle is located. And the method of adjusting weights based on target parameters can be used in combination with the attention mechanism to achieve a more flexible and effective weight adjustment strategy.

[0043] In the implementation of this solution, a comprehensive feature vector generation method compatible with large models such as Transformer is specifically designed to uniformly encode unstructured biological data and structured environmental data, solving the problem of large models adapting to multimodal heterogeneous data input.

[0044] As an optional implementation, the process of dynamically adjusting the weights of each modal feature in the baseline feature vector based on the target parameter to generate a comprehensive feature vector is as follows: Figure 2 As shown: After obtaining a baseline feature vector by fusing the initial weights with the concatenated high-dimensional feature vectors, the weights of each modal feature are dynamically adjusted based on the target large model's confidence in each modal feature, user feedback on the recognition results, newly introduced sample data during the current recognition cycle, and the current scenario. After completing these dynamic weight adjustments, a comprehensive feature vector is generated.

[0045] The following introduces the process of dynamically adjusting the weights of each modal feature based on the target parameters.

[0046] When dynamically adjusting the weights of each modal feature in the baseline eigenvector, adjustments are made based on the all-modal weights corresponding to the current system time period, enabling periodic weight adjustments based on time adaptation. Time-adaptive periodic weight adjustments determine the frequency of weight adjustments based on different environmental conditions or specific patterns (such as diurnal or seasonal variations). This approach essentially triggers weight updates at fixed intervals based on the system clock. For example, if the base period is 30 seconds, the time window can be configured to update weights every 30 seconds. During weight updates, the weights of all modalities are updated within each time window.

[0047] When dynamically adjusting the weights of each modal feature based on the target macromodel's confidence in each modal feature, the target macromodel's confidence in each modal feature is calculated during each recognition cycle. When the confidence of a modal feature falls below a first threshold, the weight of that modal feature is downgraded, and the weights of the remaining modal features are compensatorily enhanced. Confidence is calculated based on indicators such as the recognition accuracy and error rate of each modal feature by the target macromodel. For environmental features, their confidence is also related to the stability and relevance of the environmental data. When performing downgrading, an exponential decay method is used, for example, where the decay coefficient is negatively correlated with the confidence score.

[0048] When dynamically adjusting the weights of each modal feature based on user feedback on the recognition results, user feedback is obtained, including the user's evaluation of the accuracy of the recognition results and their subjective judgment of the importance of different modal features. The weights of each modal feature are adjusted based on this user feedback. This approach allows the large model to better adapt to user expectations and the needs of actual application scenarios. When adjusting the weights of each modal feature based on user feedback, the following strategies can be adopted: 1. Weight adjustment based on accuracy evaluation: If the user's evaluation of the accuracy of the recognition result is low, the weight of certain modal features in the large model can be increased to improve the performance of the large model. 2. Weight adjustment based on modal feature importance judgment: The weights of each modal feature are adjusted based on the user's subjective judgment of the importance of different modal features. 3. Comprehensive weight adjustment: Combining weight adjustments based on accuracy and importance to comprehensively adjust the weights of each modal feature.

[0049] When dynamically adjusting weights based on newly introduced sample data during the current recognition cycle, obtain the newly introduced sample data during the current recognition cycle, such as new sample data introduced during autumn migration, and the new sample data includes new biological data and environmental data. Weights are then adjusted based on the following operations: 1. Evaluate the impact of the new data; calculate the impact of the new data on model recognition, for example, by calculating the confidence level or prediction error of the new data. 2. Determine the direction of weight adjustment; determine the direction and magnitude of weight adjustment for each modal feature based on the impact of the new data. Update the weights; update the weights of each modal feature based on the determined adjustment direction and magnitude.

[0050] Dynamically adjusting weights based on the current recognition cycle requires contextual identification. For example, scenarios include the wild, laboratory, and nature reserves, where different target species exist. The importance of biological and environmental features varies across different scenarios. For example, in the wild, environmental features contribute more to species identification. Scenarios can also include sudden changes in environmental parameters, specific events (such as seasonal migration), and temporal fluctuations. Sudden changes in environmental parameters can cause a step-change in the quality of a particular modality (e.g., a sharp drop in audio signal-to-noise ratio during heavy rain). This nonlinear change requires a dynamic weight adjustment mechanism to address it. Specific events (such as seasonal migration) can trigger updates to multimodal feature weights. During migration, environmental and biological characteristics can significantly change, impacting model performance. By dynamically adjusting the weights of each modal feature, the model can better adapt to data characteristics during migration. Temporal fluctuations (such as the transition from day to night) can also trigger updates to multimodal feature weights. The transition between day and night can cause changes in light, temperature, behavioral patterns, and other factors, which require weight adjustments to accommodate these changes.

[0051] The following introduces the process of dynamically adjusting weights through several specific examples.

[0052] 1. Image feature weight adjustment driven by light intensity. The target species is birds. It relies on image features during the day and audio features at night.

[0053] 1. Input data: environmental data (light intensity L), image features (extracted image feature vector V img), audio features (audio feature vector V audio). 2. Confidence calculation: Calculate the image confidence score based on the light intensity L B img, the relationship between the image confidence score and the light intensity is an S-shaped curve, and the specific formula is as follows:

[0054] in, L 0 is the light threshold,k is the slope parameter (used to control the transition speed). L 0 is the critical value at which the species' visual ability significantly decreases when the light intensity is lower than this threshold, determined experimentally; k Nocturnal animals k Smaller values (smooth transition).

[0055] 3. Weight distribution; based on image confidence score Assign weights to image features and audio features, image feature weights = , audio feature weight =1- .

[0056] By dynamically adjusting the weights of image features and audio features according to light intensity, the model can better adapt to different conditions during the day and night. In this way, the model can more accurately identify and analyze bird behaviors and characteristics.

[0057] 2. Audio weight adjustment driven by environmental noise: strong wind or rainfall causes the audio signal-to-noise ratio to decrease.

[0058] 1. Input data: environmental data (noise decibel value N), audio features (Mel spectrum features V audio), positioning data (positioning feature vector V location).

[0059] 2. Signal-to-noise ratio evaluation; defining audio quality coefficients Q audio, is inversely proportional to the noise decibel value N. The specific formula is as follows:

[0060] in, N : Current ambient noise decibel value (measured value). N max: The maximum noise threshold allowed by the system (above which the audio is completely ineffective). This threshold is determined by experimentally measuring the audio recognition threshold of the target species. Examples include bird song recognition (above which the song is masked by ambient noise) and underwater sonar (which relies on the sound absorption properties of water).

[0061] ε is the attenuation factor (a nonlinear parameter that controls the speed of quality degradation, ε≥1). Specifically, the audio recognition rate data of the target species in a noisy environment can be obtained; the optimal ε is fitted by nonlinear regression, so that Q The Pearson correlation coefficient between audio and recognition rate is maximized. Different ε values can be set for different species (e.g., ε=2.0 for songbirds and ε=1.5 for frogs).

[0062] : Ensure that the result is non-negative (when N ≥ N When max, Q audio=0).

[0063] 3. Weight correction; based on audio quality coefficient Q audio correction audio weight W audio, if Q audio<0.5, then force the audio weight to be reduced to W audio=0.2 and increase the weight of positioning features.

[0064] By dynamically adjusting the weights of audio features based on ambient noise, the model can better adapt to adverse conditions such as strong winds or rainfall. In this way, the model can more accurately identify and analyze the behaviors and characteristics of target species.

[0065] 3. Dynamic Adjustment of Positioning Feature Weights 1. Scenario Requirements The reliability of positioning data (such as GPS / Beidou) varies significantly in the following scenarios: Time sensitivity: The positioning weight of migrating birds at night is higher than that during the day.

[0066] Path deviation: The weight needs to be reduced when deviating from the historical migration route.

[0067] Signal quality: Urban canyons or cloudy weather can increase positioning errors.

[0068] 2. Weight calculation model input data Environmental parameters: time (day or night), positioning error radius (meters), historical path deviation (%).

[0069] Biological characteristics: typical activity radius of the species (e.g. migratory birds = 50km, resident birds = 5km).

[0070] 3. Weight calculation formula

[0071] Migration data analysis revealed that the weight of nighttime needs to be increased by 20% to compensate for positioning dependency. Consider setting coefficient 1 to 1.2. The path weight is linearly positively correlated with the positioning weight. Consider setting coefficient 2 to 1. For every 100-meter increase in error penalty, the weight decreases by 0.2 (upper limit 0.5). Consider setting coefficient 3 to 0.2. However, in this case, the regression fitting coefficients are unconstrained, resulting in the following: W gps The numerical range of does not match the weights of modalities such as image and audio.

[0072] To preventW The GPS value is too large. Ensure that the magnitude of each feature is consistent during multimodal fusion, limit the total impact of time and path weights, and avoid over-reliance on positioning data. You can set adjustment constraints: coefficient 1 + coefficient 2 = 1.2, so that the relative proportion of the regression relationship can be preserved, and coefficient 3 = 0.2;

[0073] PathWeight = 1-min (1, ) ErrorPenalty = min(0.5, ).

[0074] 4. Calculation Example Scenario The time is 20:00 at night, the deviation from the historical path is 15 km, the positioning error is 80 m, and the species activity radius is 50 km; Time Weight = 0.7 (nighttime), Path Weight = 1-min(1, 15 / 50) = 0.7, Error Penalty = min(0.5, 80 / 100) = 0.5. Then the positioning weight .

[0075] In the above example, the time weight is adjusted based on the time of day (daytime or nighttime), the path weight is adjusted based on the distance from the historical path, and the error penalty is adjusted based on the positioning error. The positioning weight is calculated by combining the time weight, path weight, and error penalty. By dynamically adjusting the positioning data weight based on environmental parameters and biometric characteristics, the model can better adapt to the changes in positioning data reliability in different scenarios.

[0076] 4. Weight Optimization of Multimodal Dynamic Fusion 1. Fusion rules The input is: image weight for instance one, audio weight for instance two, and positioning weight for instance three.

[0077] Dynamic compensation: If the positioning error is greater than 200 meters, the positioning weight will be forced to W gps Limit it, for example, limit it to within 0.3. When the image and audio weights are both less than the set value (such as 0.2), increase the positioning weight by, for example, 50%.

[0078] 2. Normalization formula The normalization formula is used to adjust the weights so that their sum is 1 for fusion calculation:

[0079] Wi is the original weight of the ith mode calculated from the environmental data, C i is the compensation coefficient (weight amplification factor in emergency mode), C i The value depends on the mode, emergency compensation mode C i = 1.5, Normal mode C i = 1; is the total weighted effective contribution of all modes, W i′ is the final weight after normalization.

[0080] For example, consider the following input: image weight = 0.15 (heavy fog), audio weight = 0.1 (heavy rain), and positioning weight = 0.8 (error = 30 meters). Since both the image and audio weights are less than 0.2, the positioning weight is increased by 50% to 0.8*1.5=1.2.

[0081] In the above scenario, the new positioning weight is calculated .

[0082] In this example, in extreme weather conditions, the system can automatically switch to positioning-dominant mode, increasing the weight of positioning to adapt to different working modes.

[0083] The above four examples introduce the dynamic adjustment of feature weights in different situations. Based on the dynamic weight adjustment mechanism, a more comprehensive multimodal data fusion solution is provided, which optimizes the generation of comprehensive feature vectors.

[0084] It should be noted that in the process of adjusting weights based on target parameters, the amplitude of weight changes can be controlled based on the weight smoothing biological constraint mechanism. Weight smoothing biological constraints refer to intelligently controlling the rate and amplitude of weight changes by introducing species-specific behavioral patterns and physiological characteristics, thereby avoiding non-physiological mutations caused by pure mathematical optimization. Weight smoothing processing includes, for example, the following biological constraints: 1. Loading a preset smoothing coefficient based on the target species type; 2. Constraining the amplitude of weight changes between adjacent time frames to not exceed the maximum fitness value of the species; 3. Enabling enhanced smoothing mode during special physiological stages such as the breeding season / migration period. Among them, it should be noted that when the weight change triggers a physiological limit alarm, it can automatically switch to a conservative adjustment mode.

[0085] The rationality of weighted smoothing biological constraints is demonstrated below by comparing them with pure mathematical smoothing. The two are compared in the following dimensions: In the objective function dimension: Traditional exponential smoothing aims to minimize the mean square error, pays more attention to mathematical fitting accuracy, and is suitable for general data smoothing needs; biologically constrained smoothing aims to conform to the behavior patterns of species. When processing data, it is necessary to consider the physiological and behavioral characteristics of the species to ensure that the smoothed data can truly reflect the actual behavior of the species.

[0086] In terms of parameter source: in traditional exponential smoothing, parameters (such as smoothing coefficients) are mainly obtained through statistical fitting of data, and historical data are usually used for regression analysis or other statistical methods to determine the optimal smoothing coefficients; in biologically constrained smoothing, parameters are derived from ecological research consensus and field observation data, including not only the results of laboratory studies, but also long-term observation data of species in natural environments to ensure the biological rationality of the parameters.

[0087] In the dimension of dynamic adaptability: in traditional exponential smoothing, the smoothing coefficient is fixed. Once determined, it remains unchanged throughout the entire data processing process. It is suitable for scenarios where data changes are relatively stable. In biologically constrained smoothing, the smoothing coefficient is dynamically adjusted according to the physiological state of the species. For example, when the species is in different physiological stages (such as the breeding season, migration period, etc.), the smoothing coefficient will change accordingly to better reflect the actual behavioral changes of the species.

[0088] In the dimension of anomaly processing, traditional exponential smoothing may lead to results that violate biological laws when dealing with outliers. For example, mutations in the data may be smoothed, but this treatment may not conform to the actual physiological characteristics of the species. Biologically constrained smoothing will force compliance with physiological limits, that is, the physiological limitations of the species will be considered when processing data to avoid results that are inconsistent with biological common sense.

[0089] For example, with traditional processing, the weight of a particular bird's audio signal dropped rapidly from 0.5 to 0.2 in 0.3 seconds, interrupting song analysis. This rapid change likely exceeds the species' actual physiological capabilities, leading to anomalies in data processing and preventing effective analysis. However, the biologically constrained smoothing mechanism, with a maximum drop of 0.1 / second and a smooth transition of 1.5 seconds, maintains continuous observation. By limiting the maximum rate of data change, it ensures that data changes conform to the species' physiological characteristics, thus avoiding analysis interruptions caused by sudden changes and maintaining the continuity and reliability of data processing.

[0090] To further illustrate the weight smoothing biological constraint, a specific example of the weight smoothing biological constraint is given below: 1. Adjustment of the weight of raptors (eagles) in dark clouds Environmental conditions: Overcast sky, reduced visibility; image quality degraded due to reduced light intensity.

[0091] Biological constraints: The typical reaction time of a bird of prey (hawk) is, for example, 0.1-0.5 seconds, with a smoothing factor ranging from 0.1-0.3, allowing it to quickly switch to infrared mode when encountering sudden dark clouds.

[0092] Weight Adjustment: Due to reduced visibility and image quality, the image weight needs to be reduced. Assuming an initial image weight of 0.6, based on biological constraints, the image weight adjustment rate should be controlled between 0.1 and 0.3, gradually reducing the image weight to around 0.3. Due to the switch to infrared mode, the infrared weight needs to be increased rapidly. Assuming an initial infrared weight of 0.2, it can be increased to around 0.6 within a short period of time (e.g., 0.1 seconds). The audio weight remains unchanged or is adjusted appropriately based on ambient noise.

[0093] 2. Weight adjustment of wading birds (cranes) during light gradient Environmental conditions: Gradual changes in light intensity.

[0094] Biological constraints: Typical reaction times for wading birds (cranes) are, for example, 2-5 seconds, with a smoothing factor in the range of 0.6-0.8, which slowly reduces the visual weight as the lighting changes.

[0095] Weight Adjustment: As lighting changes, the image weight needs to be slowly reduced. Assuming an initial image weight of 0.7, based on biological constraints, it can be gradually reduced to around 0.4 over 2-5 seconds. The audio weight can be appropriately increased to compensate for the reduction in visual information. Assuming an initial audio weight of 0.2, it can be gradually increased to around 0.4. The localization weight remains unchanged or adjusted as needed.

[0096] 3. Adjustment of songbird (finch) weighting when wind noise increases Environmental conditions: Increased wind speed, increased ambient noise, and decreased audio quality.

[0097] Biological constraints: Typical reaction time of songbirds (e.g., finches) is 1-2 seconds, smoothing factors range from 0.4-0.6, and audio weighting is adjusted at a moderate rate when wind noise increases.

[0098] Weight Adjustment: Due to increased wind noise, the audio weight needs to be reduced. Assuming the initial audio weight is 0.5, it can be reduced to around 0.2 within 1-2 seconds. The image weight can be increased appropriately to compensate for the reduction in audio information. Assuming the initial image weight is 0.3, it can be gradually increased to around 0.6. The positioning weight remains unchanged or adjusted as needed.

[0099] The above example shows how to dynamically adjust the weights of image, audio, and positioning data based on the species' biology and environmental conditions. By following biological constraints, the weight adjustment process can be ensured to be consistent with the species' physiological and behavioral characteristics.

[0100] The following describes the process of species identification based on a large model. The process involves performing inference analysis on the comprehensive feature vector based on the target large model, obtaining at least one species name corresponding to the target species and the confidence level of each species name, and determining the species name of the target species based on the confidence level of each species name. The process includes: The target large model is optimized according to the comprehensive feature vector corresponding to the target species, and the comprehensive feature vector is inferred and analyzed based on the optimized target large model; according to the inference analysis results output by the optimized target large model, at least one species name corresponding to the target species and the confidence of each species name are obtained; based on the prior information of the target species, the confidence corresponding to the at least one species name is corrected; and the species name of the target species is determined based on the confidence of the corrected species name.

[0101] After obtaining the comprehensive feature vector and before using the target large model to perform inference analysis on the comprehensive feature vector, the target large model needs to be optimized based on the comprehensive feature vector corresponding to the target species. After optimizing the target large model, the optimized target large model performs inference analysis on the comprehensive feature vector and outputs the inference analysis results, providing at least one species name corresponding to the target species and the confidence level of each species name.

[0102] After obtaining at least one species name corresponding to the target species and the confidence level of each species name, the confidence level of the species name is revised based on the prior information of the target species. Prior information of the target species refers to information about the characteristics, behavior, distribution, etc. of the species that is known or assumed before specific observations or experiments are conducted. This information usually comes from historical data, literature research, expert knowledge, or statistical analysis. Prior information helps to better understand and predict the behavior and distribution of species. After using the prior information to revise the confidence level of the species name, the species name of the target species is determined based on the revised confidence level of at least one species name. For example, the species name with the highest confidence level is determined as the final species name.

[0103] Among them, when optimizing the target large model according to the comprehensive feature vector corresponding to the target species, and performing reasoning analysis on the comprehensive feature vector based on the optimized target large model, it includes: The comprehensive feature vector corresponding to the target species is fused with the feature representation of the target large model, and the weights of different modal features are dynamically adjusted to optimize the target large model; Based on the optimized target large model, the enhanced comprehensive feature vector is subjected to inference analysis, or based on the optimized target large model, the key fusion features are subjected to inference analysis; wherein the key fusion features are generated by fusing the multi-dimensional key features extracted from the comprehensive feature vector.

[0104] When optimizing the target large model, optimization can be performed from two dimensions: 1. Fusion of the comprehensive feature vector corresponding to the target species with the feature representation of the target large model to enhance the target large model's ability to understand multimodal data; for example, the fusion method is: aligning the comprehensive feature vector with the input feature format of the target large model, and mapping the comprehensive feature vector to the feature space of the target large model through linear transformation or nonlinear mapping to achieve feature fusion. 2. Dynamically adjust the weights of different modal features based on the attention mechanism to optimize the model to improve the model's recognition accuracy.

[0105] When optimizing the target large model based on feature fusion, model parameters can be adjusted or not. For the case where model parameters are not adjusted: 1. The comprehensive feature vector is directly concatenated to the feature representation of the target large model to form a longer feature vector. In this case, the model parameters remain unchanged, but the richer feature input is used in subsequent tasks. 2. The comprehensive feature vector is added to the feature representation of the target large model with a certain weight. In this case, the model parameters also remain unchanged, but the contribution of the features is adjusted by the weight.

[0106] Regarding model parameter adjustments: 1. Fine-tuning: After feature fusion, fine-tune the target large model to allow the model to learn how to better process the fused features. In this case, the model parameters of the target large model will be adjusted based on the new training data. 2. End-to-end training: Embed the feature fusion process into the model training process, allowing the model to learn how to fuse features and perform tasks. In this case, the model parameters of the target large model will also be adjusted based on the training objectives.

[0107] In the attention mechanism, weights are calculated dynamically based on the input features. When dynamically adjusting the weights of different modal features for model optimization based on the attention mechanism, additional learnable weights can be introduced to adjust the contributions of different modal features. These weights are independent of the model parameters. For example, a learnable weight matrix can be introduced to adjust the weights of different modal features, and the parameters of this weight matrix are learned independently.

[0108] When optimizing a model based on the attention mechanism, model parameters can also be adjusted. 1. Fine-tuning: In the attention mechanism, if the model parameters of the target large model are fine-tuned, the model will learn how to better adjust the weights of different modal features. In this case, the model parameters of the target large model will be adjusted based on the new training data. 2. End-to-end training: The attention mechanism is embedded in the model training process, allowing the model to learn how to adjust the weights of different modal features. In this case, the model parameters of the target large model will also be adjusted according to the training objectives.

[0109] After optimizing the target large model, when using it to process the comprehensive feature vector, the comprehensive feature vector can first be enhanced, such as through feature normalization and dimensionality reduction, to improve feature robustness. The enhanced comprehensive feature vector can then be inferred using the target large model, with the inference analysis results output. Alternatively, multidimensional key features can be extracted from the comprehensive feature vector, such as image texture features, audio frequency features, positioning coordinate features, and environmental temperature features. These extracted multidimensional key features can be fused to generate a richer feature representation, i.e., key fused features. Inference analysis can then be performed on the key fused features based on the optimized target large model, with the inference analysis results output.

[0110] During the implementation of the above-mentioned species identification based on the target large model, the target large model is optimized to enhance the target large model's ability to understand multimodal data and improve the accuracy of model identification. By processing the enhanced comprehensive feature vector based on the optimized target large model or processing the key fusion features generated based on the comprehensive feature vector, species identification can be performed relatively accurately and the efficiency of species identification can be guaranteed.

[0111] The following describes the process of training the target large model. The model training includes the following steps: Pre-training an architecture model suitable for multimodal data processing based on first sample data corresponding to the first biological sample set to determine a pre-trained model; Adjusting model parameters of the pre-trained model based on the data-labeled second sample data corresponding to the second biological sample set to fine-tune the pre-trained model; Migrating common features associated with the pre-trained model to the pre-trained model to optimize the pre-trained model; The pre-trained model that has undergone model fine-tuning and model optimization is pruned and quantized to determine the target large model for species identification.

[0112] The embodiments of the present application select Transformer architecture models suitable for multimodal data processing, such as ViT (Vision Transformer), CLIP (Contrastive Language-Image Pre-training), Audio-CLIP, and Geo-CLIP. ViT is suitable for image data processing and can learn global features in images; CLIP is suitable for multimodal data processing and can jointly learn the feature representations of images and texts, and is suitable for species identification tasks; Audio-CLIP adds the processing capabilities of the audio modality on the basis of CLIP, and can jointly learn the feature representations of images, texts, and audios, further improving the model's ability to understand multimodal data; Geo-CLIP is based on CLIP and combines positioning data and environmental data to further enhance the model's ability to understand geographic and environmental contexts. Preferably, the selected Transformer architecture model suitable for multimodal data processing has the ability to understand audio, positioning data, and environmental data while being able to jointly learn the feature representations of images and texts.

[0113] During the pre-training phase, a large-scale multimodal dataset is collected. Specifically, image datasets and image-related text data, such as image descriptions, are required; audio datasets are collected, along with location data (extracting geographic location information for images and audio), and environmental data (environmental information related to the time the images and audio were captured, such as weather, temperature, and humidity). Multimodal data is pre-processed, such as cropping, scaling, and normalizing image data to ensure consistency in size and format; text data is segmented and encoded to convert text into a numerical form that the model can process; audio data is processed, such as extracting features like mel-spectrograms and converting the audio into a format suitable for model processing; and location data is converted into a format suitable for model processing, such as embedding geographic information like longitude and latitude into image or audio features or as a standalone feature input. Environmental data such as weather, temperature, and humidity are standardized or normalized to facilitate integration with the image and audio data. After data processing, the pre-trained model is determined based on the pre-training task and the processed data.

[0114] After finalizing the pre-trained model, model fine-tuning is required. This phase requires collecting data related to plants and animals, specifically wildlife. During data collection, a large amount of wildlife image data is collected, including samples from different species and environments, ensuring that the image data covers a variety of scenarios, such as different lighting conditions, seasons, and geographic locations. Audio data corresponding to the images, such as animal calls and environmental sounds, is collected, ensuring that the audio data is aligned with the image data in both time and space. Related text descriptions, such as species names, feature descriptions, and behavioral descriptions, are collected, ensuring that the text descriptions are aligned with the image and audio data. Environmental data related to the image and audio data, such as weather, temperature, and humidity, is collected, ensuring that the environmental data is aligned with the image and audio data in both time and space. Location data, such as the latitude and longitude of the shooting location, is collected, ensuring that the location data is aligned with the image and audio data. After data collection, the data is annotated, and the model fine-tuning is performed based on this annotated data. During model fine-tuning, some layers of the pre-trained model are frozen to preserve the common feature representations learned during pre-training. The final layers of the model are trained to adapt to the wildlife recognition task, reducing computational effort while leveraging the common features of the pre-trained model. During fine-tuning, image, audio, text, environmental, and location data are combined to enhance the model's understanding of multimodal data. Cross-entropy loss is used for training to optimize model parameters.

[0115] After fine-tuning the model, common features associated with the pre-trained model are transferred to the pre-trained model to optimize it. For example, knowledge from a pre-trained model in other fields can be transferred to the task of wildlife species identification. Through transfer learning, the model's adaptability and generalization capabilities for wildlife data are improved, further enhancing the model's ability to understand multimodal data.

[0116] After fine-tuning the model and optimizing it through feature transfer, the pre-trained model is pruned and quantized to compress the model to determine the target large model for species identification. During pruning, the pre-trained model is pruned to remove unimportant weights or neurons; structured pruning methods are used to preserve the structural integrity of the model; pruning reduces the number of model parameters and improves computational efficiency. During quantization, such as using quantization techniques, the model's weights and activation values are quantized to a low-precision representation. Pruning and quantization can reduce the model's parameter count and storage space, improving computational efficiency. After pruning and quantization, the model's performance needs to be evaluated on a validation set to ensure that the pruning and quantization operations have not significantly reduced the model's accuracy.

[0117] Through the above process, we can select a Transformer architecture model suitable for multimodal data processing for pre-training and fine-tuning, use transfer learning technology to improve the generalization ability of the model, and reduce the number of model parameters through model compression, thereby improving the computational efficiency and storage efficiency of the model.

[0118] After the target large model is determined through model training, the following steps are also included: Generate comprehensive feature vectors and labels associated with the samples based on new sample data provided by the acquisition device, and update model parameters of the target large model using an online learning algorithm based on the comprehensive feature vectors and labels associated with the samples; and / or Obtain annotation information after the user annotates the recognition result output by the target large model, add the annotation information to the sample data to update the sample data, and adjust the model parameters of the target large model based on the updated sample data.

[0119] In the embodiments of the present application, an online learning algorithm and / or an incremental learning mechanism may be used to adjust model parameters to optimize model performance. When using an online learning algorithm to update model parameters, an online gradient descent algorithm or a stochastic gradient descent algorithm is selected. After acquiring new sample data provided by an acquisition device, a comprehensive feature vector and label associated with the sample are generated. Based on the comprehensive feature vector and the corresponding label, an online learning algorithm is used to update the model parameters of the target large model, such as using an online gradient descent algorithm to update the model parameters to optimize model performance.

[0120] When adjusting model parameters based on incremental learning, the model incorporates the annotated data into model training when users annotate or provide feedback on recognition results, continuously optimizing model performance. During this process, users annotate or provide feedback on recognition results through an interactive interface, collecting user-annotated data, including image, audio, location, and environmental data, along with their corresponding labels. This user-annotated data is then added to the training dataset, and the model is fine-tuned using an incremental learning algorithm, for example, mini-batch gradient descent.

[0121] Through online learning algorithms and incremental learning mechanisms, the model can be dynamically learned and updated in real time. The online learning algorithm allows the model to receive new data in real time and dynamically adjust parameters, quickly adapting to environmental changes and the emergence of new species. The incremental learning mechanism allows the model to incorporate annotated data into model training as users annotate or provide feedback on recognition results, continuously optimizing model performance. This model optimization approach can significantly improve the model's adaptability and accuracy, providing stronger support for species identification.

[0122] In addition to identifying wild animal and plant species, the embodiments of the present application can also analyze species behavior patterns and predict future behavioral trends based on species identification results and historical data, providing deeper insights for ecological research; and, by combining environmental data with wild animal and plant distribution and behavior information, evaluate the impact of environmental changes on biodiversity and provide a scientific basis for ecological protection decisions.

[0123] For identified species, the following steps need to be performed when analyzing their behavior patterns and predicting future behavior trends by combining identification results with historical data: 1. Data collection: Collect species identification results, behavioral data (such as movement trajectory, activity time) and environmental data, and annotate behavioral data, such as foraging, migration, reproduction, etc.

[0124] 2. Behavioral analysis: extracting behavioral characteristics of species, such as movement speed and activity frequency, and using timing analysis models or reinforcement learning algorithms for behavioral modeling.

[0125] 3. Behavior prediction: Use the trained model to predict behavior and output future behavior trends.

[0126] The main task of behavioral analysis and prediction is to analyze behavioral patterns based on species identification results and historical data, and predict future behavioral trends. The following is a detailed description of this process: 1. Data Collection and Labeling (1) Data collection Based on species identification, information such as species name and confidence level is obtained; behavioral data of species (taking animals as an example) is collected, such as movement trajectory, activity time, behavior type, etc.; and environmental data related to behavior is collected, such as temperature, humidity, water quality, etc.

[0127] (2) Data labeling Behavioral labeling: Label the collected behavioral data, such as foraging, migration, reproduction, etc.; Timestamp alignment: Ensure that all data (identification results, behavioral data, environmental data) are accurately timestamped for time series analysis.

[0128] 2. Feature Extraction (1) Behavioral characteristics Extract the movement trajectory of the species, including speed, acceleration, direction change, etc.; count the activity frequency of the species in different time periods; and record the duration of each behavior.

[0129] (2) Environmental characteristics Record changes in ambient temperature, record changes in ambient humidity, record changes in water quality, such as pH value, dissolved oxygen, etc.

[0130] 3. Model Selection and Training (1) Model selection Time series analysis model: Use LSTM (Long Short-Term Memory Network) or GRU (Gated Recurrent Unit) for time series analysis; Reinforcement learning model: Use reinforcement learning algorithm for behavior modeling.

[0131] (2) Model training Data preparation: Combine the extracted features and labeled behavioral data into a training dataset; during model training, use the training dataset to train the time series analysis model or reinforcement learning model.

[0132] 4. Behavior Prediction Model inference: Input current state: Input the current species behavior characteristics and environmental characteristics into the trained model. Output prediction result: The model outputs the prediction result of future behavior, such as the behavior type in the next time period.

[0133] When combining environmental data with information on the distribution and behavior of wild animals and plants to assess the impact of environmental change on biodiversity, data collection and comprehensive analysis are necessary. This analysis includes correlation and regression analysis. Correlation analysis can help identify the relationship between environmental variables and the distribution and behavior of wild animals and plants, while regression analysis can further assess the impact of environmental change on biodiversity. These analyses can provide a scientific basis for ecological protection and resource management.

[0134] During data collection, collect environmental data and wildlife data. Environmental data includes, for example, temperature (water temperature and air temperature), air humidity, water quality, and other environmental data (such as wind speed, wind direction, and light intensity). Wildlife data includes distribution data and behavioral data. Distribution data includes, for example, species distribution and species abundance. Behavioral data includes, for example, behavior type (foraging, migration, reproduction, etc.), behavior frequency, and behavior duration. All collected data must include accurate timestamps and geographic location information to facilitate subsequent analysis.

[0135] Correlation analysis and regression analysis are used in data analysis. When conducting correlation analysis, statistical analysis methods (such as the Pearson correlation coefficient) are used to analyze the correlation between environmental data and the distribution and behavior of wild animals and plants. Specifically, data preprocessing is performed, such as cleaning and standardizing the data to ensure data quality; correlation coefficients are calculated using statistical software; and correlation analysis results are presented using visualization tools such as scatter plots and heat maps. When conducting regression analysis, regression models (such as linear regression and logistic regression) are used to assess the impact of environmental changes on biodiversity. Specifically, data preprocessing is performed, such as cleaning and standardizing the data to ensure data quality; an appropriate regression model is selected and the model is trained using a training dataset; model performance is evaluated using a validation dataset and appropriate evaluation metrics are selected; the coefficients of the regression model are interpreted to assess the impact of environmental variables on biodiversity.

[0136] The embodiment of this application aims to use large model technology to build an efficient and intelligent wild animal and plant species identification system to achieve accurate species identification, behavior analysis and ecological environment monitoring, and provide strong support for biodiversity conservation.

[0137] The present application embodiment provides a system for identifying wild animal and plant species, such as Figure 3 As shown, the wildlife species identification system 300 includes: a collection device 301 and a cloud platform 302; The collection device 301 collects biological data and environmental data associated with the target species, wherein the biological data includes image data, audio data, and positioning data associated with the target species, and the environmental data includes at least meteorological data and physical environment data associated with the living environment of the target species, wherein the target species is a wild animal or a wild plant; The cloud platform 302 pre-processes the data provided by the acquisition device 301, obtains the data to be analyzed corresponding to the target species, extracts features from the data to be analyzed, and aligns the extracted multimodal features based on time. The multimodal features include biological features and environmental features. The biological features include image features extracted from the image data, audio features extracted from the audio data, and positioning features extracted from the positioning data. The cloud platform 302 performs feature fusion on the aligned multimodal features based on the dynamically adjusted weights of each modal feature to generate a comprehensive feature vector indicating comprehensive feature information of the target species. The weights of each modal feature are dynamically adjusted based on the real-time contribution of the multimodal features to species identification. The cloud platform 302 performs reasoning analysis on the comprehensive feature vector based on the deployed target large model, obtains at least one species name corresponding to the target species and the confidence of each species name, and determines the species name of the target species based on the confidence of each species name.

[0138] Optionally, when performing data preprocessing and obtaining data to be analyzed, the cloud platform 302 is also used to: perform image denoising, image enhancement, and image standardization on image data associated with the target species to obtain images to be analyzed; remove background noise from audio data associated with the target species to obtain audio to be analyzed; convert the latitude and longitude coordinates of the target species into plane coordinates, and obtain positioning data to be analyzed in combination with data measured by the inertial measurement unit; and standardize the environmental data associated with the target species to a unified range to obtain environmental data to be analyzed.

[0139] Optionally, when performing feature alignment, the cloud platform 302 is also used to: extract image features, audio features, positioning features, and environmental features with timestamps from the data to be analyzed; align the extracted features of different modalities according to timestamps to obtain multimodal features that remain consistent in the time dimension.

[0140] Optionally, when generating the comprehensive feature vector, the cloud platform 302 is further used to: perform feature splicing on the aligned image features, audio features, positioning features, and environmental features to generate a high-dimensional feature vector; assign an initial weight to each modal feature in the multimodal features, fuse the assigned initial weights into the high-dimensional feature vector, and generate a baseline feature vector, wherein the initial weight of the environmental feature is determined based on the degree of influence of the environmental data on the survival and identification of the target species, and the initial weight of each modal feature in the biological feature is assigned based on the estimated importance of each modal feature in species identification; Based on the target parameters, the weights of each modal feature in the baseline feature vector are dynamically adjusted, and the weight change amplitude is controlled based on the weight smoothing biological constraint mechanism. A comprehensive feature vector is generated according to the baseline feature vector with dynamically adjusted weights. The target parameters include at least one of the following: the full modal weight corresponding to the system time period, the confidence of the target large model for each modal feature, the user's feedback information on the recognition results output by the target large model in the previous recognition cycle, the newly introduced sample data in the current recognition cycle, and the scenario in which the current recognition cycle is located.

[0141] Optionally, when determining the species name of the target species, the cloud platform 302 is further configured to: optimize the target macro model according to the comprehensive feature vector corresponding to the target species, and perform inference analysis on the comprehensive feature vector based on the optimized target macro model; and obtain at least one species name corresponding to the target species and the confidence level of each species name according to the inference analysis result output by the optimized target macro model; Modifying the confidence level corresponding to at least one species name based on prior information of the target species; The species name of the target species is determined based on the confidence of the revised species name.

[0142] Optionally, the cloud platform 302 is further configured to: fuse the comprehensive feature vector corresponding to the target species with the feature representation of the target large model, and dynamically adjust the weights of different modal features to optimize the target large model; Based on the optimized target large model, the enhanced comprehensive feature vector is subjected to inference analysis, or based on the optimized target large model, the key fusion features are subjected to inference analysis; wherein the key fusion features are generated by fusing the multi-dimensional key features extracted from the comprehensive feature vector.

[0143] Optionally, the cloud platform 302 is also used to: pre-train an architecture model suitable for multimodal data processing based on first sample data corresponding to the first biological sample set to determine the pre-trained model; adjust model parameters of the pre-trained model based on second sample data corresponding to the second biological sample set and after data annotation to fine-tune the pre-trained model; migrate common features associated with the pre-trained model to the pre-trained model to optimize the pre-trained model; prune and quantize the pre-trained model that has undergone model fine-tuning and model optimization to determine a target large model for species identification.

[0144] Optionally, the cloud platform 302 is further configured to: generate a comprehensive feature vector and label associated with the sample based on the new sample data provided by the acquisition device 301, and update the model parameters of the target large model using an online learning algorithm based on the comprehensive feature vector and label associated with the sample; and / or Obtain annotation information after the user annotates the recognition result output by the target large model, add the annotation information to the sample data to update the sample data, and adjust the model parameters of the target large model based on the updated sample data.

[0145] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0146] An embodiment of the present application also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned embodiment of the method for identifying wild animal and plant species are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0147] For example, Figure 4 FIG. 1 shows a schematic diagram of the physical structure of an electronic device. Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. The processor 410, the communications interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call logic instructions stored in the memory 430. The processor 410 is used to execute the various processes of the method for identifying wild animal and plant species according to the embodiment of the present application, which will not be elaborated on here.

[0148] In addition, the logic instructions in the aforementioned memory 430 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0149] The present application also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the various processes of the aforementioned embodiment of the method for identifying wild animal and plant species, achieving the same technical effects. To avoid repetition, the description is omitted here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0150] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0151] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of this application.

[0152] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A method for identifying wild animal and plant species, characterized in that: include: Preprocessing the collected biological data and environmental data associated with the target species to obtain data to be analyzed corresponding to the target species, the biological data including image data, audio data, and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environment data associated with the living environment of the target species, where the target species is a wild animal or a wild plant; Extracting features from the data to be analyzed and aligning the extracted multimodal features based on time, the multimodal features including biometric features and environmental features, the biometric features including image features extracted from the image data, audio features extracted from the audio data, and positioning features extracted from the positioning data; Based on the dynamically adjusted weights of each modal feature, the aligned multimodal features are subjected to feature fusion to generate a comprehensive feature vector indicating comprehensive feature information of the target species, wherein the weights of each modal feature are dynamically adjusted based on the real-time contribution of the multimodal feature to species identification; Based on the target large model, the comprehensive feature vector is subjected to inference analysis to obtain at least one species name corresponding to the target species and the confidence of each species name, and the species name of the target species is determined based on the confidence of each species name.

2. The method for identifying wild animal and plant species according to claim 1, characterized in that: The preprocessing of the collected biological data and environmental data associated with the target species to obtain the data to be analyzed corresponding to the target species includes: Performing image denoising, image enhancement, and image standardization on image data associated with the target species to obtain an image to be analyzed; removing background noise from audio data associated with the target species to obtain audio to be analyzed; Converting the latitude and longitude coordinates of the target species into plane coordinates, and combining the data measured by the inertial measurement unit to obtain positioning data to be analyzed; The environmental data associated with the target species are standardized to a uniform range to obtain the environmental data to be analyzed.

3. The method for identifying wild animal and plant species according to claim 1, characterized in that: The extracting features from the data to be analyzed and aligning the extracted multimodal features based on time includes: Extracting image features, audio features, positioning features, and environmental features with timestamps from the data to be analyzed; The extracted features of different modalities are aligned according to the timestamps to obtain multimodal features that remain consistent in the time dimension.

4. The method for identifying wild animal and plant species according to claim 1, characterized in that: The method of fusing the aligned multimodal features based on the dynamically adjusted weights of the modal features to generate a comprehensive feature vector indicating comprehensive feature information of the target species includes: Perform feature splicing on the aligned image features, audio features, positioning features, and environmental features to generate a high-dimensional feature vector; Assigning an initial weight to each modal feature in the multimodal features, fusing the assigned initial weights into the high-dimensional feature vector to generate a baseline feature vector, wherein the initial weights of the environmental features are determined based on the degree of influence of the environmental data on the survival and identification of the target species, and the initial weights of the modal features in the biological features are assigned based on the estimated importance of each modal feature in species identification; The weights of each modal feature in the baseline feature vector are dynamically adjusted based on target parameters, and the weight change amplitude is controlled based on a weight smoothing biological constraint mechanism. The comprehensive feature vector is generated according to the baseline feature vector whose weights are dynamically adjusted. The target parameters include at least one of the following: the full modal weight corresponding to the system time period, the confidence of the target large model for each modal feature, the user's feedback information on the recognition result output by the target large model in the previous recognition cycle, the sample data newly introduced in the current recognition cycle, and the scene in which the current recognition cycle is located.

5. The method for identifying wild animal and plant species according to claim 1, characterized in that: The inference analysis of the comprehensive feature vector based on the target large model is performed to obtain at least one species name corresponding to the target species and the confidence of each species name, and the species name of the target species is determined based on the confidence of each species name, including: Optimizing the target macromodel according to the comprehensive feature vector corresponding to the target species, and performing reasoning analysis on the comprehensive feature vector based on the optimized target macromodel; Obtaining at least one species name corresponding to the target species and the confidence level of each species name based on the inference analysis results output by the optimized target large model; Modifying the confidence level corresponding to each of the at least one species name based on the prior information of the target species; The species name of the target species is determined based on the confidence of the corrected species name.

6. The method for identifying wild animal and plant species according to claim 5, characterized in that: The step of optimizing the target macromodel according to the comprehensive feature vector corresponding to the target species and performing reasoning analysis on the comprehensive feature vector based on the optimized target macromodel includes: fusing the comprehensive feature vector corresponding to the target species with the feature representation of the target large model, and dynamically adjusting the weights of different modal features to optimize the target large model; Performing reasoning analysis on the enhanced comprehensive feature vector based on the optimized target large model, or performing reasoning analysis on the key fusion features based on the optimized target large model; wherein the key fusion features are generated based on the fusion of multi-dimensional key features extracted from the comprehensive feature vector.

7. The method for identifying wild animal and plant species according to any one of claims 1 to 6, characterized in that: The method further comprises: Pre-training an architecture model suitable for multimodal data processing based on first sample data corresponding to the first biological sample set to determine a pre-trained model; Adjusting model parameters of the pre-trained model based on the data-labeled second sample data corresponding to the second biological sample set to fine-tune the pre-trained model; Migrating common features associated with the pre-trained model to the pre-trained model to optimize the pre-trained model; The pre-trained model that has undergone model fine-tuning and model optimization is pruned and quantized to determine the target large model for species identification.

8. The method for identifying wild animal and plant species according to claim 7, characterized in that: Also includes: Generate comprehensive feature vectors and labels associated with the samples based on new sample data provided by the acquisition device, and update model parameters of the target large model using an online learning algorithm based on the comprehensive feature vectors and labels associated with the samples; and / or Obtain annotation information after the user annotates the recognition result output by the target large model, add the annotation information to the sample data to update the sample data, and adjust the model parameters of the target large model based on the updated sample data.

9. A wild animal and plant species identification system, characterized in that: include: Collection equipment and cloud platform; The collection device collects biological data and environmental data associated with the target species, the biological data including image data, audio data, and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environment data associated with the living environment of the target species, wherein the target species is a wild animal or a wild plant; The cloud platform preprocesses the data provided by the acquisition device, obtains the data to be analyzed corresponding to the target species, extracts features from the data to be analyzed, and aligns the extracted multimodal features based on time, wherein the multimodal features include biological features and environmental features, and the biological features include image features extracted from the image data, audio features extracted from the audio data, and positioning features extracted from the positioning data; The cloud platform performs feature fusion on the aligned multimodal features based on the dynamically adjusted weights of each modal feature to generate a comprehensive feature vector indicating comprehensive feature information of the target species. The weights of each modal feature are dynamically adjusted based on the real-time contribution of the multimodal feature to species identification. The cloud platform performs reasoning analysis on the comprehensive feature vector based on the deployed target large model, obtains at least one species name corresponding to the target species and the confidence of each species name, and determines the species name of the target species based on the confidence of each species name.

10. The wild animal and plant species identification system according to claim 9, characterized in that: When generating the comprehensive feature vector, the cloud platform is further used to: Perform feature splicing on the aligned image features, audio features, positioning features, and environmental features to generate a high-dimensional feature vector; Assigning an initial weight to each modal feature in the multimodal features, fusing the assigned initial weights into the high-dimensional feature vector to generate a baseline feature vector, wherein the initial weights of the environmental features are determined based on the degree of influence of the environmental data on the survival and identification of the target species, and the initial weights of the modal features in the biological features are assigned based on the estimated importance of each modal feature in species identification; The weights of each modal feature in the baseline feature vector are dynamically adjusted based on target parameters, and the weight change amplitude is controlled based on a weight smoothing biological constraint mechanism. The comprehensive feature vector is generated according to the baseline feature vector whose weights are dynamically adjusted. The target parameters include at least one of the following: the full modal weight corresponding to the system time period, the confidence of the target large model for each modal feature, the user's feedback information on the recognition result output by the target large model in the previous recognition cycle, the sample data newly introduced in the current recognition cycle, and the scene in which the current recognition cycle is located.

Citation Information

Patent Citations

  • Multi-modal data fusion-based species monitoring and identification method and system

    CN118364252A

  • Multi-modal fusion bird identification method and device

    CN118430012A

  • Green seedling type identification method and system

    CN118861762A

  • Bird type identification method and identification device, and electronic equipment

    CN118861987A

  • Wood tree species identification method based on multi-feature fusion

    CN119229301A

Cited By

  • Digital twinborn body construction method and system of terrestrial ecosystem

    CN120707767A

  • Biological diversity intelligent monitoring method and system based on cloud platform

    CN120726307A

  • Cloud-based intelligent biodiversity monitoring methods and systems

    CN120726307B