A method for identifying a wild animal or plant species
By collecting and preprocessing biological and environmental data, extracting multimodal features, and performing time alignment and dynamic weight adjustment, a comprehensive feature vector is generated. Using a large model for species identification, the problem of decreased identification accuracy caused by abrupt changes in environmental parameters is solved, and more efficient species identification is achieved.
Patent Information
- Application Number
- CN202510968474.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-15
AI Technical Summary
Existing technologies do not consider the temporal dynamics of environmental data in the identification of wild flora and fauna species, which leads to a decrease in identification accuracy when environmental parameters change abruptly, and they fail to effectively integrate multimodal data to improve identification results.
By collecting biological and environmental data associated with the target species, preprocessing them, extracting multimodal features, and fusing them based on time alignment and dynamic weight adjustment, a comprehensive feature vector is generated. A large model is then used for inference analysis to identify the species.
It significantly improves the accuracy and robustness of wild animal and plant species identification, enabling intelligent identification in complex environments.
Smart Images

Figure CN120470544B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of ecology, and in particular to a method for identifying wild animal and plant species. BACKGROUND
[0002] With the increasing awareness of global ecological protection and biodiversity protection, the demand for monitoring and research of wild animals and plants is growing.
[0003] Currently, when monitoring and researching wild animals and plants, staff need to use artificial intelligence to collect pictures, videos, audio and other data related to wild animals and plants for intelligent analysis to study species. In this process, the processing of environmental data related to the survival of species is relatively simple. Existing technologies (such as pure visual classification models) usually do not consider the time dynamics of environmental data, and cannot identify species in a specific environment due to the failure to consider the influence of environmental factors on species identification. In addition, the existing technology usually performs static weight fusion on the collected species images, audio and the like, thereby affecting the accuracy of species identification under conditions such as sudden mutation of environmental parameters leading to stepwise decline in quality of a specific modality (such as sudden drop in audio signal-to-noise ratio during heavy rain). SUMMARY
[0004] In view of the above problems, the embodiments of the present application provide a method for identifying wild animal and plant species to overcome the above problems or at least partially solve the above problems.
[0005] In a first aspect, the embodiments of the present application provide a method for identifying wild animal and plant species, comprising:
[0006] Preprocessing the collected biological data and environmental data associated with the target species to obtain the analysis data corresponding to the target species, the biological data including image data, audio data and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environment data associated with the survival environment of the target species, the target species being a wild animal or a wild plant;
[0007] Extracting features from the analysis data and aligning the extracted multi-modal features based on time, the multi-modal features including biological features and environmental features, the biological features including image features extracted from the image data, audio features extracted from the audio data and positioning features extracted from the positioning data;
[0008] Performing feature fusion on the aligned multi-modal features based on the dynamically adjusted weights of the modal features to generate a comprehensive feature vector indicating the comprehensive feature information of the target species, the weights of the modal features being dynamically adjusted based on the real-time contribution of the multi-modal features to species identification;
[0009] Inference analysis is performed on the comprehensive feature vector based on a target large model to obtain at least one species name corresponding to the target species and a confidence degree of each species name, and the species name of the target species is determined based on the confidence degrees of the species names.
[0010] In a second aspect, an embodiment of the present application provides a wild animal and plant species identification system, comprising: a collection device and a cloud platform.
[0011] The collection device collects biological data and environmental data associated with a target species, the biological data including image data, audio data and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environment data associated with the living environment of the target species, the target species being a wild animal or a wild plant.
[0012] The cloud platform pre-processes the data provided by the collection device to obtain to-be-analyzed data corresponding to the target species, extracts features from the to-be-analyzed data, and aligns the extracted multi-modal features based on time, the multi-modal features including biological features and environmental features, the biological features including image features extracted from the image data, audio features extracted from the audio data, and positioning features extracted from the positioning data.
[0013] The cloud platform performs feature fusion on the aligned multi-modal features based on the dynamically adjusted weights of the modal features to generate a comprehensive feature vector indicating comprehensive feature information of the target species, and the weights of the modal features are dynamically adjusted based on the real-time contribution of the multi-modal features to species identification.
[0014] Inference analysis is performed on the comprehensive feature vector based on a target large model to obtain at least one species name corresponding to the target species and a confidence degree of each species name, and the species name of the target species is determined based on the confidence degrees of the species names.
[0015] The technical scheme of the embodiment of the application collects biological data and environmental data associated with the target species, and after obtaining the to-be-analyzed data corresponding to the target species based on data preprocessing, performs feature extraction on the to-be-analyzed data, aligns the extracted multi-modal features based on time, and performs feature fusion on the aligned multi-modal features based on the weights of the dynamically adjusted modal features, to generate a comprehensive feature vector. Based on the target large model, the species name of the target species and the confidence of each species name are inferred and analyzed, the large model technology can be used to realize intelligent identification of wild animals and plants, and by introducing the environmental feature processing and dynamic weight adjustment mechanism, a more comprehensive multi-modal data fusion scheme is provided to optimize the generation of the comprehensive feature vector, and then the accuracy and robustness of the wild animal and plant identification can be significantly improved by analyzing the optimized comprehensive feature vector. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 FIG. 1 shows a schematic diagram of a wild animal and plant species identification method provided by an embodiment of the application;
[0017] Figure 2 FIG. 4 shows an example diagram of generating a comprehensive feature vector by dynamically adjusting the weights of the modal features in the reference feature vector provided by an embodiment of the application;
[0018] Figure 3 FIG. 5 shows a schematic diagram of a wild animal and plant species identification system provided by an embodiment of the application;
[0019] Figure 4 FIG. 6 shows a schematic diagram of an electronic device structure provided by an embodiment of the application. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0021] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. The plurality of embodiments in the embodiments of the application can include two and more than two.
[0022] In various embodiments of the present application, it should be understood that the size of the serial number of the following processes does not mean the order of execution, and the execution order of the processes should be determined by their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0023] The embodiments of the present application provide a wild plant and animal species identification method, as shown in the method applied to a cloud platform, comprising the following steps: Figure 1
[0024] Step 101, pre-process the collected biological data and environmental data associated with the target species to obtain the target species corresponding to the data to be analyzed, the biological data includes image data, audio data and positioning data associated with the target species, and the environmental data at least includes meteorological data and physical environment data associated with the living environment of the target species, and the target species is a wild animal or a wild plant.
[0025] After collecting the biological data and environmental data associated with the target species by using the collection device, the collected data is pre-processed to obtain the target species corresponding to the data to be analyzed. The collected biological data at least includes image data, audio data and positioning data, and the collected environmental data at least includes meteorological data and physical environment data. Specifically, the environmental data includes but is not limited to temperature, humidity, soil conditions, light intensity, water resource conditions, air pressure, wind speed, season and other data related to the living environment of the species. By collecting the environmental data associated with the target species, the background information of the species survival and activity can be provided to more comprehensively describe the species characteristics, thereby helping to improve the accuracy of species identification.
[0026] In a specific embodiment, an unmanned device (such as a drone, an unmanned ship) equipped with various sensors is used for data collection, and various sensors on the unmanned device are controlled to collect data at a set frequency, and attributes such as time stamp and positioning information are added, realizing unified control of the sensors and synchronous data collection, and then obtaining multi-dimensional information related to the species. The unmanned device can travel according to a preset path or autonomously plan a path, and obtain image data, audio data, positioning data and environmental data during the travel. For the case that the unmanned device travels according to the preset path, a GIS (Geographic Information System) tool or path planning software is used to pre-plan the travel path of the unmanned device (such as a series of latitude and longitude coordinate points), and the navigation system of the unmanned device automatically adjusts the direction and speed according to these coordinate points. For the case of autonomous path planning, a path planning algorithm is used to dynamically generate a travel path in combination with real-time environmental data (such as obstacle position, water flow speed, etc.), and the navigation system of the unmanned device automatically adjusts the direction and speed according to the real-time positioning information and the path planning result.
[0027] After collecting multi-dimensional data associated with the target species, data preprocessing is performed on the collected data, such as denoising, enhancement, format conversion, etc. The data preprocessing can provide relatively rich data while improving data quality, laying a foundation for subsequent data analysis.
[0028] In step 102, the data to be analyzed is subjected to feature extraction, and the extracted multi-modal features are time-aligned, including biological features and environmental features. The biological features include image features extracted from image data, audio features extracted from audio data, and positioning features extracted from positioning data.
[0029] After obtaining the data to be analyzed of the target species through data preprocessing, multi-modal features are extracted from the data to be analyzed. The extracted multi-modal features include biological features and environmental features. The biological features include image features, audio features, and positioning features. For example, wild plant and animal features such as texture, shape, color, etc. are extracted from images using a convolutional neural network; features such as mel-frequency spectrum and mel-frequency cepstral coefficient (MFCC) are extracted from audio, which can reflect the frequency and time characteristics of the audio; and two-dimensional coordinates of positioning data are extracted as features. The environmental features include, for example, climate features, soil features, hydrological features (water quality, water depth, water temperature, flow rate, etc.), topographic features, vegetation features, etc.
[0030] After obtaining the multi-modal features based on feature extraction, the extracted multi-modal features are time-aligned to ensure consistency of multi-modal data in the time dimension through feature alignment. For example, image, audio, positioning, and other asynchronous biological data are aligned with dynamic environmental data such as weather and physical environment based on the time axis, solving the problem of inconsistent timing of multi-source data in natural scenes.
[0031] In step 103, the aligned multi-modal features are subjected to feature fusion based on the dynamically adjusted weights of the modal features, to generate a comprehensive feature vector indicating the comprehensive feature information of the target species. The weights of the modal features are dynamically adjusted based on the real-time contribution of the multi-modal features to species identification.
[0032] After aligning the multi-modal features, the weights of the modal features are dynamically adjusted. Based on the dynamically adjusted weights and the aligned multi-modal features, feature fusion is performed to generate a comprehensive feature vector indicating the comprehensive feature information of the target species.
[0033] The dynamic adjustment of the weight is based on the contribution of the biological features and the environmental features to the species recognition, that is, the weight of each modal feature is dynamically adjusted based on the real-time contribution of the feature to the species recognition. The essence of dynamically adjusting the weight is to quantitatively determine the reliability of each modal feature in real time and to realize adaptive fusion of the features through weight distribution; by dynamically adjusting the weight, the generation of the comprehensive feature vector can be optimized to improve the accuracy and robustness of species recognition using the comprehensive feature vector. Specifically, for example, the fusion weight is dynamically adjusted according to the real-time contribution of each modal feature to the species recognition (such as high image weight in sunny weather and high audio weight at night), thereby improving the recognition robustness in complex environments.
[0034] The process of fusing features of different modalities to generate a comprehensive feature vector is as follows: image features, audio features, positioning features, and environmental features are combined with corresponding weights to splice features to obtain high-dimensional features. This process can provide a more comprehensive multi-modal data fusion scheme by introducing environmental features and dynamically adjusting weights. This way of generating a comprehensive feature vector through multi-modal data fusion can fully utilize the advantages of different modal data and improve the accuracy and robustness of species recognition.
[0035] In the prior art, traditional biological image features and audio features are usually used for feature fusion, and such fusion cannot achieve species recognition in specific environments. In this scheme, traditional biological image features (such as bird feather texture) and audio features (such as bird song spectrum) are associated with environmental features (such as temperature and humidity, altitude) to model and narrow the range of species recognition through environmental constraints (for example, only possible species under a specific humidity).
[0036] Step 104: performing inference analysis on the comprehensive feature vector based on the target large model to obtain at least one species name corresponding to the target species and the confidence of each species name, and determining the species name of the target species based on the confidence of each species name.
[0037] After determining the comprehensive feature vector through feature fusion, inference analysis is performed on the comprehensive feature vector based on a target large model for species recognition to obtain at least one species name corresponding to the target species and the confidence of each species name output by the target large model. Based on the at least one species name and the corresponding confidence, the final species name of the target species is determined to utilize the large model for wild plant and animal species recognition.
[0038] The confidence of the species name represents the confidence level of the model in predicting the species name, which is usually represented by a probability value. Based on the confidence of each species name output by the target large model for the target species, the target species can be effectively identified.
[0039] In the above embodiments of the present application, biological data and environmental data associated with the target species are collected, and after obtaining the corresponding analysis data of the target species based on data preprocessing, feature extraction is performed on the analysis data, the extracted multi-modal features are time-aligned, the aligned multi-modal features are fused based on the dynamically adjusted weights of the modal features, a comprehensive feature vector is generated, the target species name and the confidence of each species name are predicted based on the inference analysis of the target large model on the comprehensive feature vector, and the intelligent recognition of wild animals and plants can be realized by using the large model technology. Moreover, by introducing the environmental feature processing and dynamic weight adjustment mechanism, a more comprehensive multi-modal data fusion scheme is provided, the generation of the comprehensive feature vector is optimized, and then the accuracy and robustness of the wild animal and plant species recognition can be significantly improved by analyzing the optimized comprehensive feature vector.
[0040] The process of obtaining analysis data based on data preprocessing is introduced as follows. When the biological data and environmental data associated with the target species are preprocessed, the corresponding analysis data of the target species is obtained, including:
[0041] The image data associated with the target species is subjected to image denoising, image enhancement and image standardization processing to obtain the analysis image.
[0042] The background noise in the audio data associated with the target species is removed to obtain the analysis audio.
[0043] The latitude and longitude coordinates of the target species are converted into plane coordinates, and the analysis positioning data is obtained in combination with the data measured by the inertial measurement unit.
[0044] The environmental data associated with the target species is standardized to a unified range to obtain the analysis environmental data.
[0045] When the image data is preprocessed, image denoising algorithms (such as Gaussian filtering and median filtering) are used to remove noise in the image, and image enhancement and standardization processing are performed. In image enhancement, histogram equalization, contrast adjustment and other methods are used to enhance the details of the image; in image standardization, the image data is standardized to a unified size and format. By performing image preprocessing, the collected images are denoised, enhanced and standardized, which can improve the image quality.
[0046] When the audio data is preprocessed, audio noise reduction algorithms (such as spectral subtraction and wavelet transform) are used to remove background noise. Normalization and pre-emphasis can also be used to further process the audio. The purpose of pre-emphasis is to enhance the high-frequency components in the audio signal and highlight important features; normalization is to scale the amplitude of the audio signal to a specific range to improve numerical stability. Through the above audio processing, the quality of the audio data can be gradually optimized.
[0047] In the preprocessing of positioning data, the latitude and longitude coordinates are converted into two-dimensional plane coordinates, combined with IMU (Inertial Measurement Unit) data, and more accurate positioning information is obtained through a weighted fusion algorithm. In the preprocessing of environmental data, the environmental data is standardized to a unified range.
[0048] Among them, the preprocessed image data can be stored in JPEG or PNG format, and the timestamp and positioning information are added as the file name or metadata; the preprocessed audio data is stored in WAV or MP3 format, and the timestamp and positioning information are added as the file name or metadata; the preprocessed positioning data and environmental data are stored in CSV or JSON format, containing timestamp, latitude and longitude, environmental parameters and other information.
[0049] Optionally, after obtaining the to-be-analyzed data corresponding to the target species based on data preprocessing, feature extraction is performed on the to-be-analyzed data, and the extracted multi-modal features are time-aligned.
[0050] Image features, audio features, positioning features, and environmental features carrying timestamps are extracted from the to-be-analyzed data; the extracted features of different modalities are aligned according to the timestamps, and multi-modal features that remain consistent in the time dimension are obtained.
[0051] In the feature extraction in the to-be-analyzed data, for example, a convolutional neural network is used to extract image features, mel spectrum, MFCC, etc. of audio features, positioning features of positioning data, and environmental features of environmental data. Among them, the environmental features are obtained by converting the environmental data into a feature vector that can reflect the state of the survival environment of the target species, such as temperature, humidity, etc. After standardization, the environmental data is used as the environmental feature.
[0052] After feature extraction, image features, audio features, positioning features, and environmental features with accurate timestamps are obtained, and feature alignment is performed based on the timestamps, so that data at the same time point can be corresponded, ensuring the consistency of the data in the time dimension.
[0053] Through the above data preprocessing process, the data quality can be improved, providing a better data basis and stronger data support for subsequent wild animal and plant species identification; by extracting features from the preprocessed data and time-aligning the extracted features, multi-modal features that remain consistent in the time dimension can be obtained.
[0054] In an optional embodiment of the present application, when the aligned multi-modal features are fused based on the dynamically adjusted weights of the features of each modality to generate a comprehensive feature vector indicating the comprehensive feature information of the target species, it includes:
[0055] The image features, audio features, positioning features, and environmental features after alignment processing are spliced to generate a high-dimensional feature vector.
[0056] An initial weight is assigned to each modality feature in the multi-modal feature, and the assigned initial weight is fused into the high-dimensional feature vector to generate a reference feature vector. The initial weight of the environmental feature is determined based on the influence of the environmental data on the survival and identification of the target species, and the initial weight of each modality feature in the biological feature is assigned based on the estimated importance of each modality feature in species identification.
[0057] The weights of each modality feature in the reference feature vector are dynamically adjusted based on the target parameters, and the weight variation amplitude is controlled based on the weight smoothing biological constraint mechanism. A comprehensive feature vector is generated based on the reference feature vector whose weight has been dynamically adjusted. The target parameters include at least one of the following: the full-modal weight corresponding to the system time period, the confidence of the target large model for each modality feature, the feedback information of the user on the identification result output by the target large model in the last identification period, the sample data newly introduced in the current identification period, and the scene in which the current identification period is located.
[0058] After the extracted features are aligned based on the timestamp, the image features, audio features, positioning features, and environmental features after alignment processing are spliced to generate a high-dimensional feature vector, for example, Fhigh-dimensional feature correlation = [Fimage, Faudio, Fpositioning, Fenvironment], where Fimage represents image features, Faudio represents audio features, Fpositioning represents positioning features, and Fenvironment represents environmental features.
[0059] According to the importance and correlation of different modality data, an initial weight is assigned to each modality, and the assigned initial weight is fused into the high-dimensional feature vector to generate a reference feature vector. The initial weight of the environmental feature is determined based on the influence of the environmental data on the survival and identification of the target species, and the initial weight of each modality feature in the biological feature is assigned based on the estimated importance of each modality feature in species identification. For example, Freference feature vector = αFimage + βFaudio + γFpositioning + δFenvironment, where α, β, γ, and δ are initial weight coefficients, and α + β + γ + δ = 1.
[0060] In the case of generating a reference feature vector, a comprehensive feature vector can be generated by dynamically adjusting the weight. The process of dynamically adjusting the weight is based on the contribution of biological features and environmental features to species identification. When dynamically adjusting the weight, the weight of each modality feature can be adjusted based on the attention mechanism.
[0061] As an example, through the multi-modal model of the Transformer architecture, the correlation between different modal data is learned, and the weights are adaptively adjusted. For example, using a multi-head attention mechanism to weight-sum the features of different modalities: Fintegrated feature vector = Attention(Fimage, Faudio, Flocation, Fenvironment). The generated integrated feature vector contains the integrated information of image, audio, location and environment data, which can more comprehensively describe the species characteristics; and using the attention mechanism, according to the importance and correlation of different modal data, the weights of each modal feature are dynamically adjusted, and the integrated feature vector is generated by weighted fusion, which can provide a high-reliability integrated feature vector, facilitating subsequent species analysis.
[0062] The attention mechanism is a method of dynamically adjusting weights by calculating the similarity or correlation between input features and adaptively assigning weights. In the case of generating a reference feature vector, the embodiments of the present application can also dynamically adjust the weights of each modal feature in the reference feature vector based on the target parameters, and control the weight change range based on the weight smoothing biological constraint mechanism. After adjusting the weights of each modal feature in the reference feature vector, the integrated feature vector is determined. The target parameters include at least one of the following: the full-modal weight corresponding to the system time period, the confidence of the target large model for each modal feature, the feedback information of the user on the recognition result output by the target large model in the last recognition period, the sample data newly introduced in the current recognition period, and the scene in which the current recognition period is located. And the way of adjusting the weight based on the target parameter can be combined with the attention mechanism to achieve a more flexible and effective weight adjustment strategy.
[0063] In the implementation of the present scheme, a comprehensive feature vector generation method compatible with large models such as Transformer is specifically designed to uniformly encode unstructured biological data and structured environmental data, solving the problem of adapting large models to multi-modal heterogeneous data input.
[0064] As an optional implementation, the process of generating an integrated feature vector by dynamically adjusting the weights of each modal feature in the reference feature vector based on the target parameters is as shown in Figure 2
[0065] After obtaining the reference feature vector by fusing the high-dimensional feature vector generated based on the initial weights and splicing, the weights of each modal feature are dynamically adjusted based on the confidence of the target large model for each modal feature, the feedback information of the user on the recognition result, the sample data newly introduced in the current recognition period, and the scene in which the current recognition period is located. After completing the dynamic adjustment of the weights, the integrated feature vector is generated.
[0066] The process of dynamically adjusting the weight of each modality feature based on target parameters is introduced below.
[0067] When dynamically adjusting the weight of each modality feature in the reference feature vector, the weight of each modality feature is dynamically adjusted based on the overall modality weight corresponding to the current system time period, so as to adapt to the time and periodically adjust the weight. The time-adaptive periodic weight adjustment needs to determine the frequency of weight adjustment according to different environmental conditions or specific rules (such as day-night change, seasonal change, etc.). In this way, the essence is to trigger weight update based on the fixed time interval of the system clock, for example, the basic period is 30 seconds, and the time window can be configured to update the weight once every 30 seconds. When the weight is updated, the weight of all modalities is updated within each time window.
[0068] When the weight of each modality feature is dynamically adjusted based on the confidence of each modality feature of the target large model, the confidence of each modality feature of the target large model is calculated in each recognition period. When the confidence of a certain modality feature is lower than the first threshold, the weight of the modality feature is down-weighted, and the weights of the remaining modality features are compensated and enhanced. The confidence is calculated based on the recognition accuracy, error rate and other indicators of the target large model for each modality feature. For environmental features, the confidence is also related to the stability and correlation of environmental data; when down-weighting, an exponential decay method is used, and the decay coefficient is negatively correlated with the confidence score.
[0069] When the weight of each modality feature is dynamically adjusted based on the feedback information of the user on the recognition result, the user feedback including the user's accuracy evaluation on the recognition result and the subjective judgment of the importance of different modality features is obtained, and the weight of each modality feature is adjusted based on the user feedback. This way makes the large model better adapt to the user's expectations and the needs of the actual application scenario. When adjusting the weight of each modality feature according to the user feedback, the following strategies can be used: 1. Weight adjustment based on accuracy evaluation; if the user's accuracy evaluation of the recognition result is low, the weight of the large model for certain modality features can be increased to improve the performance of the large model. 2. Weight adjustment based on modality feature importance judgment; adjust the weight of each modality feature according to the user's subjective judgment of the importance of different modality features. 3. Comprehensive weight adjustment; combine the weight adjustment based on accuracy and importance to comprehensively adjust the weight of each modality feature.
[0070] In the dynamic adjustment of weights based on newly introduced sample data in the current identification cycle, the newly introduced sample data in the current identification cycle is obtained, such as the newly introduced sample data caused by autumn migration, and the newly introduced sample data includes new biological data and environmental data. Then, the weight adjustment is performed based on the following operations: 1. Evaluate the impact of new data; calculate the impact of new data on model identification, for example, by calculating the confidence or prediction error of new data. 2. Determine the weight adjustment direction; according to the impact of new data, determine the weight adjustment direction and amplitude of each modal feature. Update the weight; according to the determined adjustment direction and amplitude, update the weight of each modal feature.
[0071] In the dynamic adjustment of weights based on the scene in which the current identification cycle is located, the scene needs to be identified. For example, the scene includes different target species living environments such as wild, laboratory, nature reserve, etc. For different scenes, the importance of biological features and environmental features is different, for example, in the wild environment, the contribution of environmental features to species identification is greater. The scene can also include, for example, environmental parameter mutation, specific events (such as seasonal migration), time variation, etc. Environmental parameter mutation will cause a stepwise decrease in the quality of a specific modality (such as a sudden drop in audio signal-to-noise ratio during heavy rain), which requires a dynamic weight adjustment mechanism to solve. Specific events (such as seasonal migration) will trigger multi-modal feature weight update. The environment and biological features during migration will change significantly, and these changes will affect the performance of the model. By dynamically adjusting the weights of each modal feature, the model can better adapt to the data characteristics during migration. Time variation (such as the change from day to night) will trigger multi-modal feature weight update. The switching between day and night will cause changes in light, temperature, behavior patterns, etc. These changes need to be adapted by adjusting the weight.
[0072] The process of dynamically adjusting the weight is introduced below through several specific examples.
[0073] I. Image feature weight adjustment driven by light intensity, target species is bird, daytime relies on image features, night relies on audio features.
[0074] 1. Input data; environmental data (light intensity L), image features (extracted image feature vector V img), audio features (audio feature vector V audio).
[0075] 2. Confidence calculation; calculate the image confidence score B img according to the light intensity L, the relationship between the image confidence score and the light intensity is an S-shaped curve, and the specific formula is as follows:
[0076]
[0077] wherein,L 0 is the illumination threshold, k is the slope parameter (used to control the transition speed). L 0 is the critical value of the species' visual ability significantly decreased when the illumination intensity is lower than this threshold, determined by experiments; k According to the diurnal activity pattern of the species, nocturnal animals k The value is smaller (smooth transition).
[0078] 3. Weight distribution; according to the image confidence score The weights of image features and audio features are assigned, and the image feature weight = , and the audio feature weight =1- .
[0079] By dynamically adjusting the weights of image features and audio features according to the illumination intensity, the model can better adapt to different conditions in the daytime and at night. In this way, the model can more accurately identify and analyze the behavior and characteristics of birds.
[0080] II. Environmental noise-driven audio weight adjustment, strong wind or rain causing audio signal-to-noise ratio to decrease.
[0081] 1. Input data; environmental data (noise decibel value N), audio features (Mel frequency spectrum features V audio), positioning data (positioning feature vector V location).
[0082] 2. Signal-to-noise ratio evaluation; define audio quality coefficient Q audio, inversely proportional to noise decibel value N, the specific formula is as follows:
[0083]
[0084] Where, N : the current environmental noise decibel value (measured value). N max: the maximum noise threshold allowed by the system (audio is completely disabled beyond this value), which is specifically measured by experimentally measuring the audio recognition threshold of the target species. For example: bird song recognition (beyond this value, the song is masked by environmental noise), underwater sonar (dependent on the sound absorption characteristics of water).
[0085] ε is the attenuation factor (a nonlinear parameter that controls the speed of quality decline, ε ≥ 1). Specifically, the audio recognition rate data of the target species in a noisy environment can be obtained; through nonlinear regression fitting, the optimal ε is obtained, so that Qaudio and recognition rate Pearson correlation coefficient maximization; while different species can be set different ε (such as ε = 2.0 for birds, ε = 1.5 for frogs).
[0086] : Ensure the result is non-negative (when N ≥ N max, Q audio=0).
[0087] 3. Weight correction; according to the audio quality coefficient Q audio corrects the audio weight W audio, if Q audio<0.5, then forced to reduce the audio weight to W audio=0.2, and increase the positioning feature weight.
[0088] By dynamically adjusting the weight of audio features according to environmental noise, the model can better adapt to harsh conditions such as strong wind or rain. In this way, the model can more accurately identify and analyze the behavior and characteristics of target species.
[0089] Three, dynamic adjustment of positioning feature weight
[0090] 1. Scene demand
[0091] Positioning data (such as GPS / Beidou) has significant changes in reliability in the following scenarios:
[0092] Time-sensitive: The positioning weight of nocturnal migratory birds is higher than that during the day.
[0093] Path deviation: When deviating from the historical migration route, the weight needs to be reduced.
[0094] Signal quality: Urban canyons or cloudy weather cause positioning error to increase.
[0095] 2. Weight calculation model input data
[0096] Environmental parameters: time (day and night), positioning error radius (meters), historical path deviation degree (%).
[0097] Biological characteristics: typical activity radius of species (such as 50km for migratory birds, 5km for resident birds).
[0098] 3. Weight calculation formula
[0099]
[0100] Where, through migration data analysis, it is found that the night time weight needs to be increased by 20% to compensate for the positioning dependence, considering setting the coefficient 1 to 1.2, the path weight and the positioning weight are linearly positively correlated, considering setting the coefficient 2 to 1, the error penalty is reduced by 0.2 (upper limit 0.5) for every 100 meters, considering setting the coefficient 3 to 0.2; but in this case, the regression fitting coefficient is unconstrained, which will lead to: W gps The numerical range of does not match the image, audio, and other modal weight.
[0101] In order to prevent W the gps value from being too large, ensure that the order of magnitude of each feature is consistent when multi-modal fusion, limit the total influence of time and path weight, and avoid excessive dependence on positioning data, you can set the adjustment constraint condition: coefficient 1 + coefficient 2 = 1.2, so you can keep the relative proportion of the regression relationship, coefficient 3 = 0.2;
[0102]
[0103] PathWeight = 1-min(1, )
[0104] ErrorPenalty = min(0.5, ).
[0105] 4. Calculation example scenario
[0106] The time is night 20:00, the deviation from the historical path is 15 km, the positioning error is 80 m, and the species activity radius is 50 km; TimeWeight = 0.7 (night), PathWeight = 1-min(1, 15 / 50) = 0.7, ErrorPenalty = min(0.5, 80 / 100) = 0.5. Then the positioning weight .
[0107] In the above example, the time weight is adjusted according to the time (day or night), the path weight is adjusted according to the distance from the historical path, the error penalty is adjusted according to the positioning error, and the positioning weight is calculated by combining the time weight, path weight and error penalty. By dynamically adjusting the weight of positioning data according to environmental parameters and biological characteristics, the model can better adapt to the reliability changes of positioning data in different scenarios.
[0108] Four, weight optimization of multi-modal dynamic fusion
[0109] 1. Fusion rules
[0110] Input: image weight of instance one, audio weight of instance two, positioning weight of instance three.
[0111] Dynamic compensation: if positioning error is greater than 200 meters, forcibly increase positioning weight W gps Limit, such as limiting within 0.3, when both image and audio weights are less than a set value (such as 0.2), increase positioning weight by, for example, 50%.
[0112] 2. Normalization formula
[0113] The normalization formula is used to adjust the weights so that their sum is 1, in order to perform fusion calculation:
[0114]
[0115] W i The original weight of the i-th modality is calculated from environmental data, C i is the compensation coefficient (weight amplification factor in emergency mode), C i The value of depends on the mode, emergency compensation mode C i = 1.5, normal mode C i = 1; is the total amount of weighted effective contribution of all modes, W i′ is the normalized maximum weight.
[0116] As an example, input: image weight = 0.15 (heavy fog), audio weight = 0.1 (heavy rain), positioning weight = 0.8 (error = 30 meters); since both image and audio weights are less than 0.2, increase positioning weight by 50% to update to 0.8*1.5 = 1.2.
[0117] In the above scenario, the new positioning weight calculated is .
[0118] In this instance, under extreme weather conditions, it can automatically switch to positioning dominant mode, increasing the weight of positioning to adapt to different working modes.
[0119] The above four instances introduce dynamic adjustment of feature weights under different conditions. Based on the dynamic weight adjustment mechanism, a more comprehensive multi-modal data fusion scheme is provided, optimizing the generation of comprehensive feature vectors.
[0120] It should be noted that in the process of adjusting the weight based on the target parameter, the weight change range can be controlled based on the weight smoothing biological constraint mechanism. The weight smoothing biological constraint refers to intelligently controlling the rate and range of weight change by introducing species-specific behavioral rules and physiological characteristics, avoiding non-physiological mutations caused by pure mathematical optimization. The weight smoothing process includes, for example, the following biological constraint conditions: 1. Load a preset smoothing coefficient according to the target species type; 2. The weight change range between adjacent time frames is constrained to be no more than the maximum adaptation value of the species; 3. An enhanced smoothing mode is enabled during the breeding period / migration period and other special physiological stages. It should be noted that when the weight change triggers a physiological limit alarm, it can automatically switch to a conservative adjustment mode.
[0121] The rationality of the weight smoothing biological constraint is illustrated below by comparing it with pure mathematical smoothing. The two are compared in the following dimensions:
[0122] In the target function dimension: Traditional exponential smoothing aims to minimize mean square error, focusing more on mathematical fitting accuracy, suitable for general data smoothing needs; biological constraint smoothing aims to conform to species behavior patterns, and when processing data, the physiological and behavioral characteristics of the species are considered to ensure that the smoothed data can truly reflect the actual behavior of the species.
[0123] In the parameter source dimension: Traditional exponential smoothing, parameters (such as smoothing coefficients) are mainly obtained through data statistical fitting, usually using historical data for regression analysis or other statistical methods to determine the optimal smoothing coefficient; biological constraint smoothing, parameters are derived from ecological research consensus and field observation data, including not only laboratory research results but also long-term observation data of species in natural environments to ensure the biological rationality of the parameters.
[0124] In the dynamic adaptability dimension: Traditional exponential smoothing, the smoothing coefficient is fixed and remains unchanged throughout the data processing process, suitable for scenarios where data changes are relatively stable; biological constraint smoothing, the smoothing coefficient is dynamically adjusted according to the physiological state of the species, for example, when the species is in different physiological stages (such as breeding period, migration period, etc.), the smoothing coefficient will change accordingly to better reflect the actual behavior changes of the species.
[0125] In the abnormality handling dimension, traditional exponential smoothing may lead to results that violate biological laws when handling outliers, for example, data mutations may be smoothed, but such processing may not conform to the actual physiological characteristics of the species; biological constraint smoothing will be forced to comply with physiological limits, i.e., when processing data, the physiological limitations of the species will be considered to avoid results that do not conform to biological common sense.
[0126] For example, in traditional processing, the weight of a certain bird's audio rapidly decreases from 0.5 to 0.2 within 0.3 seconds, causing the interruption of song analysis. This rapid change may exceed the actual physiological capabilities of the species, leading to anomalies in data processing and preventing effective analysis from continuing. The biological constraint smoothing mechanism limits the maximum decrease to 0.1 / second and smooths the transition for 1.5 seconds, maintaining continuous observation. By limiting the maximum rate of data change, the data change is ensured to conform to the physiological characteristics of the species, thereby avoiding analysis interruptions caused by data mutations and maintaining the continuity and reliability of data processing.
[0127] To further illustrate the weight smoothing biological constraint, the following gives a specific example of weight smoothing biological constraint:
[0128] I. Weight adjustment of raptors (eagles) in dense clouds
[0129] Environmental conditions: Dense clouds, reduced visibility; due to reduced light intensity, image quality decreases.
[0130] Biological constraints: The typical reaction time of raptors (eagles) is, for example, 0.1-0.5 seconds, and the smoothing factor range is 0.1-0.3. When suddenly encountering clouds, quickly switch to infrared mode.
[0131] Weight adjustment: Due to reduced visibility, image quality decreases, and image weight needs to be reduced. Assuming the initial image weight is 0.6, according to the biological constraint, the adjustment rate of image weight should be controlled between 0.1-0.3, and the image weight can be gradually reduced to around 0.3. Since switching to infrared mode, infrared weight needs to be quickly increased, assuming the initial infrared weight is 0.2, it can be increased to around 0.6 in a short time (such as 0.1 seconds). Audio weight remains unchanged or is adjusted appropriately according to environmental noise.
[0132] II. Weight adjustment of waders (cranes) in light gradual change
[0133] Environmental conditions: Light gradually changes.
[0134] Biological constraints: The typical reaction time of waders (cranes) is, for example, 2-5 seconds, and the smoothing factor range is 0.6-0.8. Slowly reduce visual weight when light gradually changes.
[0135] Weight adjustment: As light changes, image weight needs to be slowly reduced, assuming the initial image weight is 0.7, according to the biological constraint, it can be gradually reduced to around 0.4 within 2-5 seconds. Audio weight can be appropriately increased to compensate for the decrease in visual information, assuming the initial audio weight is 0.2, it can be gradually increased to around 0.4. Positioning weight remains unchanged or is adjusted appropriately as needed.
[0136] III. Weight adjustment of birds (songbirds) when wind noise increases
[0137] Environmental conditions: wind speed increases, environmental noise increases, audio quality decreases.
[0138] Biological constraints: typical reaction time of songbirds (songbirds) is, for example, 1-2 seconds, and the smoothing factor ranges from 0.4 to 0.6. The audio weight is adjusted at medium speed when wind noise increases.
[0139] Weight adjustment: due to the increase in wind noise, the audio weight needs to be reduced. Assuming the initial audio weight is 0.5, it can be reduced to about 0.2 within 1-2 seconds. The image weight can be appropriately increased to compensate for the decrease in audio information. Assuming the initial image weight is 0.3, it can be gradually increased to about 0.6. The positioning weight remains unchanged or is appropriately adjusted as needed.
[0140] The above example introduces how to dynamically adjust the weights of image, audio and positioning data according to the biological characteristics of species and environmental conditions. By following the biological constraints, it can ensure that the weight adjustment process conforms to the physiological and behavioral characteristics of the species.
[0141] The process of species identification based on large models is introduced as follows. When the target large model is based on the comprehensive feature vector for inference analysis, at least one species name corresponding to the target species and the confidence of each species name are obtained, and the species name of the target species is determined based on the confidence of each species name, including:
[0142] According to the comprehensive feature vector corresponding to the target species, the target large model is optimized, and the comprehensive feature vector is analyzed based on the optimized target large model; according to the inference analysis result output by the optimized target large model, at least one species name corresponding to the target species and the confidence of each species name are obtained; the confidence corresponding to each of the at least one species name is corrected based on the prior information of the target species; and the species name of the target species is determined based on the confidence of the corrected species name.
[0143] After obtaining the comprehensive feature vector, the target large model is optimized according to the comprehensive feature vector corresponding to the target species before the comprehensive feature vector is analyzed by the target large model. After optimizing the target large model, the comprehensive feature vector is analyzed by the optimized target large model, and the inference analysis result is output, providing at least one species name corresponding to the target species and the confidence of each species name.
[0144] After obtaining the at least one species name corresponding to the target species and the confidence of each species name, the confidence of the species name is corrected based on the prior information of the target species. The prior information of the target species refers to the information about the characteristics, behavior, distribution, etc. of the species that is known or assumed before a specific observation or experiment is performed. These information usually comes from historical data, literature research, expert knowledge or statistical analysis. Prior information helps better understand and predict the behavior and distribution of the species. After using prior information to correct the confidence of the species name, the species name of the target species is determined based on the confidence of the corrected at least one species name, such as determining the species name with the highest confidence as the final species name.
[0145] In the process of optimizing the target large model according to the comprehensive feature vector corresponding to the target species, and performing inference analysis on the comprehensive feature vector based on the optimized target large model, the following steps are included:
[0146] Fusing the comprehensive feature vector corresponding to the target species with the feature representation of the target large model, and dynamically adjusting the weights of different modal features to optimize the target large model;
[0147] Performing inference analysis on the enhanced comprehensive feature vector based on the optimized target large model, or performing inference analysis on the key fusion features based on the optimized target large model; wherein the key fusion features are generated based on the fusion of multi-dimensional key features extracted from the comprehensive feature vector.
[0148] When optimizing the target large model, it can be optimized from two dimensions: 1. Fuse the comprehensive feature vector corresponding to the target species with the feature representation of the target large model to enhance the understanding ability of the target large model to multi-modal data. The fusion method is, for example: align the input feature format of the comprehensive feature vector with the target large model, map the comprehensive feature vector to the feature space of the target large model through linear transformation or nonlinear mapping, to realize feature fusion. 2. Dynamically adjust the weights of different modal features based on the attention mechanism to optimize the model, so as to improve the recognition accuracy of the model.
[0149] When optimizing the target large model based on feature fusion, the model parameters can be adjusted or not adjusted. For the case of not adjusting the model parameters: 1. The comprehensive feature vector is directly concatenated to the feature representation of the target large model to form a longer feature vector, in this case, the model parameters remain unchanged, only more rich feature input is used in subsequent tasks. 2. The comprehensive feature vector and the feature representation of the target large model are added with certain weights, in this case, the model parameters remain unchanged, only the contribution of the features is adjusted through the weights.
[0150] For the case of adjusting the model parameters: 1. Fine-tuning; after feature fusion, fine-tune the target large model to let the model learn how to better process the fused features. In this case, the model parameters of the target large model will be adjusted according to the new training data. 2. End-to-end training; embed the feature fusion process into the model training process, let the model learn how to fuse features and perform tasks. In this case, the model parameters of the target large model will also be adjusted according to the training target.
[0151] In the attention mechanism, the weights are dynamically calculated, and the calculation of the weights is based on the input features. When dynamically adjusting the weights of different modal features based on the attention mechanism for model optimization, additional learnable weights can also be introduced, which are used to adjust the contribution of different modal features, and these weights are independent of the model parameters. For example, a learnable weight matrix is introduced to adjust the weights of different modal features, and the parameters of the weight matrix are learned independently.
[0152] When optimizing the model based on the attention mechanism, the model parameters can also be adjusted. 1. Fine-tuning; in the attention mechanism, if the model parameters of the target large model are fine-tuned, the model will learn how to better adjust the weights of different modal features. In this case, the model parameters of the target large model will be adjusted according to the new training data. 2. End-to-end training; embed the attention mechanism into the training process of the model, let the model learn how to adjust the weights of different modal features. In this case, the model parameters of the target large model will also be adjusted according to the training target.
[0153] After optimizing the target large model, when processing the comprehensive feature vector using the target large model, the comprehensive feature vector can be first enhanced, such as feature normalization, feature dimension reduction, etc., to improve the robustness of the features, and then the target large model is used to infer the enhanced comprehensive feature vector to output the inference analysis result. It can also be to extract multi-dimensional key features from the comprehensive feature vector, such as image texture features, audio frequency features, positioning coordinate features, and environmental temperature features, fuse the extracted multi-dimensional key features to generate more rich feature representations, that is, generate key fusion features; then based on the optimized target large model, the key fusion features are inferred and analyzed to output the inference analysis result.
[0154] In the above implementation process of species identification based on the target large model, the target large model is optimized to enhance the understanding ability of the target large model to multi-modal data and improve the accuracy of model identification. Based on the optimized target large model, the enhanced comprehensive feature vector or the key fusion features generated based on the comprehensive feature vector are processed, which can relatively accurately identify the species and ensure the efficiency of species identification.
[0155] The process of training the target large model is introduced as follows, including the following steps during model training:
[0156] Pre-training an architecture model suitable for multi-modal data processing based on first sample data corresponding to a first biological sample set;
[0157] Adjusting model parameters of the pre-trained model based on second sample data corresponding to a second biological sample set and subjected to data annotation, to fine-tune the pre-trained model;
[0158] Migrating general features associated with the pre-trained model to the pre-trained model to optimize the pre-trained model;
[0159] Pruning and quantifying the pre-trained model subjected to model fine-tuning and model optimization to determine a target large model for species identification.
[0160] The embodiments of the present application select a Transformer architecture model suitable for multi-modal data processing, such as ViT (Vision Transformer), CLIP (Contrastive Language-Image Pre-training), Audio-CLIP, and Geo-CLIP. ViT is suitable for image data processing and can learn global features in images; CLIP is suitable for multi-modal data processing and can jointly learn feature representations of images and texts, suitable for species identification tasks; Audio-CLIP is based on CLIP and adds processing capabilities for audio modalities, which can jointly learn feature representations of images, texts, and audio, further improving the model's understanding of multi-modal data; Geo-CLIP is based on CLIP and combines positioning data and environmental data to further enhance the model's understanding of geographical and environmental context. Preferably, the selected Transformer architecture model suitable for multi-modal data processing can jointly learn feature representations of images and texts while having understanding capabilities for audio, positioning data, and environmental data.
[0161] In the pre-training stage, a large-scale multi-modal dataset is collected, specifically image dataset, text data related to images such as image descriptions, audio dataset, positioning data (extracting geographic location information of images and audio) and environmental data (environmental information related to image and audio shooting time such as weather, temperature, humidity, etc.). Preprocess the multi-modal data, such as cropping, scaling, normalizing, etc. for image data to ensure consistency in image size and format; perform word segmentation, encoding, etc. on text data to convert text into numerical form that the model can process; perform audio signal processing on audio data, such as extracting mel-spectrogram features, etc. to convert audio into a format suitable for model processing; convert geographic information such as latitude and longitude into a format suitable for model processing, for example, embedding into image or audio features or as independent feature input; standardize or normalize weather, temperature, humidity, etc. environmental data for fusion with image and audio data. After data processing, the pre-trained model is determined based on the pre-training task and the processed data.
[0162] After determining the pre-trained model, model fine-tuning is needed, which requires collecting data related to plants and animals, specifically, for example, collecting data related to wild plants and animals. When collecting data, collect a large amount of image data of wild plants and animals, including samples of different species and different environments, to ensure that the image data covers a variety of scenarios such as different lighting conditions, seasons and geographic locations; collect audio data corresponding to the images, such as animal calls, environmental sounds, etc. to ensure that the audio data is aligned with the image data in time and space; collect relevant text descriptions such as species names, feature descriptions, behavior descriptions, etc. to ensure that the text descriptions are aligned with the image and audio data; collect environmental data related to image and audio data such as weather, temperature, humidity, etc. to ensure that the environmental data is aligned with the image and audio data in time and space, collect positioning data such as latitude and longitude information of the shooting location to ensure that the positioning data is aligned with the image and audio data. After data collection, label the data and fine-tune the model based on the labeled data. During model fine-tuning, freeze some layers of the pre-trained model to retain the general feature representation learned in the pre-training stage; train the last few layers of the model to adapt to the wild plant and animal recognition task, which can reduce the amount of calculation while utilizing the general features of the pre-trained model. In the fine-tuning process, combine image, audio, text, environmental and positioning data to enhance the model's understanding of multi-modal data. Use cross-entropy loss function for training to optimize model parameters.
[0163] After the model is fine-tuned, the general features associated with the pre-trained model are transferred to the pre-trained model to optimize the pre-trained model. For example, the knowledge of the pre-trained model in other fields is transferred to the wild plant and animal species identification task, the adaptability and generalization ability of the model to wild plant and animal data are improved through transfer learning, and the understanding ability of the model to multi-modal data is further improved.
[0164] After the model is fine-tuned and the features are transferred to optimize the model, the pre-trained model is pruned and quantized to determine the target large model for species identification through model compression. When pruning, the pre-trained model is pruned to remove unimportant weights or neurons; a structured pruning method is used to preserve the structural integrity of the model; through pruning, the parameter quantity of the model is reduced and the calculation efficiency is improved. When quantizing, the weights and activation values of the model are quantized to low-precision representations using quantization techniques. Through pruning and quantization, the parameter quantity and storage space of the model can be reduced, and the calculation efficiency can be improved. After pruning and quantization, the performance of the model needs to be evaluated on the validation set to ensure that the pruning operation and quantization operation do not significantly reduce the accuracy of the model.
[0165] Through the above process, a Transformer architecture model suitable for multi-modal data processing can be selected for pre-training and fine-tuning, the generalization ability of the model can be improved using transfer learning techniques, and the parameter quantity of the model can be reduced through model compression, thereby improving the calculation efficiency and storage efficiency of the model.
[0166] After the target large model is determined through model training, it further includes:
[0167] Based on the new sample data provided by the collection device, a comprehensive feature vector and a label associated with the sample are generated, and the model parameters of the target large model are updated using an online learning algorithm based on the comprehensive feature vector and the label associated with the sample; and / or
[0168] The annotation information obtained after the user annotates the recognition result output by the target large model is added to the sample data to update the sample data, and the model parameters of the target large model are adjusted based on the updated sample data.
[0169] The embodiments of the present application can use an online learning algorithm and / or an incremental learning mechanism to adjust the model parameters to optimize the model performance. When updating the model parameters using an online learning algorithm, an online gradient descent algorithm or a stochastic gradient descent algorithm is selected. After obtaining the new sample data provided by the collection device, a comprehensive feature vector and a label associated with the sample are generated, and the model parameters of the target large model are updated using an online learning algorithm based on the comprehensive feature vector and the corresponding label, such as using an online gradient descent algorithm to update the model parameters to optimize the model performance.
[0170] When adjusting the model parameters based on the incremental learning mechanism, the model is allowed to incorporate the labeled data into the model training when the user labels or feedbacks the recognition results, and continuously optimizes the model performance. In this process, the user labels or feedbacks the recognition results through the interactive interface, collects the user-labeled data, including images, audio, positioning and environmental data and their corresponding labels. The user-labeled data is added to the training data set, and the incremental learning algorithm is used to fine-tune the model and update the model parameters; for example, using small batch gradient descent for incremental learning.
[0171] Through the online learning algorithm and the incremental learning mechanism, real-time dynamic learning and updating of the model can be realized; the online learning algorithm allows the model to receive new data in real time and dynamically adjust the parameters, quickly adapting to environmental changes and the appearance of new species; the incremental learning mechanism allows the model to incorporate the labeled data into the model training when the user labels or feedbacks the recognition results, and continuously optimizes the model performance. The above-mentioned model optimization method can significantly improve the adaptability and accuracy of the model, and provide stronger support for species identification.
[0172] In addition to wild plant and animal species identification, the embodiments of the present application can also analyze the behavior patterns of species based on the species identification results and historical data, predict their future behavior trends, and provide deeper insights for ecological research; and, combined with environmental data and wild plant and animal distribution and behavior information, assess the impact of environmental changes on biodiversity, and provide scientific basis for ecological protection decisions.
[0173] For the identified species, when analyzing its behavior patterns and predicting future behavior trends based on the identification results and historical data, the following steps need to be performed:
[0174] 1. Data collection; collect species identification results, behavior data (such as motion trajectory, activity time) and environmental data, and label the behavior data, such as foraging, migration, breeding, etc.
[0175] 2. Behavior analysis; extract the behavior characteristics of the species, such as movement speed, activity frequency, etc., and use time series analysis models or reinforcement learning algorithms for behavior modeling.
[0176] 3. Behavior prediction; use the trained model to predict behavior, and output future behavior trends.
[0177] The main task of behavior analysis and prediction is to analyze the behavior patterns of species based on the identification results and historical data, and predict their future behavior trends, which will be described in detail as follows:
[0178] I. Data collection and labeling
[0179] (1) Data collection
[0180] Based on the species recognition, obtain the species name, confidence, etc. Collect behavior data of species (take animals as an example), such as movement trajectory, activity time, behavior type, etc. Collect environment data related to behavior, such as temperature, humidity, water quality, etc.
[0181] (2) Data annotation
[0182] Behavior annotation: annotate the collected behavior data, such as foraging, migration, breeding, etc. Timestamp alignment: ensure that all data (recognition results, behavior data, environment data) have accurate timestamps for time series analysis.
[0183] II. Feature extraction
[0184] (1) Behavior features
[0185] Extract the movement trajectory of the species, including speed, acceleration, direction change, etc. Statistics of activity frequency of species in different time periods. Record the duration of each behavior.
[0186] (2) Environmental features
[0187] Record the change of environmental temperature, record the change of environmental humidity, record the change of water quality, such as pH value, dissolved oxygen, etc.
[0188] III. Model selection and training
[0189] (1) Model selection
[0190] Time series analysis model: use LSTM (Long Short-Term Memory Network) or GRU (Gated Recurrent Unit) for time series analysis. Reinforcement learning model: use reinforcement learning algorithm for behavior modeling.
[0191] (2) Model training
[0192] Data preparation, combine the extracted features and annotated behavior data into a training data set. When training the model, use the training data set to train the time series analysis model or reinforcement learning model.
[0193] IV. Behavior prediction
[0194] Model inference; input current state: input the current behavior features and environment features of the species into the trained model. Output prediction results: the model outputs the prediction results of future behavior, such as the behavior type of the next time period.
[0195] In the process of combining environmental data and wild animal and plant distribution, behavior information to assess the impact of environmental changes on biodiversity, data collection and comprehensive data analysis are needed, including correlation analysis and regression analysis. Correlation analysis can help identify the relationship between environmental variables and wild animal and plant distribution, behavior, and regression analysis can further assess the impact of environmental changes on biodiversity. Through these analyses, scientific basis can be provided for ecological protection and resource management.
[0196] During data collection, environmental data and wild animal and plant data are collected. Environmental data includes, for example, temperature (water temperature and air temperature), air humidity, water quality, and other environmental data (such as wind speed, wind direction, light intensity, etc.). Wild animal and plant data includes distribution data and behavior data. Distribution data includes, for example, species distribution, species number, etc. Behavior data includes, for example, behavior type (foraging, migration, breeding, etc.), behavior frequency, behavior duration, etc. The collected data needs to be accompanied by accurate time stamp and geographic location information for subsequent analysis.
[0197] During data analysis, correlation analysis and regression analysis are used. In correlation analysis, statistical analysis methods (such as Pearson correlation coefficient) are used to analyze the correlation between environmental data and wild animal and plant distribution, behavior. Specifically, data preprocessing is performed, such as data cleaning and standardization to ensure data quality. Statistical software is used to calculate correlation coefficients. Scatter plots, heat maps, and other visualization tools are used to display correlation analysis results. In regression analysis, regression models (such as linear regression, logistic regression) are used to assess the impact of environmental changes on biodiversity. Specifically, data preprocessing is performed, such as data cleaning and standardization to ensure data quality. Suitable regression models are selected, and training data sets are used to train the model. Validation data sets are used to evaluate model performance, and appropriate evaluation metrics are selected. Regression model coefficients are interpreted to assess the impact of environmental variables on biodiversity.
[0198] The embodiments of the present application aim to utilize large model technology to build an efficient and intelligent wild animal and plant species identification system, achieve accurate species identification, behavior analysis, and ecological environment monitoring, and provide strong support for biodiversity protection.
[0199] The embodiments of the present application provide a wild animal and plant species identification system, as shown in Figure 3 The wild animal and plant species identification system 300 includes a collection device 301 and a cloud platform 302.
[0200] The collection device 301 collects biological data and environmental data associated with the target species, the biological data including image data, audio data and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environment data associated with the living environment of the target species, the target species being a wild animal or a wild plant;
[0201] The cloud platform 302 pre-processes the data provided by the collection device 301, obtains the to-be-analyzed data corresponding to the target species, extracts features from the to-be-analyzed data, and aligns the extracted multi-modal features based on time, the multi-modal features including biological features and environmental features, the biological features including image features extracted from image data, audio features extracted from audio data, and positioning features extracted from positioning data;
[0202] The cloud platform 302 performs feature fusion on the aligned multi-modal features based on the dynamically adjusted weights of the modal features, generates a comprehensive feature vector indicating the comprehensive feature information of the target species, and dynamically adjusts the weights of the modal features based on the real-time contribution of the multi-modal features to species identification;
[0203] The cloud platform 302 performs inference analysis on the comprehensive feature vector based on the deployed target large model, obtains at least one species name corresponding to the target species and the confidence of each species name, and determines the species name of the target species based on the confidence of each species name.
[0204] Optionally, when performing data preprocessing to obtain the to-be-analyzed data, the cloud platform 302 is further configured to: perform image denoising, image enhancement and image standardization processing on the image data associated with the target species to obtain to-be-analyzed image; remove background noise in the audio data associated with the target species to obtain to-be-analyzed audio; convert the latitude and longitude coordinates of the target species into plane coordinates, and combine the data measured by the inertial measurement unit to obtain to-be-analyzed positioning data; and standardize the environmental data associated with the target species to a unified range to obtain to-be-analyzed environmental data.
[0205] Optionally, when performing feature alignment, the cloud platform 302 is further configured to: extract image features, audio features, positioning features and environmental features carrying timestamps from the to-be-analyzed data; and align the extracted features of different modalities according to the timestamps to obtain multi-modal features consistent in the time dimension.
[0206] Optionally, the cloud platform 302 is further configured to, when generating the comprehensive feature vector, perform feature splicing on the aligned image features, audio features, positioning features, and environmental features to generate a high-dimensional feature vector, assign an initial weight to each modality feature in the multi-modal feature, and fuse the assigned initial weight into the high-dimensional feature vector to generate a baseline feature vector, wherein the initial weight of the environmental feature is determined based on the influence degree of the environmental data on the survival and identification of the target species, and the initial weight of each modality feature in the biological feature is assigned based on the estimated importance of each modality feature in species identification.
[0207] The weights of the modality features in the baseline feature vector are dynamically adjusted based on target parameters, and a weight smoothing biological constraint mechanism is used to control the weight variation amplitude, a comprehensive feature vector is generated based on the baseline feature vector with dynamically adjusted weights, and the target parameters include at least one of the following: a full-modal weight corresponding to a system time period, a confidence of the target large model for each modality feature, feedback information of a user on an identification result output by the target large model in a previous identification period, newly introduced sample data in a current identification period, and a scene in which the current identification period is located.
[0208] Optionally, the cloud platform 302 is further configured to, when determining the species name of the target species, optimize the target large model based on the comprehensive feature vector corresponding to the target species, and perform inference analysis on the comprehensive feature vector based on the optimized target large model; and obtain at least one species name corresponding to the target species and a confidence of each species name based on an inference analysis result output by the optimized target large model.
[0209] The confidence of at least one species name is corrected based on prior information of the target species.
[0210] The species name of the target species is determined based on the confidence of the corrected species name.
[0211] Optionally, the cloud platform 302 is further configured to fuse the comprehensive feature vector corresponding to the target species with the feature representation of the target large model, and dynamically adjust the weights of different modality features to optimize the target large model.
[0212] The enhanced comprehensive feature vector is analyzed based on the optimized target large model, or the key fusion feature is analyzed based on the optimized target large model, wherein the key fusion feature is generated based on the fusion of the multi-dimensional key features extracted from the comprehensive feature vector.
[0213] Optionally, the cloud platform 302 is further configured to: pre-train an architecture model suitable for multi-modal data processing based on first sample data corresponding to the first biological sample set, to determine a pre-trained model; adjust model parameters of the pre-trained model based on second sample data corresponding to the second biological sample set and subjected to data labeling, to fine-tune the pre-trained model; migrate general features associated with the pre-trained model to the pre-trained model, to optimize the pre-trained model; and prune and quantize the pre-trained model subjected to model fine-tuning and model optimization, to determine a target large model for species identification.
[0214] Optionally, the cloud platform 302 is further configured to: generate sample-associated comprehensive feature vectors and labels based on new sample data provided by the acquisition device 301, and update model parameters of the target large model according to the sample-associated comprehensive feature vectors and labels and by using an online learning algorithm; and / or
[0215] Obtain labeled information after a user labels an identification result output by the target large model, add the labeled information to the sample data to update the sample data, and adjust model parameters of the target large model based on the updated sample data.
[0216] For the system embodiment, it is basically similar to the method embodiment, and thus is described simply. For relevant parts, refer to the description of the method embodiment.
[0217] Embodiments of the present application also provide an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above-mentioned wild animal and plant species identification method embodiment and achieves the same technical effects. To avoid repetition, no further description is given here.
[0218] For example, Figure 4 An entity structure diagram of an electronic device is shown. As Figure 4 shown, the electronic device can include a processor 410, a communications interface 420, a memory 430, and a communications bus 440. The processor 410, the communications interface 420, and the memory 430 can communicate with each other through the communications bus 440. The processor 410 can invoke logical instructions in the memory 430, and the processor 410 is configured to execute each process of the wild animal and plant species identification method of embodiments of the present application. No further description is given here.
[0219] Further, the logic instructions in the memory 430 described above can be implemented in the form of software functional units and sold or used as standalone products, and can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art, or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0220] The embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement each process of the identification method for wild animal and plant species, and achieve the same technical effects. To avoid repetition, details are not described herein. The computer readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0221] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, methods, articles, or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles, or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article, or device that includes the element.
[0222] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, and of course, they can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such an understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, a magnetic disk, an optical disk), and includes a number of instructions for causing a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0223] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims.
Claims
1. A method of identifying a wild animal or plant species, characterized by, The method comprises the following steps: preprocessing biological data and environmental data associated with a target species to obtain analysis data corresponding to the target species, the biological data including image data, audio data and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environment data associated with the living environment of the target species, the target species being a wild animal or a wild plant; extracting features from the analysis data and aligning the extracted multi-modal features based on time, the multi-modal features including biological features and environmental features, the biological features including image features extracted from the image data, audio features extracted from the audio data and positioning features extracted from the positioning data; performing feature fusion on the aligned multi-modal features based on dynamically adjusted weights of the modal features to generate a comprehensive feature vector indicating comprehensive feature information of the target species, the weights of the modal features being dynamically adjusted based on real-time contributions of the multi-modal features to species identification; wherein the aligned image features, audio features, positioning features and environmental features are spliced to generate a high-dimensional feature vector; assigning an initial weight to each of the multi-modal features, fusing the assigned initial weights to the high-dimensional feature vector to generate a baseline feature vector, the initial weight of the environmental features being determined based on the influence of the environmental data on the survival and identification of the target species, and the initial weights of the modal features in the biological features being assigned based on the estimated importance of each modal feature in species identification; dynamically adjusting the weights of the modal features in the baseline feature vector based on a target parameter and controlling the weight change amplitude based on a weight smoothing biological constraint mechanism, and generating the comprehensive feature vector from the baseline feature vector with dynamically adjusted weights; performing inference analysis on the comprehensive feature vector based on a target large model to obtain at least one species name corresponding to the target species and the confidence of each species name, and determining the species name of the target species based on the confidence of each species name, comprising: optimizing the target large model according to the comprehensive feature vector corresponding to the target species, and performing inference analysis on the comprehensive feature vector based on the optimized target large model, specifically including: fusing the comprehensive feature vector corresponding to the target species with the feature representation of the target large model, and dynamically adjusting the weights of different modal features to optimize the target large model; obtaining at least one species name corresponding to the target species and the confidence of each species name according to the inference analysis result output by the optimized target large model; correcting the confidence of the at least one species name based on prior information of the target species; determining the species name of the target species based on the confidence of the corrected species name.
2. The method of claim 1, wherein the step of identifying the wild animal or plant species is performed by using a database of wild animal or plant species. The preprocessing of the biological data and the environmental data associated with the target species to obtain the analysis data corresponding to the target species comprises: image denoising, image enhancement and image standardization are performed on image data associated with the target species to obtain to-be-analyzed images; background noise in audio data associated with the target species is removed to obtain to-be-analyzed audio; latitude and longitude coordinates of the target species are converted into plane coordinates, and to-be-analyzed positioning data are obtained in combination with data measured by an inertial measurement unit; environmental data associated with the target species is standardized to a unified range to obtain to-be-analyzed environmental data.
3. The method of claim 1, wherein the step of identifying the wild animal or plant species is performed by using a database of wild animal or plant species. The feature extraction is performed on the to-be-analyzed data, and the extracted multi-modal features are time-aligned, comprising: image features, audio features, positioning features and environmental features carrying timestamps are extracted from the to-be-analyzed data; the extracted features of different modalities are aligned according to timestamps to obtain multi-modal features consistent in the time dimension.
4. The wild animal and plant species identification method according to claim 1, characterized in that the target parameters include at least one of the following: a full-modal weight corresponding to a system time period, a confidence of the target large model for each modal feature, feedback information of a user on an identification result output by the target large model in a previous identification period, sample data newly introduced in a current identification period, and a scene in which the current identification period is located.
5. The method of claim 1, wherein the step of identifying the wild animal or plant species is performed by using a database of wild animal or plant species. The target large model is optimized according to the comprehensive feature vector corresponding to the target species, and the comprehensive feature vector is analyzed by inference based on the optimized target large model, and further comprising: the enhanced comprehensive feature vector is analyzed by inference based on the optimized target large model, or the key fusion feature is analyzed by inference based on the optimized target large model; wherein the key fusion feature is generated based on fusion of multi-dimensional key features extracted in the comprehensive feature vector.
6. The method of identifying a wild animal or plant species according to any one of claims 1 to 5, wherein, The method further comprises: pre-training an architecture model suitable for multi-modal data processing based on first sample data corresponding to a first biological sample set to determine a pre-trained model; adjusting model parameters of the pre-trained model based on second sample data corresponding to a second biological sample set and subjected to data annotation to fine-tune the pre-trained model; migrating general features associated with the pre-trained model to the pre-trained model to optimize the pre-trained model; pruning and quantifying the pre-trained model subjected to model fine-tuning and model optimization to determine a target large model for species identification.
7. The method of claim 6, wherein the step of identifying the wild animal or plant species is performed by a method comprising: Further comprising: generating sample-associated comprehensive feature vectors and labels based on new sample data provided by a collection device, and updating model parameters of the target large model using an online learning algorithm according to the sample-associated comprehensive feature vectors and labels; and / or obtaining annotation information after a user annotates an identification result output by the target large model, adding the annotation information to sample data to update the sample data, and adjusting model parameters of the target large model based on the updated sample data.
8. A system for identifying a wild animal or plant species, characterized in that comprising: a collection device and a cloud platform; The collection device collects biological data associated with a target species and environmental data, the biological data including image data, audio data and positioning data associated with the target species, and the environmental data including at least meteorological data and physical environmental data associated with the living environment of the target species, the target species being a wild animal or a wild plant; The cloud platform pre-processes the data provided by the collection device, obtains the to-be-analyzed data corresponding to the target species, extracts features from the to-be-analyzed data, and aligns the extracted multi-modal features based on time, the multi-modal features including biological features and environmental features, the biological features including image features extracted from the image data, audio features extracted from the audio data, and positioning features extracted from the positioning data; The cloud platform performs feature fusion on the aligned multi-modal features based on dynamically adjusted weights of the modal features, generates a comprehensive feature vector indicating comprehensive feature information of the target species, and dynamically adjusts the weights of the modal features based on real-time contributions of the multi-modal features to species identification; wherein, when generating the comprehensive feature vector, the cloud platform is further configured to: perform feature splicing on the aligned image features, audio features, positioning features and environmental features to generate a high-dimensional feature vector; An initial weight is assigned to each of the multi-modal features, the assigned initial weight is fused into the high-dimensional feature vector to generate a baseline feature vector, the initial weight of the environmental features is determined based on the influence of the environmental data on the survival and identification of the target species, and the initial weights of the modal features in the biological features are assigned based on the estimated importance of the modal features in species identification; The weights of the modal features in the baseline feature vector are dynamically adjusted based on a target parameter, and the weight variation range is controlled based on a weight smoothing biological constraint mechanism, and the comprehensive feature vector is generated according to the baseline feature vector with dynamically adjusted weights; The cloud platform performs inference analysis on the comprehensive feature vector based on a deployed target large model, obtains at least one species name corresponding to the target species and the confidence of each species name, and determines the species name of the target species based on the confidence of each species name, including: The target large model is optimized according to the comprehensive feature vector corresponding to the target species, and inference analysis is performed on the comprehensive feature vector based on the optimized target large model, specifically including: fusing the comprehensive feature vector corresponding to the target species with the feature representation of the target large model, and dynamically adjusting the weights of different modal features to optimize the target large model; According to the inference analysis result output by the optimized target large model, at least one species name corresponding to the target species and the confidence of each species name are obtained; The confidence of the at least one species name is corrected based on prior information of the target species; The species name of the target species is determined based on the confidence of the corrected species name.
9. The system for identifying wild animal and plant species according to claim 8, characterized in that, The target parameter comprises at least one of the following: a full-modal weight corresponding to a system time period, a confidence of the target large model for each modal feature, feedback information of a user on an identification result output by the target large model in a previous identification period, sample data newly introduced in a current identification period, and a scene in which the current identification period is located.
Citation Information
Patent Citations
Multi-modal fusion bird identification method and device
CN118430012A
Ambient sound event detection method based on multi-modal data fusion
CN119446154A