Species detection and identification method based on hybrid scaling strategy and cascade architecture
The species detection and identification method using a hybrid scaling strategy and cascaded architecture solves the problem of independent optimization of the detection model and the classification model, and achieves efficient and accurate species identification in complex environments, making it suitable for intelligent and large-scale applications in ecological protection.
Patent Information
- Application Number
- CN202510955402.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-31
AI Technical Summary
In existing species identification technologies, the independent optimization of detection and classification models leads to the inability to dynamically adjust resource allocation, and a single image scaling strategy is difficult to balance the needs of multi-scale target identification, resulting in low identification accuracy, especially in complex environments where it is difficult to distinguish closely related species.
By employing a hybrid scaling strategy and a cascaded architecture, a cascaded model of detectors and classifiers is constructed by collaboratively adjusting depth, width, and resolution. Combined with transfer learning and model fine-tuning, the network structure and parameters are optimized to achieve a complete process from coarse-grained to fine-grained.
It improves the model's recognition accuracy and robustness in complex environments, reduces computational resource consumption, is suitable for species recognition across scenarios and at multiple scales, and supports intelligent and large-scale applications in ecological protection.
Smart Images

Figure CN120876945A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, and in particular to a species detection and recognition method based on a hybrid scaling strategy and a cascaded architecture. It integrates image preprocessing technology, neural network model construction and expansion technology, and model transfer learning and fine-tuning technology. It adopts a hybrid scaling strategy to optimize the network structure and organically connects the detection network and the recognition network to form an integrated technology link, thereby achieving fine-grained detection and accurate recognition of species. Background Technology
[0002] With the intensification of global biodiversity change, the social value of species identification technology is becoming increasingly prominent. In fields such as ecological protection, natural resource management, and public science participation, rapid and accurate species identification has become a core requirement supporting biodiversity monitoring. Because traditional manual identification is time-consuming, labor-intensive, and lacks guaranteed accuracy, deep learning technology has been applied to detection and classification tasks in various situations. The introduction of deep learning technology has brought about a paradigm shift in species identification. Through algorithms such as convolutional neural networks, computers can automatically learn the image features of species, achieving rapid detection and classification in massive amounts of data. This technological breakthrough has demonstrated high efficiency in scenarios such as wildlife camera trap data processing and plant remote sensing monitoring, significantly reducing labor costs and increasing monitoring coverage. However, the existing technological framework still has structural bottlenecks: the independent optimization mechanisms of detection and classification models prevent the dynamic adjustment of system resource allocation according to the actual scenario. For example, in complex habitats, the detection module may focus too much on target localization and ignore feature details, while the classification module may misclassify species due to a lack of contextual information. Furthermore, a single image scaling strategy struggles to balance the recognition needs of multi-scale targets: small-scale targets are prone to losing key features due to insufficient resolution during magnification, while large-scale targets may suffer morphological distortion due to compression; both significantly reduce recognition accuracy. In wildlife monitoring scenarios, independently optimized detection models may focus excessively on target localization while neglecting fine-grained features such as fur texture and markings, resulting in the classification module failing to acquire sufficient discriminative information. On the other hand, a single scaling strategy in recognition tasks is prone to problems such as blurred details after magnification or overall morphological distortion after compression, making it difficult for the model to distinguish closely related species. Therefore, there is an urgent need for a technology that integrates hybrid scaling strategies with adaptive cascaded optimization to achieve accurate and efficient recognition across scenarios and multiple scales, supporting large-scale applications in ecological conservation. Summary of the Invention
[0003] To address the aforementioned technical challenges, this invention provides a species detection and identification method based on a hybrid scaling strategy and a cascaded architecture. By innovatively designing a hybrid scaling mechanism that coordinates the adjustment of depth, width, and resolution, it improves accuracy while optimizing computational efficiency. Simultaneously, the cascaded architecture organically connects the detector and classifier, constructing a complete process from coarse-grained target localization to fine-grained feature analysis, achieving adaptive collaborative optimization of the detection and classification modules. This method aims to enhance the robustness of detection and identification models in complex environments, providing technical support for building high-precision, low-energy-consumption species detection and identification systems. It promotes the upgrading of biodiversity monitoring towards intelligent and large-scale applications, assists in the transformation of ecological protection from passive response to proactive intelligent management, improves the timeliness and effectiveness of response strategies, and provides reliable support for rapid response and adjustment at the policy and practical levels.
[0004] The technical solution of this invention is:
[0005] A species detection and identification method based on a hybrid scaling strategy and a cascaded architecture, including
[0006] 1. Species Image Acquisition: The equipment is deployed in the target area (such as forest areas or the boundary of protected areas), and uses thermal sensing triggering and motion detection algorithms to control the shooting, prioritizing the capture of live targets to reduce invalid images. The acquired raw images are transmitted back to the local data center via 4G networks and other means.
[0007] 2. Image Preprocessing: For the acquired dataset, the image size is first adjusted by filling in edge regions to ensure a uniform aspect ratio. Then, the images are normalized to balance pixel value distribution and format conversion is performed. When filling edge regions, a symmetrical filling strategy is used; after the long side of the image is scaled to the target size, the remaining short side is filled with gray to avoid introducing noise unrelated to the background. Normalization uses channel-wise mean-variance standardization to adapt to the input distribution of the pre-trained model. In the format conversion stage, the image tensor is adjusted to CHW (channels × height × width) format to be compatible with GPU-accelerated computation. Preprocessing eliminates data format differences, laying a standardized data foundation for efficient recognition by the subsequent network model.
[0008] 3. Model Construction: The technical solution for species detection and fine-grained recognition is based on a convolutional network model architecture. First, a baseline network is built, and its structure is optimized through scaling strategies to adapt to data inputs with different features. Then, the object detection module and the fine-grained classification module are cascaded and integrated to form a complete detection and recognition model, achieving integrated processing from object detection to accurate classification.
[0009] 4. Transfer Learning and Model Fine-tuning: Based on the pre-trained model, its learned general features are extracted. For the target species identification task, key layers are selected for parameter adjustment, and model specificity is optimized by injecting labeled data. Data augmentation strategies are combined to improve generalization ability, ultimately achieving efficient model transfer in new scenarios and reducing training costs.
[0010] 5. API Service Deployment and Platform Integration: The front-end platform calls the API via POST requests to upload image data. An asynchronous processing mechanism improves throughput efficiency in high-concurrency scenarios. The model performs object detection and fine-grained classification on the preprocessed images, generating species labels and confidence scores. The returned data is encapsulated in JSON format. The front-end platform receives the prediction results in JSON format, parses them, and dynamically renders the species information.
[0011] Furthermore,
[0012] Infrared cameras and video monitoring equipment are used to acquire images of species in the wild; image data is transmitted over the network via 4G network and wireless bridge to build a dataset.
[0013] A species detection and fine-grained recognition model is constructed based on a neural network. The MBConv structure is used to build the baseline network, and the network is improved on the baseline network through a hybrid scaling strategy. The detector network and the classifier network are cascaded to form a detection and recognition model, which performs fine-grained classification on the detected targets.
[0014] Model fine-tuning and optimization based on collected data: The collected data is divided into training set, validation set and test set in a ratio of 7:2:1. The Adam optimization algorithm is used, the initial learning rate is set, the batch size and number of iterations are adjusted to carry out transfer learning of the model, and the validation set is used for real-time evaluation. The optimal model parameters are saved, and new data is collected periodically to update the model parameters.
[0015] Let the scaling factor r be the value used in image preprocessing.
[0016]
[0017] Scale the image proportionally to the specified size. Calculate the fill amount and adjust by increments:
[0018] dw=(old_width-new_width)%stride
[0019] dh=(old_height-new_height)%stride
[0020] Center the image and fill it with color symmetrically on the top and bottom / left and right sides; normalize the scaled and filled image by using channel mean-variance standardization to normalize the pixel values from [0, 255] to [0, 1] to adapt to the input distribution of the pre-trained model; finally, adjust the channel order and format, converting the channel order to the RGB format required by the model and the tensor to CHW format.
[0021] The construction steps in the model construction are as follows:
[0022] The MBConv structure is composed of a 1*1 ordinary convolution, a 3*3 or 5*5 Depthwiss Conv convolution, an SE module, a 1*1 ordinary convolution, and a Drouβout layer. The SE module improves the network's representation ability by modeling the interdependencies between convolutional feature channels through three steps: global information embedding, adaptive recalibration, and reweighting.
[0023] By stacking MBConv structures and combining them with 1*1 convolutional layers, average pooling layers, and fully connected layers, a baseline network can be constructed.
[0024] Based on the baseline network, a hybrid scaling strategy is used to extend the network from different dimensions simultaneously; a hybrid factor is used. The parameters for uniform scaling—depth, width, and resolution—are calculated using the following formulas:
[0025] depth: d = α φ
[0026] width: w = β φ
[0027] resolutiOn:r=γ φ
[0028] stα·β 2 ·γ 2 ≈2
[0029] α≥1, β≥1, γ≥1
[0030] Where st represents the constraints, and α, β, and γ are the parameters corresponding to depth, width, and resolution, respectively; these are first set during scaling. The optimal α, β, and γ parameters are searched in the baseline network. After obtaining the optimal parameters, the parameters are kept unchanged, and the scaled network framework can be obtained by adjusting α, β, and γ. The baseline network and the scaled network are combined to obtain the network structure for feature extraction. During the scaling process, depth expansion is preferentially applied to shallow networks to preserve detailed features, width expansion is concentrated in deep networks to enhance semantic expressive power, and resolution improvement is increased synchronously with the network layers to ensure effective fusion of multi-scale features. Dropout layers and Dense layers are added to the network structure to classify the features extracted by the network, and finally, a complete model architecture is obtained.
[0031] The transfer learning and model fine-tuning steps in the model construction are as follows:
[0032] The pre-trained model should retain the structure of the backbone convolutional layers and freeze their parameters to avoid destroying the learned basic features during fine-tuning; the top fully connected layer of the pre-trained model is designed for the original dataset and should be deleted and replaced with a new classification layer; the number of neurons in the new layer should be consistent with the number of target species categories, and initial parameters should be assigned through random initialization or Xavier initialization strategy.
[0033] Initialize model parameters using pre-trained weights, and start training with a low learning rate to avoid damaging existing features. Gradually adjust the learning rate based on the training curve, and increase it later to optimize new layer parameters. Choose a cross-entropy loss function suitable for classification tasks.
[0034]
[0035] Where M is the number of categories, y ic The sign function is set to 1 if the true class of sample i equals c, and 0 otherwise; it is used in conjunction with the Adam optimizer.
[0036] m t =β1m t-1 +(1-β1)g t
[0037]
[0038] To address the class imbalance problem of the target species, weight adjustment can be introduced; on the validation set, the accuracy and loss values are monitored, and early stopping rules are set to prevent overfitting and shorten the training cycle; during the validation set monitoring phase, in addition to the regular accuracy, the macro-F1 score is tracked to more comprehensively evaluate the model's ability to identify a few classes.
[0039] The early stopping rule adopts a dual judgment mechanism: if the validation loss does not decrease for 10 consecutive epochs, or the macro average F1 score fluctuates within a range of less than 0.5% for 5 consecutive epochs, early stopping is triggered; after training is terminated, the model weights that perform best in the validation set are automatically rolled back, and a training curve report is generated, including the loss value decay trend and heatmap, for subsequent optimization reference.
[0040] Training cycle compression is achieved through dynamic batch adjustment: initially, a small batch size of 16 is used to ensure stability, and then gradually increased in the middle and later stages to accelerate convergence.
[0041] The transfer learning and model fine-tuning steps in the model construction are as follows:
[0042] The front-end platform calls the interface via POST requests to upload image data. It adopts an asynchronous processing mechanism to improve throughput efficiency in high-concurrency scenarios. The model performs object detection and fine-grained classification on the preprocessed image, generating species labels and confidence results. The returned data is encapsulated in JSON format. The front-end platform receives the prediction results in JSON format, parses them, and dynamically renders the species information.
[0043] The beneficial effects of this invention are
[0044] 1. Enhance model generalization ability: The hybrid scaling strategy adjusts the network structure (such as depth, width, and resolution) in multiple dimensions, allowing the model to be exposed to richer feature representations during training. This reduces the risk of overfitting to specific scales or scenes, enhances the robustness of identification of unknown environments and rare species, and is suitable for cross-regional and cross-habitat diversity monitoring needs.
[0045] 2. Optimize detection and classification collaboration efficiency: The cascaded architecture constructs a complete process for detection and classification modules, enabling the model to first complete coarse-grained target localization, and then perform fine-grained feature analysis based on the localization results. This reduces redundant computation while improving classification accuracy, achieving hierarchical and efficient collaboration from "target discovery" to "accurate identification".
[0046] 3. Enhance robustness in complex environments: Improve the model's adaptability to non-ideal conditions such as changes in lighting and background interference, reduce false detections and misjudgments caused by environmental noise, and ensure that the recognition method remains stable and reliable in complex scenarios such as field monitoring, thereby reducing the cost of manual verification. Attached Figure Description
[0047] Figure 1 This is a schematic diagram of the workflow of the present invention;
[0048] Figure 2 This is a schematic diagram of the image preprocessing process of the present invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0050] This invention provides a method for species detection and fine-grained identification based on a hybrid scaling strategy, comprising the following steps:
[0051] Data acquisition and preprocessing. Images of species in the wild are acquired using infrared cameras, video monitoring equipment, and other means. Image data is transmitted over the network via 4G networks and wireless bridges to construct a dataset. For the acquired dataset, image sizes are adjusted by filling in edge regions, and normalization and format adjustments are performed to unify the data format, preparing for subsequent network recognition.
[0052] A species detection and fine-grained recognition model is constructed based on neural networks. An MBConv architecture is used to build a baseline network, which is then improved through a hybrid scaling strategy. The detector network and classifier network are cascaded to form a detection and recognition model, which performs fine-grained classification of detected targets.
[0053] Model fine-tuning and optimization based on collected data: The collected data was divided into training, validation, and test sets in a 7:2:1 ratio. The Adam optimization algorithm was used, with an initial learning rate set, batch size and number of iterations adjusted for transfer learning of the model. The validation set was used for real-time evaluation, the optimal model parameters were saved, and new data was collected periodically to update the model parameters.
[0054] The data preprocessing in the data acquisition and preprocessing steps specifically includes:
[0055] First, the input image is uniformly adjusted to a fixed size, maintaining the original aspect ratio to avoid image distortion. Simultaneously, edge regions are filled; when the longer side is adjusted to the required length, the remaining portion of the shorter side is filled with gray to avoid introducing noise unrelated to the background and ensure the input size meets the model's stride constraint. The scaling ratio r is calculated.
[0056]
[0057] Scale the image proportionally to the specified size. Calculate the fill amount and adjust by increments:
[0058] dw=(old_width-new_width)%stride
[0059] dg=(old_height-new_height)%stride
[0060] The image is centered and symmetrically filled with color both top and bottom / left and right. The scaled and filled image is then normalized using channel-wise mean-variance standardization, reducing pixel values from [0,255] to [0,1] to fit the input distribution of the pre-trained model. Finally, the channel order and format are adjusted, converting the channel order to the RGB format required by the model and the tensor to CHW format. This avoids information cropping loss while ensuring the validity of the feature map size through stride constraints, balancing processing efficiency and model accuracy. Data preprocessing unifies the data format, facilitating later detection and recognition using the model.
[0061] The construction steps in the model construction are as follows:
[0062] The MBConv structure is constructed by sequentially using a 1*1 ordinary convolution, a 3*3 or 5*5 DepthwiseConv convolution, an SE module, a 1*1 ordinary convolution, and a Dropout layer. The SE module improves the network's representational ability by modeling the interdependencies between convolutional feature channels through three steps: global information embedding, adaptive recalibration, and reweighting.
[0063] A baseline network can be constructed by stacking MBConv structures and combining them with 1*1 convolutional layers, average pooling layers, and fully connected layers.
[0064]
[0065] Based on the baseline network, a hybrid scaling strategy is used to extend the network from different dimensions simultaneously. A hybrid factor is used. The parameters for uniform scaling—depth, width, and resolution—are calculated using the following formulas:
[0066] depth: d = α φ
[0067] width: w = β φ
[0068] resolution: r = γ φ
[0069] stα·β 2 ·γ 2 ≈2
[0070] α≥1, β≥1, γ≥1
[0071] Where st represents the constraints, and α, β, and γ are the parameters corresponding to depth, width, and resolution, respectively. These are first set during scaling. The optimal α, β, and γ parameters are searched in the baseline network, starting with 1. After obtaining the optimal parameters, these parameters are kept constant, and the scaled network framework is obtained by adjusting α, β, and γ. Combining the baseline network with the scaled network yields the network structure for feature extraction. During scaling, depth expansion is preferentially applied to shallower layers to preserve detailed features, width expansion is concentrated in deeper layers to enhance semantic expressiveness, and resolution improvement is increased synchronously with network layers to ensure effective fusion of multi-scale features. Dropout and Dense layers are added to the network structure for classifying the extracted features, resulting in the complete model architecture.
[0072] The transfer learning and model fine-tuning steps in the model construction are as follows:
[0073] The bottom convolutional layers of a pre-trained model typically learn basic visual features such as edges and textures, which are universal across tasks. Therefore, the structure of the backbone convolutional layers should be preserved, and their parameters frozen to avoid destroying the learned basic features during fine-tuning. The top fully connected layers of the pre-trained model, designed for the original dataset, need to be removed and replaced with new classification layers. The number of neurons in the new layer should match the number of target species categories, and initial parameters should be assigned using random initialization or a Xavier initialization strategy.
[0074] Initialize model parameters using pre-trained weights, and start training with a low learning rate to avoid damaging existing features. Gradually adjust the learning rate based on the training curve, and increase it later to optimize new layer parameters. Choose a cross-entropy loss function suitable for classification tasks.
[0075]
[0076] Where M is the number of categories, y ic The sign function is 1 if the true class of sample i equals c, and 0 otherwise. This is used in conjunction with the Adam optimizer.
[0077] m t =β1m t-1 +(1-β1)g t
[0078]
[0079] To address the class imbalance problem of the target species, weight adjustment can be introduced. Accuracy, loss, and other metrics are monitored on the validation set, and an early stopping rule is set to prevent overfitting and shorten the training cycle. During the validation set monitoring phase, in addition to the regular accuracy, the macro-average F1 score is tracked to more comprehensively evaluate the model's ability to identify a minority of classes. The early stopping rule employs a dual-judgment mechanism: if the validation loss does not decrease for 10 consecutive epochs, or the macro-average F1 score fluctuates within a range of less than 0.5% for 5 consecutive epochs, early stopping is triggered. After training terminates, the model weights that best perform on the validation set are automatically rolled back, and a training curve report is generated, including the loss value decay trend and heatmap, for subsequent optimization reference. Furthermore, training cycle compression is achieved through dynamic batch adjustment: a small batch size (batch_size = 16) is used initially to ensure stability, and the batch size is gradually increased in the mid-to-late stages to accelerate convergence.
[0080] The transfer learning and model fine-tuning steps in the model construction are as follows:
[0081] The front-end platform calls the API via a POST request to upload image data. An asynchronous processing mechanism improves throughput efficiency in high-concurrency scenarios. The model performs object detection and fine-grained classification on the preprocessed images, generating species labels and confidence scores. The returned data is encapsulated in JSON format. The front-end platform receives the prediction results in JSON format, parses them, and dynamically renders the species information.
[0082] The above description is merely a preferred embodiment of the present invention and is used only to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A species detection and identification method based on a hybrid scaling strategy and a cascaded architecture, characterized in that, include 1) Species image acquisition: The equipment is deployed in the target area and uses thermal sensing triggering and motion detection algorithms to control the shooting, prioritizing the capture of live targets to reduce invalid images; the acquired raw images are sent back to the local data center; 2) Image preprocessing: For the acquired dataset, the image size is first adjusted by filling the edge areas to make the image proportions uniform; then the image is normalized to balance the pixel value distribution and the format conversion is completed. 3) Model building: The technical solution for species detection and fine-grained recognition is based on a convolutional network model architecture. First, a baseline network is built, and the network structure is optimized through scaling strategies to adapt it to data inputs with different features. Then, the target detection module and the fine-grained classification module are cascaded and integrated to form a complete detection and recognition model, realizing integrated processing from target detection to accurate classification; 4) Transfer learning and model fine-tuning: Based on the pre-trained model, extract its learned general features; for the target species identification task, select key layers for parameter adjustment, and optimize model specificity by injecting labeled data; 5) API service deployment and platform integration: The front-end platform calls the interface via POST requests to upload image data. An asynchronous processing mechanism is used to improve throughput efficiency in high-concurrency scenarios. The model performs object detection and fine-grained classification on the preprocessed images to generate species labels and confidence results. The returned data is encapsulated in JSON format. The front-end platform receives the prediction results in JSON format, parses them, and dynamically renders the species information.
2. The method according to claim 1, characterized in that, When filling edge areas, a symmetrical filling strategy is adopted. After the long side of the image is scaled to the target size, the remaining part of the short side is filled with gray to avoid introducing noise unrelated to the background. The normalization process employs channel-wise mean-variance standardization to adapt to the input distribution of the pre-trained model; During the format conversion stage, the image tensor is adjusted to CHW format to be compatible with GPU-accelerated computing.
3. The method according to claim 1, characterized in that, Infrared cameras and video monitoring equipment are used to acquire images of species in the wild; image data is transmitted over the network via 4G network and wireless bridge to build a dataset.
4. The method according to claim 3, characterized in that, A species detection and fine-grained recognition model is constructed based on a neural network. The MBConv structure is used to build the baseline network, and the network is improved on the baseline network through a hybrid scaling strategy. The detector network and the classifier network are cascaded to form a detection and recognition model, which performs fine-grained classification on the detected targets.
5. The method according to claim 4, characterized in that, Model fine-tuning and optimization based on collected data: The collected data is divided into training set, validation set and test set in a ratio of 7:2:
1. The Adam optimization algorithm is used, the initial learning rate is set, the batch size and number of iterations are adjusted to carry out transfer learning of the model, and the validation set is used for real-time evaluation. The optimal model parameters are saved, and new data is collected periodically to update the model parameters.
6. The method according to claim 2, characterized in that, Let the scaling factor r be the value used in image preprocessing. Scale the image proportionally to the specified size. Calculate the fill amount and adjust by increments: dw=(old_width-new_width)%stride dh=(old_height-new_height)%stride Center the image and fill it with color symmetrically on the top and bottom / left and right sides; normalize the scaled and filled image by using channel mean-variance standardization to normalize the pixel values from [0,255] to [0,1] to adapt to the input distribution of the pre-trained model; finally, adjust the channel order and format, converting the channel order to the RGB format required by the model and the tensor to CHW format.
7. The method according to claim 1, characterized in that, The construction steps in the model construction are as follows: The MBConv structure is constructed by sequentially using a 1*1 ordinary convolution, a 3*3 or 5*5 DepthwiseConv convolution, an SE module, a 1*1 ordinary convolution, and a Dropout layer. The SE module improves the network's representational ability by modeling the interdependencies between convolutional feature channels through three steps: global information embedding, adaptive recalibration, and reweighting. A baseline network can be constructed by stacking MBConv structures and combining them with 1*1 convolutional layers, average pooling layers, and fully connected layers. Based on the baseline network, a hybrid scaling strategy is used to extend the network from different dimensions simultaneously; a hybrid factor is used. The parameters for uniform scaling—depth, width, and resolution—are calculated using the following formulas: depth:d=α φ width:w=β φ resolution:r=γ φ stα·b 2 ·c 2 ≈2 α≥1,β≥1,γ≥1 Where st represents the constraints, and α, β, and γ are the parameters corresponding to depth, width, and resolution, respectively; these are first set during scaling. Set the parameter to 1, and search for the optimal α, β, and γ parameters in the baseline network. After obtaining the optimal parameters, keep the parameters unchanged and adjust α, β, and γ to obtain the scaled network framework. Combine the baseline network with the scaled network to obtain the network structure for feature extraction. During scaling, depth expansion is preferentially applied to shallow networks to preserve detailed features, while width expansion is concentrated in deep networks to enhance semantic expressiveness. Resolution enhancement is increased synchronously with network layers to ensure effective fusion of multi-scale features. Dropout and Dense layers are added to the network structure to classify the features extracted by the network, resulting in a complete model architecture.
8. The method according to claim 1, characterized in that, The transfer learning and model fine-tuning steps in the model construction are as follows: The pre-trained model should retain the structure of the backbone convolutional layers and freeze their parameters to avoid destroying the learned basic features during fine-tuning. The top fully connected layer of the pre-trained model is designed for the original dataset and needs to be deleted and replaced with a new classification layer; the number of neurons in the new layer is consistent with the number of target species categories, and initial parameters are assigned through random initialization or Xavier initialization strategy; Initialize model parameters with pre-trained weights, and start training with a low learning rate to avoid damaging existing features; gradually adjust the learning rate according to the training curve, and increase the learning rate later to optimize the parameters of new layers; select the cross-entropy loss function suitable for classification tasks. Where M is the number of categories, y ic The sign function is set to 1 if the true class of sample i equals c, and 0 otherwise; it is used in conjunction with the Adam optimizer. m t =β1m t-1 +(1-β1)g t To address the class imbalance problem of the target species, weight adjustment can be introduced; on the validation set, metrics such as accuracy and loss value can be monitored, and early stopping rules can be set to prevent overfitting and shorten the training cycle. During the validation set monitoring phase, in addition to the regular accuracy, we added tracking of the macro-average F1 score (Macro-F1) to more comprehensively evaluate the model's ability to identify a minority of classes. The early stop rule employs a dual judgment mechanism: if the verification loss does not decrease for 10 consecutive epochs, or the macro average F1 score fluctuates within a range of less than 0.5% for 5 consecutive epochs, then early stop is triggered. After training terminates, the system automatically rolls back to the model weights that best perform on the validation set and generates a training curve report, including the loss value decay trend and heatmap, for subsequent optimization reference.
9. The method according to claim 8, characterized in that, Training cycle compression is achieved through dynamic batch adjustment: initially, a small batch size of 16 is used to ensure stability, and then gradually increased in the middle and later stages to accelerate convergence.
10. The method according to claim 1, characterized in that, The transfer learning and model fine-tuning steps in the model construction are as follows: The front-end platform calls the interface via POST requests to upload image data. It adopts an asynchronous processing mechanism to improve throughput efficiency in high-concurrency scenarios. The model performs object detection and fine-grained classification on the preprocessed image, generating species labels and confidence results. The returned data is encapsulated in JSON format. The front-end platform receives the prediction results in JSON format, parses them, and dynamically renders the species information.
Citation Information
Patent Citations
Space target image intelligent identification method
CN120125897A
Cited By
Channel visibility time-phased regression prediction method and system based on ResNet transfer learning
CN121686123A