Unsupervised domain migration method and system for micro target detection in agricultural scene

By generating high-fidelity synthetic datasets and lightweight networks, and combining a dual-branch structure with domain adversarial training, the problems of insufficient data requirements and generalization ability for small target detection in agricultural scenarios are solved, enabling real-time detection of small targets and edge inference in agricultural fields.

CN122223552APending Publication Date: 2026-06-16FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610443603.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-07
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

In agricultural scenarios, the detection of small targets cannot meet the needs of model training for large-scale and diverse data, resulting in insufficient model training and inadequate generalization ability.

Method used

A high-fidelity synthetic dataset is generated through parametric modeling, a lightweight detection network is constructed, a dual-branch structure and domain adversarial training mechanism are adopted, time-frequency joint features are fused, unsupervised domain adaptive training is performed, and frequency domain fine-tuning is carried out to optimize the detection performance of small target detail features. Finally, it is deployed to edge devices for real-time inference.

Benefits of technology

It enables zero-label detection and real-time edge inference of tiny targets in agricultural scenarios, improves the model's domain-independent feature extraction capability and target domain detail adaptation capability, reduces accuracy decay, and meets the real-time monitoring needs of agricultural sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223552A_ABST
    Figure CN122223552A_ABST
Patent Text Reader

Abstract

The application relates to the field of agricultural intelligent detection, and discloses an unsupervised domain migration method and system for micro target detection in an agricultural scene. The unsupervised domain migration method for micro target detection in the agricultural scene comprises the following steps: generating a high-fidelity synthetic data set as source domain data based on parameterized modeling, inputting the source domain data into a light-weight detection network for pre-training, constructing a double-branch structure and a domain adversarial training mechanism in the initial model, and forming a training model; inputting an unlabeled real field image for frequency domain fine-tuning, optimizing the micro target detail feature detection performance, and forming a fine-tuning model; and deploying the fine-tuning model to an edge device for real-time inference. The high-fidelity synthetic data set containing automatic labeling is generated through parameterized modeling, and then the light-weight network is pre-trained, the double-branch structure and the domain adversarial training are fused, and the time-frequency joint features are adaptively trained in an unsupervised domain, so that zero-label detection of micro targets in the agricultural scene and real-time inference of the edge device are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural intelligent detection technology, specifically to an unsupervised domain transfer method and system for detecting small targets in agricultural scenarios. Background Technology

[0002] With the rapid development of smart agriculture, automated monitoring of crop growth status has become a core support for improving agricultural production efficiency and the level of refined management. Among them, the early detection and counting of tiny targets such as flower buds and flower clusters are key prerequisites for determining the crop growth stage, regulating the flowering and fruiting period, and predicting yield, directly affecting the scientific nature and timeliness of subsequent planting decisions.

[0003] In existing technologies, the detection of small targets in agricultural scenarios mainly relies on supervised training schemes based on real field data. This involves manually taking a large number of high-resolution images of the field under different seasons, lighting, and weather conditions, manually labeling the location and category information of small targets, and then using general object detection networks such as YOLO and Faster R-CNN for model training before finally deploying them for actual detection. However, agricultural targets are small in size and high in density. Manually labeling a single flower bud in a 4K high-resolution image takes a lot of time. Furthermore, the collection of real field images is difficult and the sample coverage is limited due to natural factors such as season, region, and weather. This makes it difficult to meet the needs of model training for large-scale and diverse data, resulting in insufficient model training and inadequate generalization ability. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides an unsupervised domain transfer method and system for small target detection in agricultural scenarios. This solves the problem that existing technologies for small target detection in agricultural scenarios cannot meet the needs of model training for large-scale and diverse data, resulting in insufficient model training and inadequate generalization ability.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an unsupervised domain transfer method for detecting small targets in agricultural scenarios, comprising the following steps: A high-fidelity synthetic dataset is generated based on parametric modeling as the source domain data, and the synthetic dataset contains annotation information of tiny targets; Using a lightweight detection network as the backbone, the source domain data is input for pre-training to construct an initial model with basic detection capabilities for small targets; In the initial model, a dual-branch structure and a domain adversarial training mechanism are constructed, and time-frequency joint features are integrated. The source domain data and unlabeled real field images are input for unsupervised domain adaptation training to form a training model. By fixing the network parameters of the trained model and inputting unlabeled real field images for frequency domain fine-tuning, the performance of detecting small target detail features is optimized to form a fine-tuned model. The fine-tuned model is lightweighted and adapted for interface compatibility, and then deployed to edge devices for real-time inference.

[0006] By adopting the above technical solution, a high-fidelity synthetic dataset with automatic annotation is generated through parametric modeling, which meets the model training requirements for large-scale and diverse data. Then, through lightweight network pre-training, unsupervised domain adaptation training that integrates time-frequency joint features with dual-branch structure and domain adversarial training, and frequency domain fine-tuning to optimize the detail detection performance, the model is endowed with the ability to extract domain-independent features and adapt to target domain details. This enables zero-label detection of small targets in agricultural scenarios and real-time inference at the edge, solving the problem in existing technologies that the detection of small targets in agricultural scenarios cannot meet the requirements of large-scale and diverse data for model training, resulting in insufficient model training and inadequate generalization ability.

[0007] Preferably, the step of generating a high-fidelity synthetic dataset based on parametric modeling as source domain data specifically includes the following steps: A three-level topology of trunk-lateral branches-organs is constructed, branches are instantiated using Bézier curves, and size changes are controlled by exponential decay. Based on the aforementioned three-level topology, a three-dimensional composite form of the micro-target is designed, and mesh models of flower buds and leaves are constructed and refined in detail. The target distribution is optimized by using random sampling and noise perturbation strategies at the target locations corresponding to the grid model, so that the target arrangement conforms to the natural growth law; The camera parameters are adaptively adjusted using a Bayesian optimization algorithm, and a batch saving strategy is adopted. After each preset number of frames are rendered, the file is saved and the dependent graph resources are released. At the same time, the peak value of the video memory is controlled within a preset range. High-resolution, high-fidelity synthetic datasets with automatic annotation information are generated in batches as source domain data. The camera parameters include position, angle and focal length.

[0008] Preferably, the construction of the three-level topology of trunk-lateral branches-organs specifically includes the following steps: A dictionary data structure is used to store the hierarchical relationships between root nodes, trunks, lateral branches, and organs; Each branch is instantiated using a Bézier curve containing a preset number of control points, and the branch radius and length are calculated using an exponential decay method.

[0009] Preferably, the step of constructing mesh models of flower buds and leaves respectively and performing detail refinement specifically includes the following steps: Using a UV ellipsoid as the base mesh, insert edge rings and extrude them along the normal direction to form longitudinal edges. Select the bottom ring surface to perform extrusion and rotation transformation to form sepals. Optimize the surface specular transition through the automatic smoothing function to construct a flower bud mesh model. Using a subdivided plane as the base mesh, the central rib structure is pulled out by scaling, bending deformation is performed to simulate natural shape, thickness is added and edge characteristics are optimized to build the blade mesh model.

[0010] Preferably, the optimization of the target distribution using a random sampling and noise perturbation strategy specifically includes the following steps: The initial position of the target is determined by a golden angle sampling strategy, the polar angle is calculated at fixed angle intervals, and the radius is allocated according to the square root ratio. Gaussian noise is added to the polar angle and radius of the target location, and the noise standard deviation decreases exponentially with the branch level. By coordinating the control of sampling parameters and noise parameters, the target distribution exhibits both orderliness and local randomness, and the target spirals are densely packed without overlap.

[0011] Preferably, the construction of the initial model with basic detection capabilities for small targets specifically includes the following steps: A lightweight YOLO series network was selected as the backbone network, with only the detection branch enabled. Set the preset iteration rounds, learning rate, and optimizer type, and input labeled source domain data for training; The model is trained until its detection accuracy meets a preset threshold, forming an initial model with basic detection capabilities for small targets.

[0012] Preferably, the formation of the training model specifically includes the following steps: After dimensionality reduction of the backbone network features of the initial model, a detection branch and a domain discrimination branch are constructed. The detection branch outputs the target detection result, and the domain discrimination branch outputs the feature domain attribution discrimination result. A gradient inversion layer is inserted into both the forward and backward paths of the domain discrimination branch, and a negative coefficient is applied to the gradient during backpropagation. A two-dimensional fast Fourier transform is performed on the time-domain feature map output by the backbone network to extract the frequency-domain amplitude spectrum and concatenate it with the time-domain features to form a time-frequency joint feature. Training is guided by a joint loss function, and the weight coefficients are dynamically adjusted during training to form a training model. The joint loss function is a weighted sum of detection loss and domain adversarial loss.

[0013] Preferably, the formation of the fine-tuning model specifically includes the following steps: The parameters of the first half of the backbone network in the fixed training model are frozen, and only the parameters of the second half of the backbone network and the domain discriminant branch are unfrozen. Set the fine-tuning iteration rounds and learning rate, and only input unlabeled real field images; The unfrozen parameters of the trained model are iteratively fine-tuned until the model output converges, forming a fine-tuned model.

[0014] Preferably, the deployment to edge devices for real-time inference specifically includes the following steps: The total number of parameters and computational cost of the fine-tuning model are controlled within a preset range, and a command-line interface is designed to make the fine-tuning model compatible with the native interface of the detection framework. Set the input image size and inference accuracy for edge devices, and control the device's operating power consumption; The adapted, fine-tuned model is deployed to edge devices for small target detection.

[0015] An unsupervised domain transfer system for small target detection in agricultural scenarios, applied to the aforementioned unsupervised domain transfer method for small target detection in agricultural scenarios, includes: Generation module: used to generate a high-fidelity synthetic dataset as source domain data based on parametric modeling, wherein the synthetic dataset contains annotation information of small targets; Pre-training module: Used to pre-train the source domain data with a lightweight detection network as the backbone to build an initial model with basic detection capabilities for small targets; Domain adaptation training module: used to construct a dual-branch structure and domain adversarial training mechanism in the initial model, integrate time-frequency joint features, input the source domain data and unlabeled real field images to perform unsupervised domain adaptation training, and form a training model; Frequency domain fine-tuning module: used to fix the network parameters of the trained model, input unlabeled real field images for frequency domain fine-tuning, optimize the detection performance of small target detail features, and form a fine-tuned model; Deployment module: Used to perform lightweight adaptation and interface compatibility processing on fine-tuned models, and deploy them to edge devices for real-time inference.

[0016] This invention provides an unsupervised domain transfer method and system for detecting small targets in agricultural scenarios. It offers the following advantages: 1. This invention generates a high-fidelity synthetic dataset with automatic annotations through parametric modeling, meeting the model training requirements for large-scale and diverse data. Then, through lightweight network pre-training, unsupervised domain adaptation training that integrates time-frequency joint features with dual-branch structure and domain adversarial training, and frequency domain fine-tuning to optimize detail detection performance, the model is equipped with domain-independent feature extraction capability and target domain detail adaptation capability. This enables zero-label detection of small targets in agricultural scenarios and real-time edge inference, solving the problem in existing technologies where the detection of small targets in agricultural scenarios cannot meet the model training requirements for large-scale and diverse data, resulting in insufficient model training and inadequate generalization ability.

[0017] 2. This invention constructs a dual-branch structure and a domain adversarial training mechanism, combined with a gradient reversal layer to guide the model to learn domain-independent features, and integrates time-frequency joint features to retain minute target details. This enables the model to adapt to the feature differences between the source and target domains with the assistance of unlabeled real field images, reducing the accuracy decay when the model is deployed to real scenes and improving the model's generalization ability.

[0018] 3. This invention builds a model based on a lightweight YOLO series network. The addition of domain adaptation and frequency domain feature modules only brings a small increase in the number of parameters and does not increase the amount of computation. At the same time, through targeted lightweight adaptation processing, the total number of model parameters and computation are controlled, enabling the model to be adapted to edge devices such as Jetson. While ensuring detection accuracy, it achieves real-time inference and low-power operation, meeting the deployment needs of mobile monitoring and real-time feedback in agricultural fields. Attached Figure Description

[0019] Figure 1 This is a flowchart of the unsupervised domain migration method for small target detection in agricultural scenarios proposed in this invention; Figure 2 This is an architecture diagram of the unsupervised domain migration system for small target detection in agricultural scenarios proposed in this invention. Detailed Implementation

[0020] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Example 1: In a first embodiment of the present invention, the present invention provides an unsupervised domain transfer method for detecting small targets in agricultural scenarios, such as... Figure 1 As shown, it includes the following steps: A high-fidelity synthetic dataset is generated based on parametric modeling as the source domain data. The synthetic dataset contains annotation information of tiny targets. Furthermore, a high-fidelity synthetic dataset is generated based on parametric modeling as the source domain data, specifically including the following steps: A three-level topology of trunk-lateral branches-organs is constructed, branches are instantiated using Bézier curves, and size changes are controlled by exponential decay. Based on a three-level topology, a three-dimensional composite form of a micro-target is designed, and mesh models of flower buds and leaves are constructed and refined in detail. For the target locations corresponding to the grid model, a random sampling and noise perturbation strategy is used to optimize the target distribution so that the target arrangement conforms to the natural growth law; The camera parameters are adaptively adjusted using a Bayesian optimization algorithm, and a batch saving strategy is adopted. After rendering a preset number of frames, the file is saved and the dependent graph resources are released. At the same time, the peak memory usage is controlled within a preset range to generate high-resolution, high-fidelity synthetic datasets with automatic annotation information in batches as source domain data. The camera parameters include position, angle and focal length.

[0022] Furthermore, a three-tiered topological structure of trunk-lateral branches-organs is constructed, specifically including the following steps: A dictionary data structure is used to store the hierarchical relationships between root nodes, trunks, lateral branches, and organs; Each branch is instantiated using a Bézier curve containing a preset number of control points, and the branch radius and length are calculated using an exponential decay method.

[0023] Furthermore, separate mesh models of the flower buds and leaves are constructed and detailed, specifically including the following steps: Using a UV ellipsoid as the base mesh, insert edge rings and extrude them along the normal direction to form longitudinal edges. Select the bottom ring surface to perform extrusion and rotation transformation to form sepals. Optimize the surface specular transition through the automatic smoothing function to construct a flower bud mesh model. Using a subdivided plane as the base mesh, the central rib structure is pulled out by scaling, bending deformation is performed to simulate natural shape, thickness is added and edge characteristics are optimized to build the blade mesh model.

[0024] Furthermore, a random sampling and noise perturbation strategy is employed to optimize the target distribution, specifically including the following steps: The initial position of the target is determined by a golden angle sampling strategy, the polar angle is calculated at fixed angle intervals, and the radius is allocated according to the square root ratio. Gaussian noise is added to the polar angle and radius of the target location, and the noise standard deviation decreases exponentially with the branch level. By coordinating the control of sampling parameters and noise parameters, the target distribution exhibits both orderliness and local randomness, and the target spirals are densely packed without overlap.

[0025] Specifically, the core of generating high-fidelity synthetic datasets based on parametric modeling is to generate source domain data that is highly consistent with real field micro-targets in terms of structure, morphology, and distribution, and is accompanied by automatic annotations, through structured hierarchical modeling, natural morphology construction, ordered distribution optimization, and efficient batch rendering.

[0026] Generally, constructing a three-level topological structure of trunk-lateral branches-organs is the foundation for simulating the hierarchical relationships of natural plant growth. Its purpose is to ensure that the growth structure of synthesized micro-targets conforms to the physiological characteristics of crops. A dictionary data structure is used to store the hierarchical relationships between root nodes, trunk, lateral branches, and organs. This data structure clearly traces the subordinate logic of each level, facilitating flexible adjustment and batch modification of subsequent morphological parameters. Each branch is instantiated using a Bézier curve with four control points. These four control points precisely control the curvature and growth direction of the branch, adapting to the differences in branch growth morphology among different crops. The branch radius and length are calculated using an exponential decay method to simulate the natural size decay law of plant branches from the trunk to the lateral branches. The formula is: ; ; in For the first The radius of the hierarchical branch, The initial radius of the main trunk, It is the radius attenuation coefficient and its value ranges from (0,1). For the first The length of the hierarchical branch, The initial length of the main stem. This is the length attenuation coefficient, and its value ranges from (0,1). In some embodiments, the initial radius of the main trunk is input. Initial length and preset attenuation coefficient and The radius and length parameters of each level of branches are calculated using the above formula, and the size data of each level of branches are output, so as to accurately replicate the natural growth and decay law of plant branches and ensure the rationality of the topological structure.

[0027] This design utilizes a three-dimensional composite form based on a three-level topological structure to closely approximate the morphological details of flower buds and leaves. When constructing the flower bud mesh model, a UV ellipsoid is used as the base mesh. This base mesh possesses uniform surface distribution characteristics, providing a stable foundation for subsequent detail refinement. Edge rings are inserted at specific locations along the latitudes of the UV ellipsoid, and longitudinal ridges are formed by extrusion along the normal direction. The longitudinal ridge structure simulates the natural texture and grooves of the flower bud surface. Specific faces of the ellipsoid's bottom ring are selected for extrusion and rotation transformation to form the sepal structure. By adjusting the extrusion length and rotation angle, the sepals are made to appear naturally extended. The surface highlight transition is optimized using an automatic smoothing function, ensuring that the highlight distribution on the flower bud surface conforms to realistic light and shadow patterns, thus constructing a lifelike flower bud mesh model.

[0028] When constructing the leaf mesh model, a subdivided plane is used as the base mesh. Subdivision improves the smoothness and morphological flexibility of the leaf surface. The central rib structure is created using the scaling function. As the supporting structure of the leaf, the shape of the central rib directly determines the natural growth posture of the leaf. Bending deformation is performed along a specific axis to give the leaf a natural crescent shape during plant growth. Thickness parameters are added and edge characteristics are optimized. Thickness parameters ensure the leaf has a three-dimensional texture, while edge optimization avoids jagged or blurry edges, creating a mesh model that is highly consistent with the structure and morphology of a real leaf.

[0029] A random sampling and noise perturbation strategy is adopted for the target positions corresponding to the grid model to simulate the natural growth distribution characteristics of microscopic plant targets. Specifically, a golden angle sampling strategy is used to determine the initial position of the target. This strategy can achieve a spiral dense distribution of the target on the branches, as shown in the formula: ; ; in For the first The polar angle of the target. For sampling sequence number, For the first The radius distance from the target to the root of the branch The maximum radius of the target distribution on a single branch. This represents the total number of targets on a single branch. With the maximum distribution radius The initial polar angle and radius coordinates of each target are calculated using the above formula, and the initial position data of the target are output, so as to achieve a non-overlapping spiral dense distribution of targets on the branches, which conforms to the growth distribution law of real plant targets.

[0030] Gaussian noise is added to the polar angle and radius of the target location to simulate the randomness of natural growth, as shown in the formula: ;

[0031] in The target polar angle after adding noise. It is polar Gaussian noise and follows a normal distribution with a mean of 0 and a standard deviation of 5°. The distance to the target radius after adding noise. The noise is radius Gaussian noise and follows a normal distribution with a mean of 0 and a standard deviation of 0.3 mm. The noise standard deviation decays exponentially with increasing branch level, as shown in the formula: ,in For the first The noise standard deviation of hierarchical branches, The initial noise standard deviation, For branch levels. Input initial noise standard deviation. This formula dynamically adjusts the noise intensity of different levels of branches, making the target noise on higher-level side branches smaller, maintaining the overall orderliness of the distribution, while achieving natural randomness through local noise. Finally, through the coordinated control of sampling parameters and noise parameters, the target distribution presents a unity of orderliness and local randomness, and ensures that the target spiral tessellation is non-overlapping.

[0032] By adaptively adjusting camera parameters using a Bayesian optimization algorithm and employing a batch saving strategy, the Bayesian optimization algorithm iteratively optimizes the camera's position, angle, and focal length parameters based on the quality feedback of rendered images. This ensures that small targets are rendered in a proportionate manner with clear and discernible details. The batch saving strategy saves files and releases dependent graph resources immediately after rendering a preset number of frames. This strategy effectively reduces peak GPU memory usage and avoids rendering interruptions due to insufficient memory. By controlling peak GPU memory usage within a preset range, continuous and stable batch rendering on a single device is ensured. The final result is a high-resolution, high-fidelity synthetic dataset with automatically labeled information. The labeled information in this dataset includes the bounding box coordinates and category information of small targets, which can be directly used as source domain data for subsequent model training without manual intervention, reducing data preparation costs.

[0033] Using a lightweight detection network as the backbone, the model is pre-trained with input source domain data to build an initial model with basic detection capabilities for small targets. Furthermore, an initial model with basic detection capabilities for small targets is constructed, specifically including the following steps: A lightweight YOLO series network was selected as the backbone network, with only the detection branch enabled. Set the preset iteration rounds, learning rate, and optimizer type, and input labeled source domain data for training; The model is trained until its detection accuracy meets a preset threshold, forming an initial model with basic detection capabilities for small targets.

[0034] Specifically, an initial model with basic detection capabilities for small targets is constructed. Relying on a lightweight backbone network and high-fidelity source domain data, and through scientific training parameter configuration and accuracy control, the model first masters the basic morphological features and detection logic of small targets, providing a stable basic model for subsequent unsupervised adaptive training, while also taking into account the lightweight requirements for edge deployment.

[0035] Choosing a lightweight YOLO series network as the backbone is a key decision considering both detection accuracy and deployment efficiency. Lightweight YOLO series networks inherently possess the characteristics of fewer parameters and higher computational efficiency; for example, YOLOv8n can adapt to the real-time inference needs of subsequent edge devices. Furthermore, its native detection head design has good adaptability for small target detection, allowing it to be used for agricultural micro-target detection without significant additional modifications. Enabling only the detection branch of the backbone network and disabling other redundant functional modules concentrates computational resources on the target detection task, reducing irrelevant computational consumption and improving pre-training efficiency. Selecting YOLOv8n as the backbone network, its feature extraction network uses a C2f module, which can control the number of parameters while ensuring feature representation capabilities. Its detection head adopts a decoupled head design, which can optimize classification and regression tasks separately, making it more conducive to accurate bounding box localization and category determination of micro-targets.

[0036] Setting reasonable training parameters ensures stable model convergence and the formation of basic detection capabilities. The preset number of iteration rounds is determined based on the size of the source domain data. When the source domain data consists of 10K 4K images, the number of iteration rounds is set to 100. The optimizer type can be either a stochastic gradient descent optimizer or an adaptive moment estimation optimizer. The stochastic gradient descent optimizer is more suitable for large-scale data training and has strong convergence stability, while the adaptive moment estimation optimizer has a faster convergence speed. The choice can be flexibly made based on the feature distribution of the source domain data. The learning rate is set using a cosine annealing adjustment strategy, with the formula: ,in For the first Learning rate of the round, The initial maximum learning rate, This is the lower bound of the learning rate. For the current training round, The total number of iterations is given. Input the initial maximum learning rate of 0.01, the lower bound of the learning rate of 0.0001, and the total number of iterations of 100. The learning rate for each iteration is calculated using the formula above, and a dynamically adjusted learning rate sequence is output. This achieves a smooth decrease in the learning rate, enabling rapid parameter updates in the early stages of training and precise fine-tuning in the later stages, avoiding gradient oscillations and improving model convergence stability.

[0037] The training process uses YOLO's native detection loss function to guide parameter updates. The loss function formula is as follows: ,in For the total loss, This is the classification loss, used to optimize the model's accuracy in classifying small object categories. The regression loss is used to improve the localization accuracy of the target bounding box. The confidence loss is used to enhance the model's ability to judge the reliability of detection results. The classification loss is calculated using cross-entropy loss, which can measure the difference between the model's predicted category and the category labeled in the source domain data; the regression loss is calculated using CIoU loss, which considers factors such as the overlap of bounding boxes, the distance between center points, and the aspect ratio, and is more suitable for bounding box regression optimization of small targets; the confidence loss is calculated using binary cross-entropy loss, which is used to distinguish between target regions and background regions.

[0038] During training, the model's detection accuracy on the validation set is monitored in real time. The validation set is divided from the source domain data according to a preset ratio to ensure that the distribution of validation data is consistent with that of training data. Training stops when the model's detection accuracy meets a preset threshold, forming the initial model. This preset threshold is set based on the model having a stable basic ability to detect small targets, ensuring that the model can accurately identify small targets in the source domain data and distinguish targets from the background, laying the foundation for learning domain-independent features in subsequent domain adaptation training. The preset threshold is set to mAP50 ≥ 0.6. When the model's mAP50 on the validation set is consistently above this threshold for several consecutive epochs, the model is considered to have met the training criteria, and iteration stops. For example, after 5 epochs, the model has been able to capture the morphological features and positional information of small targets in the source domain data, possessing the basic conditions for entering subsequent unsupervised domain adaptation training.

[0039] In the initial model, a dual-branch structure and a domain adversarial training mechanism are constructed. Time-frequency joint features are integrated, and unsupervised domain adaptation training is performed by inputting source domain data and unlabeled real field images to form a training model. Furthermore, the training model is developed, which specifically includes the following steps: After dimensionality reduction of the backbone network features of the initial model, a detection branch and a domain discrimination branch are constructed. The detection branch outputs the target detection result, and the domain discrimination branch outputs the feature domain attribution discrimination result. A gradient inversion layer is inserted into both the forward and backward paths of the domain discrimination branch, and a negative coefficient is applied to the gradient during backpropagation. A two-dimensional fast Fourier transform is performed on the time-domain feature map output by the backbone network to extract the frequency-domain amplitude spectrum and concatenate it with the time-domain features to form a time-frequency joint feature. Training is guided by a joint loss function, and the weight coefficients are dynamically adjusted during training to form a training model. The joint loss function is a weighted sum of detection loss and domain adversarial loss.

[0040] Specifically, the training model is formed by constructing a dual-branch structure, introducing a domain adversarial training mechanism, and fusing time-frequency joint features. This allows the initial model to learn common features of the source and target domains with the assistance of unlabeled real field images, alleviating domain drift and preserving the ability to detect details of small targets, thus laying the foundation for domain adaptation for subsequent frequency domain fine-tuning.

[0041] After dimensionality reduction of the backbone network features in the initial model, a dual-branch structure is constructed to enable parallel optimization of the detection and domain adaptation tasks. Specifically, the high-dimensional feature map output by the backbone network needs to be reduced in dimensionality using a 1x1 convolutional kernel. This ensures the compactness of the feature representation while reducing subsequent computational complexity, making the number of feature channels after dimensionality reduction suitable for the parallel computation requirements of the dual-branch structure. The dual branches include a detection branch and a domain discrimination branch. The detection branch uses the YOLO decoupled head structure of the initial model to further optimize the classification of small targets and bounding box regression capabilities, outputting the target detection results. The domain discrimination branch consists of three fully connected layers, which distinguishes whether the input features come from the source domain (high-fidelity synthetic data) or the target domain (unlabeled real field images), outputting the feature domain attribution result. In some embodiments, the fully connected layer activation function of the domain discrimination branch uses LeakyReLU to avoid gradient vanishing, while the Sigmoid function is used in the output layer to map the discrimination result to the 0-1 interval, intuitively reflecting the domain attribution probability of the feature.

[0042] A gradient inversion layer is inserted into both the forward and backward paths of the domain discrimination branch. The gradient inversion layer does not change the feature data during forward propagation; it only applies a negative coefficient to the gradient during backward propagation, forcing the backbone network to learn domain-independent features. Its mathematical expression is: ; ; in The input features are those of the gradient inversion layer. These are the trainable parameters for the backbone network. The negative coefficients are carried by the gradient during backpropagation after the features of the input domain discriminative branch are processed by the gradient reversal layer. Feedback is sent to the backbone network, which, when updating parameters, must both minimize the loss of the detection branch and avoid accurate discrimination of the domain discrimination branch, making it difficult to distinguish between source and target domain features, thereby enabling the extraction of domain-independent features and mitigating domain drift.

[0043] A two-dimensional fast Fourier transform is performed on the time-domain feature map output by the backbone network to extract the frequency-domain amplitude spectrum and concatenate it with the time-domain feature map. This preserves the detailed features of small targets. The time-domain feature map mainly carries the spatial location and overall morphological information of the target, but it is easily affected by background interference during domain adaptation, resulting in smoothing of small target details. The frequency-domain amplitude spectrum, on the other hand, can highlight the high-frequency details of small targets, such as the longitudinal edges of flower buds and the edges of sepals. The fusion of the two can achieve feature complementarity. The formula for the two-dimensional fast Fourier transform is: ; ; in This is the temporal feature map output by the backbone network. , These represent the height and width of the temporal feature map, respectively. The complex frequency domain characteristics are obtained after the two-dimensional Fourier transform. and These represent the real and imaginary parts of a complex number, respectively. The frequency domain amplitude spectrum. Input time domain feature map. The frequency domain amplitude spectrum is calculated using the above formula, and then concatenated with the time domain feature map along the channel dimension to form a time-frequency joint feature. This feature preserves the spatial structure information in the time domain while incorporating detailed features in the frequency domain.

[0044] Training is guided by a joint loss function, the formula of which is: ,in For the total loss, To detect the loss, the original YOLO loss of the initial model is used, which includes classification loss, regression loss and confidence loss, to ensure that the model does not lose its basic ability to detect small targets; For the domain adversarial loss, a binary cross-entropy loss is calculated based on time-frequency joint features to quantify the discrimination error of the domain discrimination branch; These are weighting coefficients used to dynamically adjust the optimization priority between the detection task and the domain adaptation task. They are dynamically adjusted during training. Values: During the initial training phase Setting it to 0.3 prioritizes ensuring that the detection accuracy of small targets does not degrade; during the training phase... Increased to 0.5, strengthening domain adversarial training, and promoting the backbone network to learn domain-independent features; in the later stages of training The value was reduced to 0.2 to fine-tune the balance between detection performance and domain adaptation.

[0045] The training process requires simultaneous input of source domain data and unlabeled real field images. The source domain data provides annotation information for calculating the detection loss, while the unlabeled real field images indirectly guide the backbone network to adapt to the target domain feature distribution through the domain discrimination branch and gradient inversion layer. A batch-alternating input strategy is used during training, with the ratio of source domain data to target domain data set to 1:1 in each batch. This ensures the model is exposed to features from both domains simultaneously, promoting the learning of common features. During training iterations, the convergence trends of the detection loss and domain adversarial loss are monitored in real time. When both losses stabilize without significant oscillations, domain adaptation training is stopped, resulting in a trained model with preliminary domain adaptation capabilities. This model can now initially adapt to the feature distribution of real field scenes while retaining the potential for detecting details of small targets.

[0046] By fixing the network parameters of the training model and inputting unlabeled real field images for frequency domain fine-tuning, the performance of detecting small target detail features is optimized, thus forming a fine-tuned model. Furthermore, a fine-tuning model is developed, which specifically includes the following steps: The parameters of the first half of the backbone network in the fixed training model are frozen, and only the parameters of the second half of the backbone network and the domain discriminant branch are unfrozen. Set the fine-tuning iteration rounds and learning rate, and only input unlabeled real field images; The unfrozen parameters of the trained model are iteratively fine-tuned until the model output converges, forming a fine-tuned model.

[0047] Specifically, the fine-tuning model involves fixing the domain-independent basic features learned in the training model and iteratively fine-tuning only the network parameters related to the adaptation of details in the target domain. This enhances the model's ability to perceive and detect high-frequency detail features of small targets in real field scenarios, ensuring the synergistic optimization of domain adaptation effect and small target detection accuracy.

[0048] The parameters of the first half of the backbone network in the training model are fixed because the first half of the backbone network is mainly responsible for extracting general low-dimensional basic features. These features have already acquired domain independence through domain adaptation training, and fixing them prevents subsequent fine-tuning from destroying this property. Only the parameters of the second half of the backbone network and the domain discrimination branch are unfrozen. The second half of the backbone network is responsible for extracting high-dimensional semantic features and target detail features, while the domain discrimination branch needs to be further adapted to the target domain feature distribution. Joint fine-tuning of the two can accurately align the feature representation of the target domain. The backbone network adopts the C2f module stacking structure of YOLOv8n. The parameters of the first 10 layers of C2f modules are fixed, and all parameters of the last 5 layers of C2f modules, the detection head, and the domain discrimination branch are unfrozen. This preserves the basic domain-independent features while reserving parameter adjustment space for target domain detail adaptation.

[0049] The settings for the number of fine-tuning iterations and the learning rate need to balance convergence efficiency and parameter fine-tuning accuracy, avoiding either excessively large learning rates that lead to the degradation of learned features or excessively small learning rates that render fine-tuning ineffective. The number of fine-tuning iterations is set to 20, and a dynamic decay strategy is used for the learning rate, with the following formula: ,in For the first Fine-tuning the learning rate in the round. To initially fine-tune the learning rate, The attenuation coefficient is... This is the current fine-tuning round. Input the initial fine-tuning learning rate of 0.0001 and the decay coefficient of 0.95. Calculate the learning rate for each round using the above formula, and output the learning rate sequence that decays round by round. This allows for rapid adaptation to the target domain features in the early stages of fine-tuning, and precise fine-tuning of detailed parameters in the later stages, reducing parameter oscillations.

[0050] The fine-tuning phase uses only unlabeled real field images as input, allowing the model to autonomously learn the feature distribution and minute details of the target domain under unlabeled constraints, thus avoiding interference from source domain data in target domain adaptation. The input unlabeled real field images undergo preprocessing consistent with the source domain data to ensure uniform feature input scale, focusing the fine-tuning process on adapting to domain feature differences rather than data format adaptation. Simultaneously, the previously constructed time-frequency joint feature extraction mechanism is still used during fine-tuning; the unfrozen parameters in the latter half of the backbone network are optimized for frequency domain details of the target domain, making the time-frequency joint features more discriminative in the target domain scenario.

[0051] During fine-tuning, the model's detection performance on the target domain validation set needs to be monitored in real time. The convergence of the model's output results is used as the stopping condition for fine-tuning. The core of convergence judgment is the stability of the detection accuracy. The convergence judgment formula is: ,in This represents the difference in detection accuracy between two adjacent rounds. For the first The target domain validation set mAP50 value of the round. For the first The target domain validation set mAP50 value for each round. When three consecutive rounds... When all values ​​are less than a preset threshold, the model output is considered converged, and fine-tuning stops. This convergence criterion ensures that the model has adapted to the minute target detail features of the target domain, resulting in stable detection performance and avoiding overfitting due to excessive fine-tuning.

[0052] The weight coefficients of the domain discrimination branch during the fine-tuning phase remain at the final values ​​of the trained model. Through continuous domain discrimination feedback, it ensures that the fine-tuned features retain domain independence while enhancing the discriminative power of detailed features of small targets. After the above parameter fixing, dynamic learning rate adjustment, input of unlabeled target domain data, and convergence judgment, the unfrozen parameters of the trained model are fine-tuned to form a fine-tuned model. This model retains the domain-independent feature extraction capability obtained from domain adaptation training and improves the detection accuracy of small targets in real field scenarios through fine-tuning of target domain details.

[0053] The fine-tuned model is lightweighted and adapted for interface compatibility, and then deployed to edge devices for real-time inference.

[0054] Furthermore, deployment to edge devices for real-time inference includes the following steps: The total number of parameters and computational cost of the fine-tuning model are controlled within a preset range, and a command-line interface is designed to make the fine-tuning model compatible with the native interface of the detection framework. Set the input image size and inference accuracy for edge devices, and control the device's operating power consumption; The adapted, fine-tuned model is deployed to edge devices for small target detection.

[0055] Specifically, real-time inference deployed to edge devices is achieved through lightweight adaptation and interface compatibility design. This ensures the accuracy of small target detection while adapting the fine-tuned model to the computing power and power consumption constraints of the edge, thus enabling efficient real-time inference.

[0056] Generally, the total number of parameters and computational cost of the fine-tuned model are controlled within preset ranges, with the total number of parameters not exceeding 7.8M and the computational cost maintained at 27.68 GFLOPs to avoid exceeding the computing power limit at the edge. A command-line interface is designed to make the model compatible with the native interfaces of mainstream detection frameworks, allowing deployment without modifying existing scripts.

[0057] For edge devices, the input image size is set to 1600×1600, and FP32 inference accuracy is used to balance accuracy and efficiency. Power consumption control follows the formula: ,in For inference power consumption, This represents the maximum allowable power consumption of edge devices. For model computational cost, This represents the maximum computational load supported by the device. Input =35W, power consumption is constrained by formula to ensure stable operation of the equipment.

[0058] The adapted model is loaded onto the Jetson edge device, the real-time inference process is started, field images are input, and the detection results of small targets are output to meet the needs of real-time monitoring in agricultural fields.

[0059] Example 2: In a second embodiment of the present invention, the present invention provides an unsupervised domain migration system for small target detection in agricultural scenarios, such as... Figure 2 As shown, it includes the following modules: The generation module is used to generate a high-fidelity synthetic dataset as source domain data based on parametric modeling. The synthetic dataset contains annotation information for small targets. Pre-training module: Used to pre-train a lightweight detection network with input source domain data to build an initial model with basic detection capabilities for small targets; Domain Adaptation Training Module: Used to build a dual-branch structure and domain adversarial training mechanism in the initial model, integrate time-frequency joint features, input source domain data and unlabeled real field images for unsupervised domain adaptation training, and form a training model; Frequency domain fine-tuning module: used to fix the network parameters of the training model, input unlabeled real field images for frequency domain fine-tuning, optimize the performance of small target detail feature detection, and form a fine-tuned model; Deployment module: Used to perform lightweight adaptation and interface compatibility processing on fine-tuned models, and deploy them to edge devices for real-time inference.

[0060] A large-scale jasmine plantation needs automated detection of flower buds in the field to provide data support for flowering period regulation and yield prediction. The field environment is complex, with significant light variations, dense background weeds, and small, densely distributed flower buds. Manually annotating a single 4K image takes 5-10 minutes, and sample collection is limited by seasonality and cannot be scaled up. Existing models trained on synthetic data suffer a significant drop in detection accuracy when deployed to edge computing due to domain drift. To address these issues, this invention employs an unsupervised domain transfer system for small target detection in agricultural scenarios, the architecture of which is as follows: Figure 2 As shown. The specific implementation process of this system is as follows: First, the generation module, based on parametric modeling, generates 10K high-fidelity jasmine flower bud synthetic datasets with automatic annotations as source domain data through three-level topology construction, three-dimensional morphology modification, natural distribution optimization, and batch rendering.

[0061] Subsequently, the pre-training module uses YOLOv8n as the backbone network, inputs source domain data for pre-training, and constructs an initial model with basic detection capabilities.

[0062] Next, the domain adaptation training module constructs a dual-branch structure and a gradient inversion layer in the initial model, integrates time-frequency joint features, inputs source domain data and unlabeled field images for domain adaptation training, and forms a training model.

[0063] The frequency domain fine-tuning module fixes some parameters of the training model and fine-tunes them by only inputting unlabeled field images, thereby optimizing the performance of flower bud detail detection and forming a fine-tuned model.

[0064] The final deployment module performs lightweight adaptation of the fine-tuned model, controls the number of parameters and computation, is compatible with mainstream frameworks through the command-line interface, and is deployed to Jetson devices to achieve real-time inference at 30ms / frame and complete field flower bud detection.

[0065] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An unsupervised domain transfer method for small target detection in agricultural scenarios, characterized in that, Includes the following steps: A high-fidelity synthetic dataset is generated based on parametric modeling as the source domain data, and the synthetic dataset contains annotation information of tiny targets; Using a lightweight detection network as the backbone, the source domain data is input for pre-training to construct an initial model with basic detection capabilities for small targets; In the initial model, a dual-branch structure and a domain adversarial training mechanism are constructed, and time-frequency joint features are integrated. The source domain data and unlabeled real field images are input for unsupervised domain adaptation training to form a training model. By fixing the network parameters of the trained model and inputting unlabeled real field images for frequency domain fine-tuning, the performance of detecting small target detail features is optimized to form a fine-tuned model. The fine-tuned model is lightweighted and adapted for interface compatibility, and then deployed to edge devices for real-time inference.

2. The unsupervised domain migration method for small target detection in agricultural scenarios according to claim 1, characterized in that: The process of generating a high-fidelity synthetic dataset based on parametric modeling as the source domain data specifically includes the following steps: A three-level topology of trunk-lateral branches-organs is constructed, branches are instantiated using Bézier curves, and size changes are controlled by exponential decay. Based on the aforementioned three-level topology, a three-dimensional composite form of the micro-target is designed, and mesh models of flower buds and leaves are constructed and refined in detail. The target distribution is optimized by using random sampling and noise perturbation strategies at the target locations corresponding to the grid model, so that the target arrangement conforms to the natural growth law; The camera parameters are adaptively adjusted using a Bayesian optimization algorithm, and a batch saving strategy is adopted. After each preset number of frames are rendered, the file is saved and the dependent graph resources are released. At the same time, the peak value of the video memory is controlled within a preset range. High-resolution, high-fidelity synthetic datasets with automatic annotation information are generated in batches as source domain data. The camera parameters include position, angle and focal length.

3. The unsupervised domain migration method for small target detection in agricultural scenarios according to claim 2, characterized in that: The construction of the three-level topology of trunk-lateral branches-organs specifically includes the following steps: A dictionary data structure is used to store the hierarchical relationships between root nodes, trunks, lateral branches, and organs; Each branch is instantiated using a Bézier curve containing a preset number of control points, and the branch radius and length are calculated using an exponential decay method.

4. The unsupervised domain migration method for small target detection in agricultural scenarios according to claim 2, characterized in that: The process of constructing mesh models of flower buds and leaves and refining their details includes the following steps: Using a UV ellipsoid as the base mesh, insert edge rings and extrude them along the normal direction to form longitudinal edges. Select the bottom ring surface to perform extrusion and rotation transformation to form sepals. Optimize the surface specular transition through the automatic smoothing function to construct a flower bud mesh model. Using a subdivided plane as the base mesh, the central rib structure is pulled out by scaling, bending deformation is performed to simulate natural shape, thickness is added and edge characteristics are optimized to build the blade mesh model.

5. The unsupervised domain migration method for small target detection in agricultural scenarios according to claim 2, characterized in that: The optimization of the target distribution using a random sampling and noise perturbation strategy specifically includes the following steps: The initial position of the target is determined by a golden angle sampling strategy, the polar angle is calculated at fixed angle intervals, and the radius is allocated according to the square root ratio. Gaussian noise is added to the polar angle and radius of the target location, and the noise standard deviation decreases exponentially with the branch level. By coordinating the control of sampling parameters and noise parameters, the target distribution exhibits both orderliness and local randomness, and the target spirals are densely packed without overlap.

6. The unsupervised domain migration method for small target detection in agricultural scenarios according to claim 1, characterized in that: The construction of an initial model with basic detection capabilities for small targets specifically includes the following steps: A lightweight YOLO series network was selected as the backbone network, with only the detection branch enabled. Set the preset iteration rounds, learning rate, and optimizer type, and input labeled source domain data for training; The model is trained until its detection accuracy meets a preset threshold, forming an initial model with basic detection capabilities for small targets.

7. The unsupervised domain migration method for small target detection in agricultural scenarios according to claim 1, characterized in that: The formation of the training model specifically includes the following steps: After dimensionality reduction of the backbone network features of the initial model, a detection branch and a domain discrimination branch are constructed. The detection branch outputs the target detection result, and the domain discrimination branch outputs the feature domain attribution discrimination result. A gradient inversion layer is inserted into both the forward and backward paths of the domain discrimination branch, and a negative coefficient is applied to the gradient during backpropagation. A two-dimensional fast Fourier transform is performed on the time-domain feature map output by the backbone network to extract the frequency-domain amplitude spectrum and concatenate it with the time-domain features to form a time-frequency joint feature. Training is guided by a joint loss function, and the weight coefficients are dynamically adjusted during training to form a training model. The joint loss function is a weighted sum of detection loss and domain adversarial loss.

8. The unsupervised domain migration method for small target detection in agricultural scenarios according to claim 1, characterized in that: The formation of the fine-tuning model specifically includes the following steps: The parameters of the first half of the backbone network in the fixed training model are frozen, and only the parameters of the second half of the backbone network and the domain discriminant branch are unfrozen. Set the fine-tuning iteration rounds and learning rate, and only input unlabeled real field images; The unfrozen parameters of the trained model are iteratively fine-tuned until the model output converges, forming a fine-tuned model.

9. The unsupervised domain migration method for small target detection in agricultural scenarios according to claim 1, characterized in that: The deployment to edge devices for real-time inference specifically includes the following steps: The total number of parameters and computational cost of the fine-tuning model are controlled within a preset range, and a command-line interface is designed to make the fine-tuning model compatible with the native interface of the detection framework. Set the input image size and inference accuracy for edge devices, and control the device's operating power consumption; The adapted, fine-tuned model is deployed to edge devices for small target detection.

10. An unsupervised domain migration system for small target detection in agricultural scenarios, characterized by: An unsupervised domain transfer method for detecting small targets in an agricultural scenario as described in any one of claims 1-9, comprising: Generation module: used to generate a high-fidelity synthetic dataset as source domain data based on parametric modeling, wherein the synthetic dataset contains annotation information of small targets; Pre-training module: Used to pre-train the source domain data with a lightweight detection network as the backbone to build an initial model with basic detection capabilities for small targets; Domain adaptation training module: used to construct a dual-branch structure and domain adversarial training mechanism in the initial model, integrate time-frequency joint features, input the source domain data and unlabeled real field images to perform unsupervised domain adaptation training, and form a training model; Frequency domain fine-tuning module: used to fix the network parameters of the trained model, input unlabeled real field images for frequency domain fine-tuning, optimize the detection performance of small target detail features, and form a fine-tuned model; Deployment module: Used to perform lightweight adaptation and interface compatibility processing on fine-tuned models, and deploy them to edge devices for real-time inference.