Camellia oleifera fruit recognition optimization system based on attention mechanism

Through the fusion of dynamic occlusion perception and cross-modal spectral characteristics, the problem of weak feature discrimination and small object detection accuracy-speed imbalance under occlusion interference is solved, and efficient and real-time occlusion recognition is achieved, improving the recognition accuracy and robustness.

CN120510355AActive Publication Date: 2025-08-19RES INST OF FOREST RESOURCE INFORMATION TECHN CHINESE ACADEMY OF FORESTRY

Patent Information

Application Number
CN202510545974.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-19
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The prior art faces the problems of weak characteristic discrimination under occlusion interference, inefficient cross-spectrum fusion and small object detection accuracy-speed imbalance, especially in environments with dense branches and leaves and variable light, traditional methods are difficult to effectively identify young fruits.

Method used

The oil tea fruit recognition optimization system based on attention mechanism is adopted, and differentiated processing of the occlusion area and efficient fusion of cross-modal features through dynamic occlusion perception attention module, cross-modal spectral feature collaborative fusion network and small-objective adaptive optimization engine are combined with space-time dual-flow occlusion reasoning, cross-attention gating and dynamic loss allocation to achieve differentiated processing of the occlusion area and efficient fusion of cross-modal features.

Benefits of technology

It significantly improves the feature discrimination ability in dense branches and leaves, enhances the robustness of target recognition in low-contrast scenarios, reduces the risk of mismatch caused by light mutations, accurately captures tiny gradient changes in young fruits, realizes efficient parallel computing and real-time response, and improves the stability and generalization capabilities of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510355A_ABST
    Figure CN120510355A_ABST
Patent Text Reader

Abstract

The invention discloses a camellia oleifera fruit identification optimization system based on an attention mechanism, and relates to the technical field of agricultural computer vision, and the system comprises a dynamic occlusion perception attention module which is connected to an input layer of a cross-modal spectral feature collaborative fusion network through a space-time double-flow occlusion reasoning unit. According to the oil tea fruit recognition optimization system based on the attention mechanism, the feature discrimination capability in an environment with dense branches and leaves is remarkably improved, the target recognition robustness in a low-contrast scene is effectively enhanced, the mismatching risk caused by illumination mutation is reduced, the tiny gradient change of young fruits is accurately captured, semantic and detail features are balanced, and the recognition efficiency is improved. Meanwhile, efficient parallel computing is achieved through embedded hardware deployment, real-time response under a complex scene is ensured, weak links of the model are iteratively strengthened, the stability and generalization ability of the system are overall improved, and a high-precision and low-delay camellia oleifera fruit recognition solution is provided for intelligent agriculture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of agricultural computer vision technology, and in particular to an oil-tea camellia fruit recognition optimization system based on an attention mechanism. Background Art

[0002] Camellia oleifera is an important cash crop in southern my country, and intelligent monitoring and harvesting technologies are crucial to its development. However, the complex growing environment of camellia oleifera, characterized by dense occlusion from branches and leaves, variable lighting conditions (such as backlighting and interlaced shadows), and dynamic changes in fruit morphology. The small size of young fruit and the color of fruit that approaches the same level of maturity as the leaves and branches, pose significant challenges to traditional computer vision algorithms. Existing mainstream methods, designed based on RGB images, suffer from significant drawbacks: Low-contrast scene recognition is insufficient: young fruit and leaves are similar in color, and uneven lighting further weakens the target outline, making it difficult for traditional convolutional networks to extract discriminative features, resulting in high rates of false detection and missed detection. Dense occlusion leads to feature confusion: Under interlaced occlusion from branches and leaves, static attention mechanisms, such as the SE module, cannot dynamically adjust their focus area, are susceptible to background interference, and the discriminability of occluded boundary features is reduced. There is also an imbalance in the accuracy and efficiency of small target detection: young fruit accounts for only 0.1% to 1%, making it difficult for fixed feature pyramid strategies to balance semantic and detail information. Standard loss functions ignore size differences, resulting in insufficient coordinated optimization of localization and classification accuracy. Existing improvements attempt to incorporate multimodal data, such as near-infrared imaging or attention mechanisms. However, these approaches simply concatenate multispectral features, fail to exploit cross-modal complementarity, and lack the ability to dynamically respond to occlusion. Furthermore, small-target optimization methods, such as super-resolution reconstruction, increase computational complexity and struggle to meet real-time field demands. A system integrating dynamic multimodal attention, adaptive loss allocation, and efficient cross-spectral fusion is urgently needed to overcome the technical bottleneck of Camellia oleifera fruit recognition in complex environments. Summary of the Invention

[0003] (1) Technical problems solved

[0004] In response to the shortcomings of the existing technology, the present invention provides a camellia fruit recognition optimization system based on the attention mechanism, which solves the problems of weak feature discriminability under occlusion interference, inefficient cross-spectral fusion, and accuracy-speed imbalance in small target detection.

[0005] (2) Technical solution

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a tea fruit recognition optimization system based on attention mechanism, comprising:

[0007] The dynamic occlusion perception attention module is connected to the input layer of the cross-modal spectral feature collaborative fusion network through the spatiotemporal dual-stream occlusion inference unit. The output of the cross-modal spectral feature collaborative fusion network is connected to the heterogeneous pyramid fusion layer of the small target adaptive optimization engine. The positioning loss calculation unit of the small target adaptive optimization engine is fed back to the competitive loss function module of the collaborative verification framework.

[0008] The dynamic occlusion perception attention module includes:

[0009] Spatial dimension processing unit: Input C3-C5 layer multi-scale feature map, generate occlusion probability heat map through deformable convolution layer, where the convolution kernel offset is dynamically calculated by the local gradient of the feature map;

[0010] Time dimension processing unit: Based on lightweight LSTM, it fuses the current frame heat map and the historical frame optical flow features to output the occlusion trend prediction for the next 3-5 frames;

[0011] The cross-modal spectral feature collaborative fusion network includes a spectral reflectance matching verification module for screening effective spectral features through physical prior constraints, and also includes a cross-attention gating mechanism between the RGB branch and the near-infrared branch, wherein:

[0012] The RGB branch uses the inverted residual structure of MobileNetV3 to extract spatial features;

[0013] The near-infrared branch uses deep separable convolution to extract spectral reflectance features;

[0014] The micro-scale feature enhancement pyramid of the small target adaptive optimization engine inserts a dilated convolution layer with a dilation rate of 3 / 6 / 9 between the C3-C5 layers of the FPN, and dynamically allocates the layer weights through the feature reweighting module;

[0015] The collaborative verification framework connects various modules by sharing the underlying feature extractor and adopts a two-stage adversarial verification process to iteratively optimize the weak areas of model decision making.

[0016] Preferably, the dynamic occlusion-aware attention module has a divide-and-conquer attention strategy, including occlusion area division and differentiation processing:

[0017] The occlusion area division includes: using the occlusion probability heat map to set a high threshold of 0.7 and a low threshold of 0.3, and dividing the feature map into occlusion area, transition area, and visible area;

[0018] Differentiation includes:

[0019] a) The occluded area uses a channel attention module to suppress the background channel response, followed by a spatial attention module that uses a dilated convolution with a dilation rate of 6 for local context aggregation;

[0020] b) A local contrast enhancement module is deployed in the transition region. The parameters of its learnable threshold function are jointly optimized by the regional average gradient and standard deviation, and are connected in parallel with the histogram of oriented gradients feature extractor;

[0021] c) superimposing a non-local neural network in the visible region to establish long-range semantic associations between fruits and branches and leaves;

[0022] During the training phase of the dynamic occlusion-aware attention module, random leaf masks with a coverage of 10% to 60% are injected through an adversarial occlusion simulator. During the inference phase, a dual-threshold voting mechanism is used to retain the overlapping areas predicted by the 0.6 and 0.4 thresholds.

[0023] Preferably, the cross-modal spectral feature collaborative fusion network includes:

[0024] Spectral guidance unit: The spectral features output by the near-infrared branch are compressed by 1×1 convolution to generate a Sigmoid spatial mask, which is then multiplied channel by channel with the feature map of the corresponding level of the RGB branch;

[0025] Spatial constraint unit: The semantic features of the RGB branch are converted into a spatial weight matrix through 3×3 convolution, which performs regional weighting on the deep features of the near-infrared branch;

[0026] Dynamic routing selector: Based on the weighted score of target confidence and Sobel gradient amplitude, the formula is: Score = 0.7 × confidence + 0.3 × Sobel gradient amplitude, it selects the bottom-level super-resolution fusion path or the high-level joint upsampling path;

[0027] In the heterogeneous pyramid fusion layer:

[0028] High-level features, i.e., the C5 layer, use deformable convolution to align spectral-spatial features and then perform bilinear interpolation upsampling;

[0029] The bottom layer features, i.e., the C3 layer, fuses low-resolution NIR and high-resolution RGB features through sub-pixel convolution;

[0030] During training, the single-modal input is shielded with a probability of 50%, and the near-infrared features are constrained to be in the range of 0.65-0.82 through the spectral reflectance matching verification module, among which the reflectance of young fruit is 0.65-0.75, and that of mature fruit is 0.70-0.82, and the union is taken during the expansion period.

[0031] Preferably, the micro-scale feature enhancement pyramid of the small target adaptive optimization engine includes:

[0032] Atrous convolution layer configuration: C3 layer atrous ratio 3, C4 layer atrous ratio 6, C5 layer atrous ratio 9;

[0033] Feature reweighting module: Generates a normalized weight vector based on the target relative area input to the fully connected layer. The calculation formula is:

[0034]

[0035] where s i is the average confidence of the small target in the i-th layer, g i is the Sobel gradient amplitude, Ⅱ is the indicator function, when g i When <0.4, it is 1, otherwise it is 0;

[0036] Gaussian-exponential mixture loss function: Classification loss weight formula:

[0037] L cls =-(1+e -D )ylog(p);

[0038] D=α(1-Area)+βOR+γ(1-Contrast);

[0039] The positioning loss adopts the center point Gaussian weighting, σ is inversely proportional to the target size.

[0040] Preferably, the spectral reflectance matching verification module includes:

[0041] Differentiable constraint unit: performs piecewise linear activation on the 61-channel feature map output by the near-infrared branch (corresponding to the 650-950nm band). When the channel value is within the measured threshold [0.65, 0.82] of the band, the original gradient is retained, and the gradient exceeding the threshold is reset to zero.

[0042] Penalty loss formula:

[0043]

[0044] Inference verification mechanism: The fused features are sampled in a 16×16 sliding window. When the effective reflectance channel ratio of three consecutive windows is less than 60%, the C3 layer super-resolution reconstruction compensation of the RGB branch is triggered.

[0045] Among them, the spectral reflectance matching verification module builds a database based on the ASD FieldSpec4 spectrometer.

[0046] Among them, the calculation method of the effective reflectivity channel ratio is:

[0047]

[0048] Preferably, the sharing-competition mechanism of the collaborative verification framework includes:

[0049] Parameter freezing strategy: The underlying parameters of MobileNetV3 and depthwise separable convolution are fixed in the first three training stages;

[0050] Dynamic weight competitive loss formula:

[0051] L total =αL doaa +βL cmsf +γL stae ;

[0052] The weight adjustment rules are:

[0053] When Δ=(R doaa -R cmsf ) / R all When >0.05, α←α×0.9, β←β×1.1;

[0054] When AP stae When <80%, γ←γ+0.05·(80%-AP stae );

[0055] Gradient norm constraint:

[0056] like Then L doaa Applies a 0.7 attenuation factor.

[0057] Preferably, the two-stage adversarial verification process includes:

[0058] Generation of extreme scenario simulation sets:

[0059] a) Occlusion synthesis: Generate dynamic occlusion patches based on CycleGAN, with a coverage rate of 60% to 80%;

[0060] b) Lighting simulation: Combined with the Phong lighting model, a mixed image of overexposed (brightness > 240) and low-light (brightness < 50) images is generated;

[0061] c) Dense small objects: Generate 50-100 young fruits with an area <0.5% per image through super-resolution reconstruction;

[0062] Error feedback mechanism: Grad-CAM is used to locate feature confusion areas and intensive training is carried out in three stages through curriculum learning:

[0063] In the first phase, CMSF and STAE parameters were frozen, and DOAA optimization was focused on.

[0064] In the second stage, the CMSF is unfrozen and the cross-modal similarity threshold is set to 0.7;

[0065] The third stage jointly optimizes the STAE loss weight and feature pyramid parameters;

[0066] Among them, the positioning loss gradient is back-propagated to the error feedback mechanism of the adversarial verification process to dynamically adjust the optimization weights in the course learning stage.

[0067] Preferably, the physical data driven method of the spectral reflectance matching verification module includes:

[0068] Spectral database construction: An ASD FieldSpec4 spectrometer was used to collect reflectance data in the 650-950nm band at different growth stages, namely young fruit, expansion stage, and maturity stage;

[0069] Back-propagation linkage optimization: When the near-infrared feature in the 780nm channel exceeds the threshold [0.68, 0.79], the spatial attention weight of the corresponding position in the RGB branch is increased by a factor of 1.3;

[0070] Real-time verification exception handling process: After freezing the near-infrared branch parameters, enable the C3 layer hole convolution feature of the RGB branch to perform three sub-pixel convolution reconstructions until the reflectivity effective channel ratio is ≥ 60%.

[0071] Preferably, the hardware deployment of the tea fruit recognition optimization system based on the attention mechanism includes:

[0072] Image acquisition unit: IMX477 RGB sensor and SWIR-640 near-infrared camera synchronized trigger acquisition;

[0073] Processing unit: Jetson Xavier NX embedded platform, configured with a multi-threaded task scheduler:

[0074] a) The first thread runs the dynamic occlusion awareness attention module and is allocated to two CPU cores;

[0075] b) The second thread runs the cross-modal fusion network and allocates 128 CUDA cores;

[0076] c) The third thread runs the small target optimization engine and allocates 64 CUDA cores;

[0077] Output interface: The recognition results are transmitted to the agricultural machinery controller via the CAN bus at a frequency of 100Hz, with a response delay of <200ms.

[0078] Preferably, the implementation of the system in the orchard monitoring scenario includes:

[0079] UAV dynamic acquisition unit: A quadrotor UAV equipped with the image acquisition unit cruises in a serpentine path at a constant altitude, achieving centimeter-level track tracking through an RTK positioning module, and dynamically adapting the image acquisition frequency to the flight speed;

[0080] Real-time processing - feedback control loop:

[0081] a) The processing unit completes single-frame camellia fruit recognition within 150ms and transmits the positioning coordinates to the agricultural machinery via the LoRa wireless module;

[0082] b) when the robotic arm controller generates a picking path based on the fruit coordinates, it simultaneously receives the occlusion trend of the next three frames predicted by the dynamic occlusion perception module and dynamically adjusts the motion trajectory of the end effector;

[0083] Spectral-spatial joint verification mechanism: When the spectral reflectance matching verification module detects reflectance anomalies for five consecutive frames, it triggers the drone to hover and start multi-angle compensation shooting, including:

[0084] The pitch angle is ±15° and 3 sets of images are taken;

[0085] Turn on the near-infrared fill light to recapture features;

[0086] The multi-view features are input into the collaborative verification framework for confidence voting, and only the recognition results predicted by the two sets of models are adopted.

[0087] (3) Beneficial effects

[0088] The present invention provides an attention-based tea fruit recognition optimization system, which has the following beneficial effects:

[0089] This attention-based tea fruit recognition optimization system uses a dynamic divide-and-conquer attention mechanism to enhance fruit edge features during occlusion region division and differentiation processing. It combines a spatiotemporal prediction model to capture dynamic changes in occlusion, significantly improving feature discrimination capabilities in densely populated environments. The cross-modal spectral fusion network uses cross-attention gating and physical prior constraints to deeply fuse RGB texture and near-infrared reflectivity features, effectively enhancing the robustness of target recognition in low-contrast scenes and reducing the risk of mismatching caused by sudden changes in illumination. To address the challenge of small target detection, the system uses multi-scale receptive field enhancement and dynamic loss weight distribution to accurately capture subtle gradient changes in young fruits, balancing semantic and detail features. Embedded hardware deployment enables efficient parallel computing, ensuring real-time response in complex scenarios. The collaborative verification framework combines adversarial optimization with a module competition mechanism to iteratively strengthen weak links in the model, improving the overall stability and generalization capabilities of the system and providing a high-precision, low-latency tea fruit recognition solution for smart agriculture. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 It is a schematic diagram of the overall framework of the present invention;

[0091] Figure 2 This is a control logic timing diagram of the present invention. DETAILED DESCRIPTION

[0092] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0093] See also Figure 1 and Figure 2 The present invention provides a technical solution: a tea fruit recognition optimization system based on an attention mechanism, comprising:

[0094] The dynamic occlusion perception attention module is connected to the input layer of the cross-modal spectral feature collaborative fusion network through the spatiotemporal dual-stream occlusion inference unit. The output of the cross-modal spectral feature collaborative fusion network is connected to the heterogeneous pyramid fusion layer of the small target adaptive optimization engine. The positioning loss calculation unit of the small target adaptive optimization engine is fed back to the competitive loss function module of the collaborative verification framework.

[0095] The dynamic occlusion-aware attention module includes:

[0096] Spatial dimension processing unit: Input C3-C5 layer multi-scale feature map, generate occlusion probability heat map through deformable convolution layer, where the convolution kernel offset is dynamically calculated by the local gradient of the feature map;

[0097] Time dimension processing unit: Based on lightweight LSTM, it fuses the current frame heat map and the historical frame optical flow features to output the occlusion trend prediction for the next 3-5 frames;

[0098] The cross-modal spectral feature collaborative fusion network includes a spectral reflectance matching verification module for screening effective spectral features through physical prior constraints, and also includes a cross-attention gating mechanism between the RGB branch and the near-infrared branch, where:

[0099] The RGB branch uses the inverted residual structure of MobileNetV3 to extract spatial features;

[0100] The near-infrared branch uses deep separable convolution to extract spectral reflectance features;

[0101] The Microscale Feature Enhancement Pyramid (MSE-FPN) of the Small Target Adaptive Optimization Engine inserts dilated convolutional layers with dilation rates of 3 / 6 / 9 between the C3-C5 layers of FPN and dynamically allocates layer weights through a feature reweighting module.

[0102] The collaborative verification framework connects each module by sharing the underlying feature extractor, and adopts a two-stage adversarial verification process to iteratively optimize the weak areas of model decision-making.

[0103] It should be further explained that in the specific implementation process, the dynamic occlusion-aware attention module has a divide-and-conquer attention strategy, including occlusion area division and differentiated processing:

[0104] The occlusion area division includes: inputting the multi-scale feature map of the C3-C5 layer to the deformable convolution layer, calculating the convolution kernel offset through the local gradient, and generating the occlusion probability heat map.

[0105] A double threshold is set, i.e., a high threshold of 0.7 and a low threshold of 0.3, to divide the heat map into: an occlusion zone with a probability greater than or equal to 0.7, i.e., the background is complex and the fruit edge features need to be enhanced; a transition zone with a probability of 0.3 ≤ probability < 0.7, i.e., the occlusion boundary is blurred and the contrast needs to be enhanced; and a visible zone with a probability less than 0.3, i.e., the background is simple and global semantics need to be associated.

[0106] Differentiation includes:

[0107] a) For the occluded area, a channel attention module, namely the SE module, is used to suppress the background channel response. For example, the response of the branch and leaf texture channel is reduced by 30%-50%. Then, a spatial attention module is connected. The spatial attention module uses a dilated convolution with a dilation rate of 6 to perform local context aggregation and extract the edges of the remaining fruit.

[0108] b) A local contrast enhancement module is deployed in the transition region. The parameters of its learnable threshold function are jointly optimized by the regional average gradient and standard deviation, and are connected in parallel with the Histogram of Oriented Gradients (HOG) feature extractor. The optimization algorithm formula of the learnable threshold function is:

[0109] Among them, k is a learnable parameter, is the regional average gradient, σ is the standard deviation, and μ is the mean;

[0110] c) A non-local neural network is superimposed on the visible area to establish long-range semantic associations between fruits and branches and leaves;

[0111] During the training phase of the dynamic occlusion-aware attention module, random branch and leaf masks with a coverage of 10% to 60% are injected through an adversarial occlusion simulator. During the inference phase, a dual-threshold voting mechanism is used to retain the overlapping areas predicted by the 0.6 and 0.4 thresholds.

[0112] It should be further explained that in the dynamic occlusion perception attention module, a multi-dimensional occlusion prediction mechanism is first constructed through a spatiotemporal dual-stream occlusion inference unit. In the spatial dimension, a multi-scale feature map of the C3-C5 layer is used as input, and a deformable convolution layer is used to adaptively sample the occlusion boundary. The convolution kernel offset is dynamically calculated by the local gradient of the feature map, thereby accurately extracting the contour features of the fruit obscured by branches and leaves. At the same time, an occlusion probability heat map is generated through three layers of parallel convolution to quantify the occlusion possibility of each pixel area. In the temporal dimension, for video streams or continuous image sequences, a lightweight LSTM unit is introduced to fuse the spatial occlusion heat map of the current frame with the temporal change features of the historical frames (the motion vectors of adjacent frames are extracted by the optical flow method) to predict the dynamic expansion or contraction trend of the occluded area in the next 3-5 frames, thereby achieving spatiotemporal consistency of occlusion compensation. Based on the occlusion probability heatmap, a dual threshold (high threshold 0.7, low threshold 0.3) is set, dividing the feature map into occluded, transition, and visible regions. Differentiated attention mechanisms are then applied to each region: In the occluded region, a channel-spatial dual-dimensional attention is applied, using a SE module to suppress background-related channel responses. This is then combined with dilated convolutions to aggregate local context and enhance the edge features of the remaining fruit. In the transition region, a local contrast enhancement module is deployed, applying a nonlinear transformation of the feature map using a learnable threshold function and incorporating HOG prior information to enhance edge responses. In the visible region, a non-local neural network is superimposed to capture long-range dependencies and establish semantic associations between the fruit and surrounding branches and leaves. During the training phase, realistic branch and leaf masks (coverage of 10% to 60%) are dynamically injected through adversarial occlusion simulation. During the inference phase, a dual-threshold voting mechanism is used to filter out misclassified regions. This scheme improves the recognition accuracy of occluded regions and reduces the confusion rate of occluded boundary features.

[0113] Through the spatiotemporal dual-stream reasoning and divide-and-conquer strategy of the dynamic occlusion perception attention module, the system achieves accurate recognition in an orchard environment with severe occlusion by branches and leaves. Based on the dual-threshold occlusion probability heat map, the feature map is divided into occlusion area, transition area, and visible area. Channel-spatial attention, local contrast enhancement, and global semantic association strategies are used respectively to effectively solve the problem of occlusion boundary feature confusion. Combined with LSTM time series prediction and adversarial occlusion simulation training, the system's adaptability to dynamic occlusion is significantly enhanced, the accuracy of fruit recognition in occluded scenes is improved, and the false detection rate is reduced.

[0114] It is important to note that, in its implementation, the cross-modal spectral feature collaborative fusion network employs a dual-branch heterogeneous feature extraction architecture: the RGB branch extracts fruit texture and shadow distribution features based on the inverted residual structure of MobileNetV3, while the near-infrared (NIR) branch extracts the reflectance features of oil-tea camellia fruit in the NIR band (650-950nm) through depthwise separable convolution. Bidirectional feature guidance is achieved through a cross-attention gating mechanism: the spectral features output by the NIR branch are compressed using a 1×1 convolution to generate a sigmoid spatial mask, which is then multiplied channel-by-channel with the feature map of the corresponding level in the RGB branch to dynamically enhance low-contrast areas, such as backlit fruit outlines. Simultaneously, the semantic features of the RGB branch are converted into a spatial weight matrix through a 3×3 convolution, which regionally weights the deep features of the NIR branch to suppress redundant responses where the reflectance of branches and leaves is similar to that of the fruit. In the heterogeneous pyramid fusion layer, the high-level features at level C5 are aligned through deformable convolutions and then upsampled using bilinear interpolation. The low-resolution NIR and high-resolution RGB features are fused at level C3 through sub-pixel convolution. The dynamic routing algorithm selects the optimal fusion path according to the target confidence and Sobel gradient amplitude, which improves the target contrast and reduces the cross-modal mismatch rate in strong backlight scenes.

[0115] The spatial mask generated by the near-infrared branch dynamically enhances the low-contrast areas in the RGB image, and the target contrast is improved in strong backlighting scenes. The spectral reflectance matching verification module suppresses the spectral mismatch between branches and fruits through the differentiable threshold function and hard-constrained feature fusion, thereby reducing the cross-modal misjudgment rate. That is, through the collaborative fusion mechanism of heterogeneous features of RGB and near-infrared spectra, the system remains robust under complex lighting conditions.

[0116] It should be further explained that, during the specific implementation process, the small target adaptive optimization engine improves the detection capability of young fruit by reconstructing the micro-scale feature enhancement pyramid (MSE-FPN). Dilated convolution layers with dilation rates of 3, 6, and 9 are inserted between the C3-C5 layers of the standard FPN. The dilated convolution layer of the C3 layer (receptive field of 7×7) captures the tiny gradient changes at the edge of the young fruit, and the C5 layer (receptive field of 19×19) associates the global semantics of the occluded target. The feature reweighting module dynamically assigns layer weights according to the target size. The calculation formula is:

[0117]

[0118] Among them, wi: the normalized weight coefficient of the i-th level feature map, which is used to dynamically allocate the contribution of different levels to small target detection; s i is the average confidence of the small targets in the i-th layer, which is calculated by the weighted summation of the foreground probabilities output by the classification head; the formula is:

[0119] Where pk is the confidence of the k-th prediction box, area k is the target relative image area, Ⅱ is the indicator function, when the condition satisfies g i When <0.4, it is 1, otherwise it is 0;

[0120] g i : The mean Sobel gradient amplitude of the i-th level feature map reflects the edge clarity of the feature. The calculation formula is:

[0121]

[0122] ; Among them G x , G y is the gradient value of the Sobel operator in the horizontal and vertical directions; Ⅱ(g i <0.4): Gradient constraint indicator function, when g i It takes 1 when <0.4, otherwise it takes 0, which is used to apply a decay factor of 0.8 to the weight of the fuzzy feature level;

[0123] In the Gaussian-exponential mixture loss function, the classification loss is calculated by the difficulty coefficient:

[0124] L cls =-(1+e -D )y log(p);

[0125] D = α(l-Area) + βOR + γ(1-Contrast); for difficult samples, exponential weighting is applied, and the positioning loss is weighted by the center point Gaussian. The formula is:

[0126] Prioritize center offset optimization, where D is the sample difficulty coefficient, which is a linear combination of the following parameters:

[0127] Area: The target area relative to the image, such as young fruit Area <1%.

[0128] OR: Occlusion Rate, calculated as: OR = number of occluded pixels / total number of target pixels.

[0129] Contrast: Local contrast, calculated as the ratio of the grayscale standard deviation of the fruit area and the background area: Contrast = σ obj / (σ bg +e); e = 1e-5 is a small constant to prevent division by zero; α, β, γ: weight coefficients, with default values of α = 0.6, β = 0.3, and γ = 0.1, which are determined by grid search optimization on the validation set.

[0130] w reg: Positioning loss weight, related to the offset of the center point of the predicted box: Δx, Δy: Euclidean distance between the center point of the predicted box and the true box (unit: pixel); σ: Gaussian kernel scale parameter, σ is inversely proportional to the target size: σ = k / √Area, K = 0.1 is an empirical constant, Area is the relative area of the target; This scheme improves the IoU of young fruit detection.

[0131] By expanding the receptive field through the C3-C5 layer dilated convolution and combining it with the feature reweighting mechanism, the bounding box intersection-over-union ratio of young fruit detection is improved. The young fruit, medium fruit, and mature fruit detection heads are optimized in stages, and the pseudo-labels generated by knowledge distillation are used to reduce the young fruit missed detection rate. This has enabled the small target adaptive optimization engine to achieve a breakthrough in the difficult problem of young tea fruit detection through multi-scale feature enhancement and dynamic loss allocation.

[0132] It should be further explained that, during the specific implementation, the spectral reflectance matching verification module ensures the rationality of cross-modal fusion through physical prior constraints. The 61-channel feature map output by the near-infrared branch (corresponding to the 650-950nm band, 5nm interval) is piecewise linearly activated: when the channel value is within the measured threshold of [0.65, 0.82] in this band, the original gradient is retained, and the gradient exceeding the threshold is reset to zero and included in the penalty loss term;

[0133] Among them, the penalty loss formula is:

[0134]

[0135] Among them, f c : The characteristic value of the cth channel of the near-infrared branch, corresponding to a wavelength of 650+5(c-1)nm, when c=1 it is 650nm, and c=61 it is 950nm; λ: balance coefficient, the default value λ=0.3, determined by cross-validation; 0.65-0.82: The effective threshold of the measured reflectivity of Camellia oleifera in the 650-950nm band.

[0136] During the inference phase, a 16×16 sliding window is used to verify the proportion of valid reflectance channels. If the percentage is below 60% for three consecutive windows, the near-infrared branch parameters are frozen and the RGB branch C3 layer super-resolution reconstruction compensation is enabled. This mechanism imposes a hard penalty on features outside the physical reflectance range, ensuring that the spectral response conforms to the biological characteristics of the fruit, reducing the cross-modal mismatch rate and maintaining spectral validity under strong backlighting.

[0137] It should be further explained that in the specific implementation process, the sharing-competition mechanism of the collaborative verification framework realizes module parameter reuse by sharing the underlying feature extractor (MobileNetV3 and depth-wise separable convolution), and freezes the underlying parameters in the first three training stages to accelerate convergence;

[0138] Dynamic weight competitive loss function:

[0139] L ttotal αL aoaa +βL cmsf γL stae Among them, α, β, γ: dynamic weight coefficients, the adjustment rule is: occlusion module weight α: α new =α old × 0.9, ifΔ=[(R doaa -R cmsf ) / R all ]>0.05; modal module weight β: β new =β old ×1.1, ifΔ>0.05; small target module weight γ: γ new =γ old +0.05×(80\%-AP stae ), when AP stae <80% triggers linear compensation; R doaa : Recall rate of dynamic occlusion perception module in occlusion scene, R cmsf : Recall rate of the cross-modal fusion module in uneven lighting scenes, R all : Overall recall rate of the system; AP stae : Average Precision of the small object optimization engine on the young fruit test set; the test set contains 20,000 annotated images, covering 6 lighting conditions and 4 occlusion types, and the training-validation-test set is divided into 6:2:2.

[0140] L total :Total Loss Function, which is the overall loss function of the system, is composed of the weighted sum of the loss functions of multiple sub-modules and is used to globally optimize the model parameters; the dynamic weights α, β, and γ are used to balance the training objectives of different modules to avoid a single module dominating the optimization direction, and the gradient is distributed in the back propagation to ensure the coordinated optimization of each module and improve the overall performance of the system.

[0141] L doaa : Dynamic Occlusion-Aware Attention Loss can improve the recognition accuracy of occluded areas and optimize boundary positioning capabilities; the dedicated loss function of the dynamic occlusion-aware attention module includes occlusion area classification loss and boundary regression loss. The classification loss is the occlusion probability prediction formula:

[0142] Among them, y i is the occlusion label (0 / 1), p i is the predicted probability;

[0143] Regression loss, that is, the occlusion boundary positioning formula:

[0144] Among them, b i pred is the predicted bounding box, b i gt is the true bounding box;

[0145] Total loss: L doaa =L cls +0.5L reg .

[0146] L cmsf : Cross-Modal Spectral Fusion Loss is used to ensure that spectral features conform to physical laws and enhance the consistency of cross-modal features. The loss function of the cross-modal fusion network includes spectral matching loss and feature alignment loss, and is composed of:

[0147] Spectral matching loss, that is, the physical constraint formula of reflectivity:

[0148]

[0149] Feature alignment loss, that is, cross-modal consistency formula:

[0150]

[0151] Among them, f rgb and f nir It is the feature map of the same level of RGB and near-infrared branches.

[0152] Total loss:

[0153] L cmsf =L spec +0.3L align ;

[0154] Lstae: Small target adaptive optimization engine loss, used to enhance small target classification confidence and improve center point positioning accuracy; dedicated loss function for small target detection, including classification loss and positioning loss, focusing on optimizing small target detection; components include:

[0155] Classification loss, that is, Gaussian-exponential weighted focal loss formula:

[0156]

[0157] Among them, D i =α(1-Area i )+βOR i+γ(1-Contrast i ) is the sample difficulty coefficient.

[0158] Positioning loss, that is, the center point Gaussian weighted loss formula:

[0159]

[0160] in,

[0161] Total loss: L stae =L cls +0.7L reg .

[0162] The weight adjustment is driven by module performance differences to suppress the dominance of a single module. Combined with the cosine annealing strategy, the initial weights α = 0.4, β = 0.3, and γ = 0.3 are reset every 20 epochs to avoid local optimality.

[0163] The weight is dynamically adjusted according to the module performance, that is, the dynamic weight adjustment logic:

[0164] When the occlusion module L doaa When the training is dominated, reduce its weight to force the model to focus on cross-modal fusion; for example, when the occlusion module L doaa When the recall rate advantage exceeds 5%, reduce its weight coefficient α and increase the cross-modal module (CMSF) weight β, that is: when Δ=(R doaa -R cmsf ) / R all When >0.05, α←α×0.9, β←β×1.1.

[0165] When the small target module L stae When the performance is insufficient, increase its weight; for example, when the small target module L stae Weight γ varies with AP stae When <80%, γ←γ+0.05·(80%-AP stae );

[0166] The gradient norm constraint mechanism imposes a 0.7 attenuation factor on the dominant module gradient, that is: if Then Applying a 0.7 attenuation factor reduces the false detection rate in extreme scenarios and the module conflict rate.

[0167] It should be further explained that, in the specific implementation process, the two-stage adversarial verification process includes standard test set evaluation and extreme scenario stress testing; the extreme scenario simulation set uses CycleGAN to generate dynamic occlusion patches (coverage rate 60% to 80%); the Phong illumination model simulates a mixed environment of overexposure (brightness > 240) and low illumination (brightness < 50); and super-resolution is used to generate dense young fruits, 50-100 per image, with an area of <0.5%.

[0168] After locating the feature confusion area through Grad-CAM, course learning is used to optimize in three stages: the first stage strengthens the occlusion compensation capability and improves the recall rate; the second stage optimizes cross-modal alignment and reduces feature error; the final stage jointly adjusts the small target loss weight, improves AP, and ultimately reduces the fluctuation range of model performance.

[0169] A dynamic weighted competitive loss function is used to balance the contribution of each module, reduce the false detection rate in extreme scenarios, generate dynamic occlusion patches and dense small targets based on CycleGAN, strengthen the weak links of the model in stages through course learning, and reduce the feature conflict rate between modules. In other words, the stability of the model in extreme scenarios is ensured through a dynamic competition mechanism and adversarial iterative optimization through a collaborative verification framework. Among them, the dynamic occlusion patches generated by CycleGAN may generate noise that does not conform to the real branch and leaf morphology. To this end, additional constraints are added: when generating occlusion patches, they are forced to follow the real branch and leaf morphology database, such as vein direction and branch angle.

[0170] It should be further explained that, in its implementation, the physical data-driven approach of the spectral reflectance matching verification module uses an ASDFieldSpec4 spectrometer to construct a Camellia oleifera reflectance database. This database covers data from the 650-950nm band under different growth stages (young fruit, expansion stage, and maturity stage) and lighting conditions (midday, cloudy, and dusk), with a sampling interval of 5nm. When the 780nm channel feature exceeds the threshold [0.68, 0.79], the spatial attention weight of the corresponding position in the RGB branch is increased by 1.3 times. Three sub-pixel convolutions are then used to reconstruct the reflectance, restoring the effective channel ratio to 82%. This method reduces the spectral misclassification rate and shortens the compensation response time.

[0171] It is worth further explaining the hardware deployment of the attention-based camellia fruit recognition optimization system during its implementation: the system hardware uses an IMX477 RGB sensor and a SWIR-640 near-infrared camera to synchronously capture images. The Jetson Xavier NX embedded platform allocates two CPU cores to run the occlusion perception module (latency <30ms) and 128 CUDA cores to handle cross-modal fusion (occupancy ≤85%). Recognition results are transmitted to the robotic arm controller via the CAN bus, with a response latency of <200ms and overall low power consumption compared to traditional solutions. The output interface transmits recognition results to the agricultural machinery controller via the CAN bus at a frequency of 100Hz, with a response latency of <200ms.

[0172] It should be further explained that, in the specific implementation process, the implementation method of the system in the orchard monitoring scenario includes:

[0173] UAV dynamic acquisition unit: A quadrotor drone equipped with an image acquisition unit cruises in a serpentine path at a constant altitude, achieving centimeter-level track tracking through an RTK positioning module. The image acquisition frequency is dynamically adapted to the flight speed.

[0174] Real-time processing - feedback control loop:

[0175] a) The processing unit completes single-frame camellia fruit recognition within 150ms and transmits the positioning coordinates to the agricultural machinery via the LoRa wireless module;

[0176] b) When the robotic arm controller generates a picking path based on the fruit coordinates, it simultaneously receives the occlusion trend predicted by the dynamic occlusion perception module for the next three frames and dynamically adjusts the motion trajectory of the end effector;

[0177] Spectral-spatial joint verification mechanism: When the spectral reflectance matching verification module detects reflectance anomalies for five consecutive frames, it triggers the drone to hover and start multi-angle compensation shooting, including:

[0178] The pitch angle is ±15° and 3 sets of images are taken;

[0179] Turn on the near-infrared fill light to recapture features;

[0180] Multi-view features are input into the collaborative verification framework for confidence voting, and only recognition results that are consistent between the two sets of models are adopted.

[0181] It is important to further clarify that during implementation, in an orchard deployment, the quadrotor drone cruised at an altitude of 3±0.5m, at a speed of 1-2m / s, with an adaptively adjusted frame rate (10-15fps), and an RTK positioning error of less than 2cm. Recognition results were transmitted to the robotic arm via LoRa within 150ms, and the picking path was dynamically adjusted based on occlusion predictions for the next three frames. When an abnormal reflectivity was detected for five consecutive frames, the drone hovered and deflected ±15° to capture three sets of multi-view images. The features were recaptured using an 850nm fill light (300lux). Confidence voting reduced the missed detection rate and the probability of the robotic arm accidentally touching branches and leaves.

[0182] Through the collaborative design of a dynamic divide-and-conquer attention mechanism, cross-modal spectral fusion, a small target optimization engine, and a collaborative verification framework, this system improves the accuracy of camellia fruit recognition in occluded scenarios, reduces the cross-modal mismatch rate, and increases the IoU of young fruit detection. Embedded hardware deployment enables real-time response, improves the efficiency of orchard inspections, and provides a high-precision, low-power solution for smart agriculture.

[0183] The system utilizes a dynamic divide-and-conquer attention mechanism to enhance fruit edge features during occlusion region partitioning and differentiation processing. Combined with a spatiotemporal prediction model to capture dynamic changes in occlusion, it significantly improves feature discrimination in dense foliage environments. A cross-modal spectral fusion network employs cross-attention gating and physical prior constraints to deeply fuse RGB texture and near-infrared reflectance features, effectively enhancing robustness in object recognition in low-contrast scenes and reducing the risk of mismatches caused by sudden changes in illumination.

[0184] To address the challenge of small object detection, the system utilizes multi-scale receptive field enhancement and dynamic loss weighting to accurately capture subtle gradient changes in young fruit, balancing semantic and detailed features. Embedded hardware deployment enables efficient parallel computing, ensuring real-time response in complex scenarios. A collaborative verification framework, combined with adversarial optimization and a module competition mechanism, iteratively strengthens model weaknesses, improving overall system stability and generalization capabilities, providing a high-precision, low-latency Camellia oleifera fruit recognition solution for smart agriculture.

[0185] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0186] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A tea fruit recognition optimization system based on attention mechanism, characterized in that: include: The dynamic occlusion perception attention module is connected to the input layer of the cross-modal spectral feature collaborative fusion network through the spatiotemporal dual-stream occlusion inference unit. The output of the cross-modal spectral feature collaborative fusion network is connected to the heterogeneous pyramid fusion layer of the small target adaptive optimization engine. The positioning loss calculation unit of the small target adaptive optimization engine is fed back to the competitive loss function module of the collaborative verification framework. The dynamic occlusion perception attention module includes: Spatial dimension processing unit: inputs multi-scale feature maps and generates occlusion probability heat maps through deformable convolution layers; Time dimension processing unit: Based on lightweight LSTM, it fuses the current frame heat map and the historical frame optical flow features to output the occlusion trend prediction for the next 3-5 frames; The cross-modal spectral feature collaborative fusion network includes a spectral reflectance matching verification module for screening effective spectral features through physical prior constraints, and also includes a cross-attention gating mechanism between the RGB branch and the near-infrared branch; The micro-scale feature enhancement pyramid of the small target adaptive optimization engine inserts a multi-scale receptive field enhancement layer into the feature pyramid network and dynamically allocates layer weights through the feature reweighting module; The collaborative verification framework connects the modules by sharing the underlying feature extractor and adopts an adversarial verification process to iteratively optimize the weak areas of model decision making.

2. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 1, characterized in that: The dynamic occlusion-aware attention module has a divide-and-conquer attention strategy, including occlusion region division and differentiation processing: The occlusion area division includes: using the occlusion probability heat map to set a double threshold to divide the feature map into occlusion area, transition area, and visible area; Differentiation includes: a) The channel attention module is used to suppress the background channel response in the occluded area, followed by a spatial attention module; b) A local contrast enhancement module is deployed in the transition region and connected in parallel with the histogram of oriented gradients feature extractor; c) superimposing non-local neural networks in the visible area to establish long-range semantic associations; The training phase of the dynamic occlusion-aware attention module injects random masks through an adversarial occlusion simulator, and the inference phase adopts a dual-threshold voting mechanism to retain overlapping areas.

3. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 1, characterized in that: The cross-modal spectral feature collaborative fusion network includes: Spectral guidance unit: The spectral features output by the near-infrared branch generate a spatial mask and multiply it with the RGB branch features; Spatial constraint unit: The semantic features of the RGB branch generate a weight matrix and perform regional weighting on the near-infrared features; Dynamic routing selector: selects the fusion path based on the target confidence and gradient magnitude score; In the heterogeneous pyramid fusion layer, high-level features are upsampled after feature alignment, and low-level features are fused with low-resolution spectral features and high-resolution spatial features through super-resolution.

4. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 1, characterized in that: The micro-scale feature enhancement pyramid of the small target adaptive optimization engine includes: The multi-scale receptive field enhancement layer is configured with different receptive field sizes; The feature reweighting module dynamically generates layer weights based on the target size; Hybrid loss function: classification loss is dynamically weighted based on target attributes, and localization loss adopts center point weighting strategy.

5. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 1, characterized in that: The spectral reflectance matching verification module includes: Differentiable constraint unit: For the 61-channel feature map output by the near-infrared branch, corresponding to the 650-950nm band, piecewise linear activation is performed. When the channel value is within the measured threshold [0.65, 0.82] of this band, the original gradient is retained, and the gradient exceeding the threshold is reset to zero; Inference verification mechanism: The fused features are sampled in a 16×16 sliding window. When the effective reflectance channel ratio of three consecutive windows is less than 60%, the C3 layer super-resolution reconstruction compensation of the RGB branch is triggered. Among them, the spectral reflectance matching verification module builds a database based on the ASD FieldSpec4 spectrometer.

6. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 1, characterized in that: The sharing-competition mechanism of the collaborative verification framework includes: Parameter freezing strategy: The underlying parameters of MobileNetV3 and depth-wise separable convolution are fixed in the first three training stages.

7. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 6, characterized in that: The adversarial verification process includes extreme scenario simulation set generation and error feedback mechanism; The extreme scene simulation set generation includes generating dynamic occlusion patches based on CycleGAN with a coverage rate of 60% to 80%; generating overexposed and low-light mixed images in combination with the Phong illumination model; and generating 50-100 young fruits with an area of less than 0.5% per image through super-resolution reconstruction; The error feedback mechanism includes: using Grad-CAM to locate feature confusion areas, and strengthening training in three stages through curriculum learning: the first stage freezes CMSF and STAE parameters and focuses on DOAA optimization; the second stage unfreezes CMSF and sets a cross-modal similarity threshold of 0.7; the third stage jointly optimizes STAE loss weights and feature pyramid parameters; Among them, the positioning loss gradient is back-propagated to the error feedback mechanism of the adversarial verification process to dynamically adjust the optimization weights in the course learning stage.

8. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 5, characterized in that: The physical data driven method of the spectral reflectance matching verification module includes: Spectral database construction: An ASD FieldSpec4 spectrometer was used to collect reflectance data in the 650-950nm band at different growth stages; Back-propagation linkage optimization: When the near-infrared feature exceeds the threshold [0.68, 0.79] in the 780nm channel, the spatial attention weight of the corresponding position in the RGB branch is increased by a factor of 1.3; Real-time verification exception handling process: After freezing the near-infrared branch parameters, enable the C3 layer hole convolution feature of the RGB branch to perform three sub-pixel convolution reconstructions until the reflectivity effective channel ratio is ≥ 60%.

9. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 1, characterized in that: The hardware deployment of the tea fruit recognition optimization system based on the attention mechanism includes: Image acquisition unit: IMX477 RGB sensor and SWIR-640 near-infrared camera synchronized trigger acquisition; Processing unit: Jetson Xavier NX embedded platform, configured with a multi-threaded task scheduler: a) The first thread runs the dynamic occlusion awareness attention module and is allocated to two CPU cores; b) The second thread runs the cross-modal fusion network and allocates 128 CUDA cores; c) The third thread runs the small target optimization engine and allocates 64 CUDA cores; Output interface: transmits the recognition results to the agricultural machinery controller via the CAN bus.

10. The oil-tea camellia fruit recognition optimization system based on the attention mechanism according to claim 9, characterized in that: The implementation of the system in the orchard monitoring scenario includes: UAV dynamic acquisition unit: A quadrotor drone equipped with the image acquisition unit cruises in a serpentine path at a constant altitude (3±0.5m), achieving centimeter-level track tracking through the RTK positioning module. The image acquisition frequency is dynamically adapted to the flight speed (10fps at 1m / s, 15fps at 2m / s). Real-time processing - feedback control loop: a) The processing unit completes single-frame camellia fruit recognition within 150ms and transmits the positioning coordinates to the agricultural machinery via the LoRa wireless module; b) when the robotic arm controller generates a picking path based on the fruit coordinates, it simultaneously receives the occlusion trend of the next three frames predicted by the dynamic occlusion perception module and dynamically adjusts the motion trajectory of the end effector; Spectral-spatial joint verification mechanism: When the spectral reflectance matching verification module detects reflectance anomalies for five consecutive frames, it triggers the drone to hover and start multi-angle compensation shooting, including: The pitch angle is ±15° and 3 sets of images are taken; Turn on the near-infrared fill light (wavelength 850nm, intensity 300lux) for feature recapture; The multi-view features are input into the collaborative verification framework for confidence voting, and only the recognition results predicted by the two sets of models are adopted.

Citation Information

Patent Citations

  • Image sequence motion occlusion detection method based on optical flow and multi-scale context

    CN111612825A

  • Picking robot fruit identification method based on millimeter wave radar

    CN118279761A

  • Smart city camera multi-target tracking method and system

    CN118297984A

  • Audio-driven speaking face synthesis method based on point-edge-surface space-time alignment

    CN118351586A

  • Scene space three-dimensional model dynamic modeling method based on multi-modal data

    CN119339008A

Cited By

  • A profile back-spraying defect detection method and system

    CN122391236A