YOLO11n-g-based oil tea fruit space position detection method and system and medium

By improving the backbone network and neck network of the YOLO11n-g model, the real-time and resource occupation problems of oleifera fruit detection are solved, and efficient and accurate oleifera fruit detection is achieved, which improves the detection accuracy and real-timeness.

CN120544006APending Publication Date: 2025-08-26ANHUI AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510670154.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing oil tea fruit detection methods have shortcomings in real-time, detection efficiency and resource utilization, which are difficult to meet the real-time harvesting needs and are highly complex in calculations, so they cannot adapt to efficient operations of large-scale oil tea plantations.

Method used

The improved YOLO11n-g model is adopted, and the C3k2 module of the backbone network is replaced as the C3k2-LGDConv module, and a dynamic upsampling module is introduced into the neck network to optimize feature extraction and fusion, enhancing the robustness of the model in lighting changes and background complex scenarios.

Benefits of technology

It significantly improves the feature extraction capability of small-size and densely distributed tea fruit, improves detection accuracy and real-time performance, reduces computing resource occupation, and meets the real-time detection requirements in agricultural scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544006A_ABST
    Figure CN120544006A_ABST
Patent Text Reader

Abstract

The invention discloses a YOLO11n-g-based oil-tea camellia fruit space position detection method and system and a medium, and the method comprises the steps: collecting image data of oil-tea camellia fruits, screening and marking the image data, generating an image data set, constructing a YOLO11n-g model, carrying out the feature extraction through a backbone network, and obtaining a YOLO11n-g model; carrying out dynamic up-sampling on the image after feature extraction based on a neck network; setting a data set path, inputting a training set and a test set into the YOLO11n-g model to complete model training, inputting a to-be-detected camellia oleifera fruit image obtained in real time into the model for feature extraction and multi-scale feature fusion, and then outputting the confidence coefficient, the bounding box and the coordinate result of the generated to-be-detected camellia oleifera fruit by the head. Therefore, the spatial position of the camellia oleifera fruit can be detected. According to the method, the reasoning efficiency of the model is optimized, and the precision and real-time performance of oil tea fruit detection in an agricultural scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and deep learning technology, and in particular to a method, system, and medium for detecting the spatiotemporal position of camellia fruit based on YOLO11n-g. Background Art

[0002] Camellia oleifera is one of the world's four major woody oil plants and a species unique to my country, with a history of cultivation and utilization spanning over 2,300 years. Tea oil, known as the "Olive Oil of the East," is rich in essential trace elements and possesses high nutritional and health benefits. It is the Food and Agriculture Organization's top recommendation for a health-promoting plant oil. Since mature camellia berries easily fall off, prompt harvesting is crucial to ensure a high-quality harvest.

[0003] Current traditional camellia fruit detection methods mainly rely on video streams or collected static images for identification. However, this approach has significant drawbacks: First, video stream processing requires frame-by-frame analysis, which has high computational complexity, resulting in slow detection speed and difficulty meeting the needs of real-time harvesting; second, static image-based detection requires manual sample collection, which is a cumbersome and inefficient process and cannot meet the requirements of efficient operation in large-scale camellia plantations; in addition, traditional methods usually rely on complex model structures, occupy large computing resources, and have long model inference times, which are not conducive to deployment on resource-constrained harvesting equipment. At the same time, energy consumption is high and equipment endurance is insufficient. Therefore, there is an urgent need for an efficient, real-time camellia fruit detection method that is suitable for resource-constrained equipment to improve harvesting efficiency and reduce energy consumption. Summary of the Invention

[0004] In order to solve the technical problems of existing tea fruit detection methods in terms of real-time performance, detection efficiency, adaptability to complex scenarios and resource occupation, the present invention provides a tea fruit spatiotemporal position detection method, system and medium based on YOLO11n-g, aiming to achieve efficient and accurate tea fruit detection, meet real-time harvesting needs and reduce computing resource occupation.

[0005] In a first aspect, the present invention proposes a method for detecting the temporal and spatial position of camellia fruit based on YOLO11n-g, the method comprising the following steps: S1 collects image data of Camellia oleifera, filters and annotates the collected image data to generate an image dataset, which includes a training set, a validation set, and a test set; S2. Build a YOLO11n-g model, extract features through the backbone network of the YOLO11n-g model, and dynamically upsample the feature-extracted image based on the neck network of the YOLO11n-g model; The YOLO11n-g model is an improvement based on the YOLO11n model. The C3k2 module of the YOLO11n model backbone network is replaced with the C3k2-LGDConv module, and during the dynamic upsampling process of the YOLO11n-g model, the upsampling module of the YOLO11n model is replaced with a dynamic upsampling module. S3. Set the dataset path, input the training set and test set into the YOLO11n-g model to complete model training, input the real-time image of the oil-tea camellia fruit to be detected into the model for feature extraction and multi-scale feature fusion, and then output the confidence score, bounding box, and coordinate results of the oil-tea camellia fruit to be detected by the head, thereby realizing the detection of the spatial and temporal position of the oil-tea camellia fruit.

[0006] Furthermore, in S1, the image is manually labeled using Robotflow, and the labeled category is camellia fruit, and the labeled format is txt format.

[0007] Furthermore, the C3k2-LGDConv module includes a Bottleneck_GhostDynamicConv submodule, a liquid neural network LNN layer and a first 1×1 convolution module connected in sequence; wherein, the Bottleneck_GhostDynamicConv submodule includes a second 1×1 convolution module, a GhostDynamicConv module and a third 1×1 convolution module connected in sequence, which are used to extract lightweight multi-scale features; the liquid neural network LNN layer receives the output features of the Bottleneck_GhostDynamicConv submodule, which is used to dynamically model the timing information through differential equations to enhance the robustness of the features in scenarios of occlusion or light changes; the first 1×1 convolution module is used to perform channel adjustment and fusion on the output features of the liquid neural network layer to generate a final fusion feature map.

[0008] Furthermore, the processing process of the LGDConv module is as follows: performing a group convolution operation on the feature map of the input image data, extracting spatial features and generating a main feature map; performing semantic distribution extraction on the input feature map through a dynamic weight generator to generate a convolution kernel weight; using the convolution kernel weight to perform a convolution operation on the main feature map to generate a ghost feature map; splicing the ghost feature map and the main feature map at the channel level to generate an initial fused feature map; inputting the initial fused feature map into the liquid neural network LNN layer, dynamically modeling the timing information through differential equations, and generating a dynamic enhanced feature map to improve the robustness of the feature in scenes with occlusion or light changes; finally, performing channel fusion on the dynamic enhanced feature map through the first 1×1 convolution module to generate a final fused feature map.

[0009] Furthermore, the dynamic weight generator consists of a fourth 1×1 convolution module and a sigmoid activation function, which is used to capture the contextual information of the features in the feature map and generate weight coefficients to enhance the adaptability of the YOLO11n-g model to the oil tea fruit target under different lighting and backgrounds.

[0010] Furthermore, the step of feature extraction by the backbone network includes: performing preliminary feature extraction on the input image after convolution and compressing it to 64 channels, and gradually increasing the number of channels to 1024 through multiple layers of the C3k2-LGDConv module; passing the image that has undergone preliminary feature extraction layer by layer through the C3k2-LGDConv module, wherein each layer of the C3k2-LGDConv module contains multiple layers of Bottleneck_GhostDynamicConv sub-modules and liquid neural network (LNN) layers, and the liquid neural network (LNN) layer dynamically models the timing information through differential equations to improve the robustness of the features in scenes of occlusion or light changes; after upsampling the multiple feature fusion maps through the dynamic upsampling module, a final output feature map is generated.

[0011] Furthermore, the dynamic upsampling module includes a convolution layer, an interpolation layer and an activation function connected in sequence, wherein the interpolation layer adopts bilinear interpolation or nearest neighbor interpolation to optimize the spatial resolution and feature expression of the feature map.

[0012] Furthermore, the processing process of the dynamic upsampling module is: receiving the input fused feature map, extracting the semantic information of the fused feature map after convolution and generating dynamic interpolation weights using a dynamic interpolation kernel generator; upsampling the dynamic interpolation weights, and generating an output feature map after processing through an activation function.

[0013] In a second aspect, the present invention further provides a system for detecting the temporal and spatial position of oil-tea camellia fruit, for implementing the method described in the first aspect, comprising: A data set generation module is used to collect image data of oil-tea camellia fruits, screen and annotate the collected image data, and generate an image data set, wherein the image data set includes a training set, a validation set, and a test set; A model construction module is used to build a YOLO11n-g model, extract features through the backbone network of the YOLO11n-g model, and dynamically upsample the image after feature extraction based on the neck network of the YOLO11n-g model; The YOLO11n-g model is an improvement based on the YOLO11n model. The C3k2 module of the YOLO11n model backbone network is replaced with the C3k2-LGDConv module, and during the dynamic upsampling process of the YOLO11n-g model, the upsampling module of the YOLO11n model is replaced with a dynamic upsampling module. The target detection module is used to set the dataset path, input the training set and test set into the YOLO11n-g model to complete model training, input the real-time oil-tea fruit image to be detected into the model for feature extraction and multi-scale feature fusion, and then output the confidence, bounding box and coordinate results of the oil-tea fruit to be detected by the head, thereby realizing the detection of the spatial and temporal position of the oil-tea fruit.

[0014] In a third aspect, the present invention further proposes a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and the instructions are used to be read by a machine so that the machine executes the detection method as described in the first aspect.

[0015] The present invention has the following beneficial effects: The present invention proposes a method for detecting the spatiotemporal position of camellia fruit based on the YOLO11n-g model. Based on the existing YOLO11n backbone network, the C3k2 module is improved to obtain the C3k2-LGDConv module in the backbone network of the YOLO11n-g model. The method solves the problem of insufficient feature extraction caused by illumination changes and complex background when detecting and positioning camellia fruit in actual environments, and significantly improves the model's feature extraction capability for small-sized and densely distributed camellia fruit. The present invention introduces a dynamic upsampling module into the neck network of the existing YOLO11n model to replace the original fixed upsampling operation, and optimizes the multi-scale feature fusion effect by dynamically adjusting the upsampling strategy, thereby reducing the problem of detail loss of small-sized camellia fruit in the feature fusion process and further improving the detection accuracy. The present invention optimizes the inference efficiency of the model according to the real-time requirements of the camellia fruit detection task, meets the real-time detection requirements in agricultural scenarios, while maintaining high precision, and effectively improves the precision and real-time performance of camellia fruit detection in agricultural scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flow chart of the temporal and spatial position detection method of Camellia oleifera fruit based on YOLO11n-g proposed in the present invention.

[0017] Figure 2 This is a schematic diagram of the YOLO11n-g model structure in the YOLO11n-g-based tea fruit temporal and spatial position detection method proposed in the present invention.

[0018] Figure 3This is a schematic diagram of the structure of the LGDConv module in the YOLO11n-g model proposed in this invention.

[0019] Figure 4 This is a schematic diagram of the structure of the C3K2-LGDConv module in the YOLO11n-g model proposed in this invention.

[0020] Figure 5 This is a schematic diagram of the structure of the dynamic upsampling module in the YOLO11n-g model proposed in this invention.

[0021] Figure 6 This is the overall workflow diagram of the YOLO11n-g model proposed in this invention for real-time detection of oil-tea fruit spatial positioning in real-time scenarios. DETAILED DESCRIPTION

[0022] The present application is described in further detail below in conjunction with the accompanying drawings. It is necessary to point out that the following specific implementation methods are only used to further illustrate the present application and cannot be understood as limiting the scope of protection of the present application. Technicians in this field can make some non-essential improvements and adjustments to the present application based on the above application content.

[0023] like Figure 1 As shown in FIG, a method for detecting the temporal and spatial position of camellia fruit based on YOLO11n-g is proposed, which includes the following steps: S1. Collecting image data of oil-tea camellia fruit, screening and annotating the collected image data to generate an image dataset, wherein the image dataset includes a training set, a validation set, and a test set.

[0024] In this example, the spatial position of tea fruit is detected by improving the existing YOLO11n model. This process requires a dataset for model training and testing. Therefore, high-quality, clear photos of tea fruit from tea orchards are collected and screened for high-quality, clear images under different lighting conditions. Images are manually annotated using Robotflow, primarily labeling tea fruit. The annotated labels are formatted in txt format to better meet training requirements. The annotated dataset is divided into a training set, a validation set, and a test set in a 7:2:1 ratio.

[0025] S2. Construct a YOLO11n-g model, extract features using the backbone network of the YOLO11n-g model, and dynamically upsample the feature-extracted image based on the neck network of the YOLO11n-g model.

[0026] The existing YOLO11n model was improved to obtain the YOLO11n-g model. Specifically, the original YOLO11n model used the C3k2 module in the backbone network and a fixed upsampling module in the neck network. The improved C3k2-LGDConv module replaces the C3k2 module, upgrading the original stacked structure to an improved Bottleneck structure. The C3k2-LGDConv module contains multiple Bottleneck_GhostDynamicConv submodules and a liquid neural network (LNN) layer. The LNN layer dynamically models temporal information through differential equations to enhance feature robustness in scenarios with occlusion or changing lighting. Furthermore, a dynamic upsampling module replaces the original fixed upsampling module to optimize the feature fusion process.

[0027] The improved model structure is as follows Figure 2 As shown in the figure, the improved GhostDynamicConv is used to replace the traditional convolution operation and the new LGDConv module is obtained by improving C3k2. This solves the problem of insufficient feature extraction caused by lighting changes and complex background when detecting and positioning tea oil fruits in actual environments, and significantly improves the model's feature extraction ability for small-sized and densely distributed tea oil fruits.

[0028] Secondly, a dynamic upsampling module (Dynamic Upsampling Module) was introduced into the neck network of the original YOLO11n to replace the original fixed upsampling operation. By dynamically adjusting the upsampling strategy, the multi-scale feature fusion effect was optimized, reducing the problem of detail loss in the feature fusion process of small-sized tea-oil fruits, and further improving the detection accuracy.

[0029] In this embodiment, if Figure 3As shown in the figure, the C3k2-LGDConv module replaces the Conv2D layers in multiple Bottleneck structures in the original C3k2 module with GhostDynamicConv layers to form multiple Bottleneck_GhostDynamicConv sub-modules. The GhostDynamicConv units enhance the feature extraction capability through stacking, improving computational efficiency and feature expression capabilities. At the same time, a liquid neural network (LNN) layer is introduced after each Bottleneck_GhostDynamicConv sub-module to dynamically model timing information through differential equations to enhance the robustness of features in scenarios with occlusion or light changes. The stacked structure of the C3k2-LGDConv module includes a first 1×1 convolution module to perform dimensionality reduction on the feature map, compressing the number of channels of the feature map to 1 / 2 of the original to reduce the amount of subsequent calculations; among them, the Bottleneck_GhostDynamicConv submodule includes a second 1×1 convolution module, a GhostDynamicConv module and a third 1×1 convolution module connected in sequence to efficiently extract the features of the oil-tea fruit target; the stacked structure also finally includes a 1×1 convolution module to perform dimensionality increase on the feature map, restoring the number of channels of the feature map to the original number of channels to ensure compatibility with subsequent network structures.

[0030] In this embodiment, if Figure 4 As shown in the figure, the processing process of the LGDConv module is as follows: performing a group convolution operation on the feature map of the input image data, extracting the spatial features and generating a main feature map; performing semantic distribution extraction on the input feature map through a dynamic weight generator to generate a convolution kernel weight, wherein the dynamic weight generator is composed of a fourth 1×1 convolution layer and a Sigmoid activation function, which is used to capture the context information of the features in the feature map and generate a weight coefficient to enhance the adaptability of the YOLO11n-g model to the oil tea fruit target under different lighting and backgrounds; using the convolution kernel weight to perform a convolution operation on the main feature map to generate a ghost feature map, wherein the generated ghost feature map is A linear transformation is generated from the main feature map, further reducing computational complexity while retaining rich feature information. The ghost feature map and the main feature map are concatenated at the channel level to generate an initial fused feature map. The initial fused feature map is input into the liquid neural network (LNN) layer, and the timing information is dynamically modeled by differential equations to generate a dynamic enhanced feature map to improve the robustness of features in scenarios with occlusion or changing lighting. Finally, a 1×1 convolution operation is performed on the dynamic enhanced feature map to perform channel fusion to generate the final fused feature map, thereby enhancing the feature expression capability of the YOLO11n-g model for small-sized and densely distributed tea-oil fruits while maintaining efficient computation.

[0031] In this embodiment, the formula of Ghostdynamic is as follows: ; ; ; ; ; ; In this embodiment, if Figure 4 As shown, the backbone network performs feature extraction steps including: performing a group convolution operation on the input image data feature map X, i.e. , after preliminary feature extraction, the main feature map is generated ; The main feature map is generated by the dynamic weight generator Perform semantic distribution extraction and generate convolution kernel weights , where the dynamic weight generator consists of a 1×1 convolution layer and a Sigmoid activation function, which is used to capture the contextual information of the feature map and enhance the adaptability of the YOLO11n-g model to the oil tea fruit target under different lighting and backgrounds; using the convolution kernel weight Main feature map Perform a linear transformation, that is Generate ghost feature map ; The ghost feature map and main feature map Perform channel-level splicing to generate the initial fusion feature map ; The initial fusion feature map Input the liquid neural network (LNN) layer, through the differential equation Dynamic modeling timing information, where is the initial condition, and evolves along time step t to Generate dynamic enhancement feature map , to improve the robustness of features in scenes with occlusion or light changes; finally, through a 1×1 convolution operation, i.e. , perform channel fusion on the dynamic enhancement feature map to generate the final output feature map The backbone network then gradually increases the number of channels by stacking multiple C3k2-LGDConv modules and uses a dynamic upsampling module to upsample multiple feature fusion maps to further optimize feature representation. The C3k2-LGDConv module improves computational efficiency and feature extraction capabilities by using the Bottleneck_GhostDynamicConv submodule. Combined with LNN temporal modeling, it significantly enhances the model's spatial position detection performance for oil-tea fruit.

[0032] In this embodiment, if Figure 5 As shown, the dynamic upsampling module processes as follows: receiving the input fused feature map, extracting its semantic information after convolution, and generating dynamic interpolation weights using a dynamic interpolation kernel generator; upsampling the dynamic interpolation weights, and processing them through an activation function to generate an output feature map. The dynamic upsampling module includes a sequentially connected convolutional layer, an interpolation layer, and an activation function. The interpolation layer uses bilinear interpolation or nearest neighbor interpolation to optimize the spatial resolution and feature expression of the feature map.

[0033] The formula of the dynamic upsampling module is as follows: ; ; ; ; in is the input feature map, It is a semantic feature map extracted by 1×1 convolution, is the interpolation weight generated by the dynamic interpolation kernel generator, is the upsampled feature map, It is the final output feature map.

[0034] S3. Set the dataset path, input the training set and test set into the YOLO11n-g model to complete model training, input the real-time image of the oil-tea camellia fruit to be detected into the model for feature extraction and multi-scale feature fusion, and then output the confidence score, bounding box, and coordinate results of the oil-tea camellia fruit to be detected by the head, thereby realizing the detection of the spatial and temporal position of the oil-tea camellia fruit.

[0035] In this embodiment, if Figure 6 As shown in Figure 2, the process of feature extraction using the backbone network is as follows: First, a convolution module is used to extract the initial features of the input image and compress the image data to 64 channels. The number of channels is gradually increased to 1024 through the multi-layer C3k2-LGDConv module to enhance the model's feature extraction ability for small-sized and densely distributed camellia fruits. Figure 1 Each layer of the C3k2-LGDConv module is stacked with multiple layers of Bottleneck_GhostDynamicConv submodules and liquid neural network (LNN) layers. The liquid neural network (LNN) layer dynamically models the timing information through differential equations to improve the robustness of features in scenes with occlusion or light changes. The C3k2-LGDConv module is stacked with multiple layers of Bottleneck_GhostDynamicConv submodules to gradually extract high-level feature information from low-level features. Subsequently, a dynamic upsampling module is introduced into the neck network to perform multi-scale feature extraction and fusion on the feature map to further optimize the feature expression of small-sized tea fruit.

[0036] The process of multi-scale feature extraction and fusion using dynamic upsampling in the neck network is as follows: First, semantic information is extracted from the high-resolution feature map, and the 1×1 convolution module extracts the semantic distribution of the feature map. Then, the dynamic interpolation kernel generator generates interpolation weights according to the semantic distribution, and the upsampling strategy is dynamically adjusted. Finally, the upsampled feature map is fused with the low-resolution feature map to generate a multi-scale feature map to improve the detection accuracy of small-sized and densely distributed camellia fruits.

[0037] The process of multi-scale feature extraction and fusion in natural scenes using the dynamic upsampling module is as follows: Firstly, the dynamic upsampling module is used to extract and fuse the features of camellia oleifera fruits of different scales in natural scenes to generate a multi-scale feature map. The upsampling ratio is adjusted by dynamic interpolation weights to enhance the model's ability to express the features of small-sized camellia oleifera fruits. Secondly, the original convolution module is used to further extract features from the multi-scale feature map. Finally, the multi-scale feature map is fused with the high-level semantic feature map to generate the final feature map, thereby improving the detection performance of the model in complex agricultural scenarios.

[0038] It should be noted that the process of feature extraction and multi-scale feature fusion in the oil-tea fruit image input model of this embodiment is also included in the training stage when the model is established. The model training and testing process is realized by feature extraction and multi-scale feature fusion of the data. When detecting the image acquired in real time, the feature extraction and multi-scale feature fusion process is also included, thereby achieving high detection accuracy and real-time performance of the oil-tea fruit image to be detected acquired in real time.

[0039] According to the purpose of the present invention, the present invention also proposes a system for detecting the temporal and spatial position of oil-tea camellia fruit, which is used to implement the above-mentioned detection method, comprising: A data set generation module is used to collect image data of oil-tea camellia fruits, screen and annotate the collected image data, and generate an image data set, wherein the image data set includes a training set, a validation set, and a test set; A model construction module is used to build a YOLO11n-g model, extract features through the backbone network of the YOLO11n-g model, and dynamically upsample the image after feature extraction based on the neck network of the YOLO11n-g model; The YOLO11n-g model is an improvement based on the YOLO11n model. The C3k2 module of the YOLO11n model backbone network is replaced with the C3k2-LGDConv module, and during the dynamic upsampling process of the YOLO11n-g model, the upsampling module of the YOLO11n model is replaced with a dynamic upsampling module. The target detection module is used to set the dataset path, input the training set and test set into the YOLO11n-g model to complete model training, input the real-time oil-tea fruit image to be detected into the model for feature extraction and multi-scale feature fusion, and then output the confidence, bounding box and coordinate results of the oil-tea fruit to be detected by the head, thereby realizing the detection of the spatial and temporal position of the oil-tea fruit.

[0040] In a third aspect, the present invention further proposes a computer-readable storage medium storing a computer program, which is called by a processor to implement: the steps of the above-mentioned method for detecting the spatiotemporal position of tea fruit based on YOLO11n-g.

[0041] For the specific implementation process of each step, please refer to the description of the above method.

[0042] It should be understood that in the embodiments of the present invention, the processor referred to may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0043] The performance parameters of the camellia fruit detection model based on YOLO11n-g were analyzed and compared with existing target detection methods. The comparison results are shown in Table 1: Table 1 Comparison of detection performance of different detection methods ; It can be seen from the test results that the present invention further improves the accuracy of camellia fruit detection on the basis of meeting the requirements of camellia fruit temporal detection.

[0044] The method proposed in this paper can efficiently detect camellia fruit in agricultural scenarios. First, the YOLO11n backbone network is improved by combining a multi-layer C3k2-LGDConv module, mitigating the problem of insufficient feature extraction of camellia fruit in agricultural scenarios due to varying lighting and complex backgrounds. Next, a dynamic upsampling module is added to the neck network to optimize multi-scale feature fusion by dynamically adjusting the upsampling strategy, reducing detail loss during the feature fusion process for small and densely distributed camellia fruit. Finally, the YOLO11n detection head is used to support feature information extraction from multi-scale feature maps, improving the detection accuracy of camellia fruit in agricultural scenarios while also reducing the model complexity of camellia fruit detection in agricultural scenarios. Ultimately, building on the original YOLO11n detection head, the method achieves both detection accuracy and real-time performance for camellia fruit detection in agricultural scenarios.

[0045] This embodiment also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for detecting the spatial position of camellia oleifera fruits provided by the above methods is implemented.

[0046] In the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0047] In the present invention, unless otherwise expressly specified or limited, when a first feature is "above" or "below" a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediary. Furthermore, when a first feature is "above," "above," or "above" a second feature, it may mean that the first feature is directly above or diagonally above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is "below," "below," or "below" a second feature, it may mean that the first feature is directly below or diagonally below the second feature, or simply means that the first feature is at a lower level than the second feature.

[0048] The above-described embodiments merely illustrate several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.

Claims

1. A method for detecting the temporal and spatial position of camellia fruit based on YOLO11n-g, characterized in that: The specific steps include: S1 collects image data of Camellia fruit, filters and annotates the collected image data to generate an image dataset, which includes a training set, a validation set, and a test set; S2. Build a YOLO11n-g model, extract features through the backbone network of the YOLO11n-g model, and dynamically upsample the feature-extracted image based on the neck network of the YOLO11n-g model; The YOLO11n-g model is an improvement based on the YOLO11n model. The C3k2 module of the YOLO11n model backbone network is replaced with the C3k2-LGDConv module, and during the dynamic upsampling process of the YOLO11n-g model, the upsampling module of the YOLO11n model is replaced with a dynamic upsampling module. S3. Set the dataset path, input the training set and test set into the YOLO11n-g model to complete model training, input the real-time image of the oil-tea camellia fruit to be detected into the model for feature extraction and multi-scale feature fusion, and then output the confidence score, bounding box, and coordinate results of the oil-tea camellia fruit to be detected by the head, thereby realizing the detection of the spatial and temporal position of the oil-tea camellia fruit.

2. The method for detecting the temporal and spatial position of camellia fruit according to claim 1, wherein: In S1, the images were manually labeled using Robotflow. The labeled category was Camellia oleifera and the labeled format was txt.

3. The method for detecting the temporal and spatial position of camellia fruit according to claim 1, wherein: The C3k2-LGDConv module includes a Bottleneck_GhostDynamicConv submodule, a liquid neural network LNN layer and a first 1×1 convolution module connected in sequence; wherein, the Bottleneck_GhostDynamicConv submodule includes a second 1×1 convolution module, a GhostDynamicConv module and a third 1×1 convolution module connected in sequence, which is used to extract lightweight multi-scale features; the liquid neural network LNN layer receives the output features of the Bottleneck_GhostDynamicConv submodule, which is used to dynamically model timing information through differential equations to enhance the robustness of features in scenarios of occlusion or light changes; the first 1×1 convolution module is used to perform channel adjustment and fusion on the output features of the liquid neural network layer to generate a final fused feature map.

4. The method for detecting the temporal and spatial position of camellia fruit according to claim 3, wherein: The processing process of the LGDConv module is as follows: performing a group convolution operation on the feature map of the input image data to extract spatial features and generate a main feature map; performing semantic distribution extraction on the input feature map through a dynamic weight generator to generate convolution kernel weights; performing a convolution operation on the main feature map using the convolution kernel weights to generate a ghost feature map; and performing channel-level splicing on the ghost feature map and the main feature map to generate an initial fused feature map. The initial fused feature map is input into the liquid neural network LNN layer, and the timing information is dynamically modeled through differential equations to generate a dynamic enhanced feature map to improve the robustness of the feature in scenes with occlusion or light changes. Finally, the dynamic enhanced feature map is channel-fused through the first 1×1 convolution module to generate the final fused feature map.

5. The method for detecting the temporal and spatial position of camellia fruit according to claim 4, wherein: The dynamic weight generator consists of a fourth 1×1 convolution module and a sigmoid activation function, which is used to capture the contextual information of the features in the feature map and generate weight coefficients to enhance the adaptability of the YOLO11n-g model to the oil tea fruit target under different lighting and backgrounds.

6. The method for detecting the temporal and spatial position of camellia fruit according to claim 4, wherein: The steps of performing feature extraction on the backbone network include: The input image is convolved for preliminary feature extraction and compressed to 64 channels, and the number of channels is gradually increased to 1024 through multiple layers of the C3k2-LGDConv module. The image that has undergone preliminary feature extraction is passed through the C3k2-LGDConv module layer by layer, where each layer of the C3k2-LGDConv module contains multiple layers of Bottleneck_GhostDynamicConv sub-modules and liquid neural network (LNN) layers. The liquid neural network (LNN) layer dynamically models timing information through differential equations to improve the robustness of features in scenarios with occlusion or light changes. After upsampling the multiple feature fusion maps through the dynamic upsampling module, the final output feature map is generated.

7. The method for detecting the temporal and spatial position of oil-tea camellia fruit according to claim 6, wherein: The dynamic upsampling module includes a convolution layer, an interpolation layer and an activation function connected in sequence, wherein the interpolation layer adopts bilinear interpolation or nearest neighbor interpolation to optimize the spatial resolution and feature expression of the feature map.

8. The method for detecting the temporal and spatial position of camellia fruit according to claim 7, wherein: The processing process of the dynamic upsampling module is as follows: receiving the input fused feature map, extracting the semantic information of the fused feature map after convolution and generating dynamic interpolation weights using a dynamic interpolation kernel generator; upsampling the dynamic interpolation weights and generating an output feature map after processing through an activation function.

9. A system for detecting the temporal and spatial position of oil-tea camellia fruit, for implementing the detection method according to any one of claims 1 to 8, characterized in that: include: A data set generation module is used to collect image data of oil-tea camellia fruits, screen and annotate the collected image data, and generate an image data set, wherein the image data set includes a training set, a validation set, and a test set; A model construction module is used to build a YOLO11n-g model, extract features through the backbone network of the YOLO11n-g model, and dynamically upsample the image after feature extraction based on the neck network of the YOLO11n-g model; The YOLO11n-g model is an improvement based on the YOLO11n model. The C3k2 module of the YOLO11n model backbone network is replaced with the LGDConv module, and during the dynamic upsampling process of the YOLO11n-g model, the upsampling module of the YOLO11n model is replaced with a dynamic upsampling module. The target detection module is used to set the dataset path, input the training set and test set into the YOLO11n-g model to complete model training, input the real-time oil-tea fruit image to be detected into the model for feature extraction and multi-scale feature fusion, and then output the confidence, bounding box and coordinate results of the oil-tea fruit to be detected by the head, thereby realizing the detection of the spatial and temporal position of the oil-tea fruit.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and the instructions are used to be read by a machine so as to enable the machine to execute the detection method according to any one of claims 1 to 8.