Method and system for automatically segmenting and extracting hyperspectral information of camellia oleifera fruits in situ

Through the CA-TransUNet++ model and spectral clustering method, automatic segmentation and extraction of hyperspectral information of oleifera fruit is achieved, solving the problems of low efficiency and insufficient accuracy in the existing technology, and adapting to the real-time and scale needs of smart agriculture and forestry.

CN120298700APending Publication Date: 2025-07-11NANJING FORESTRY UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510459085.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the extraction of hyperspectral information of oil tea fruit depends on manual operation, is inefficient and has operator technical level and subjective judgment deviations, making it difficult to achieve large-scale real-time application in complex environments.

Method used

The CA-TransUNet++ model based on transfer learning is used to automatically segment the in-situ segmentation of the oleifera fruit, combined with a depth camera and a portable hyperspectral camera to acquire images, and automatically extract hyperspectral information through semantic segmentation and spectral clustering methods, including image preprocessing, model training and segmentation refinement.

Benefits of technology

It realizes rapid and efficient segmentation of fruits and extracts high-spectral information in complex field environments, improves efficiency several times and improves accuracy, and adapts to the real-time and large-scale application needs of smart agriculture and forestry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298700A_ABST
    Figure CN120298700A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for automatically segmenting and extracting hyperspectral information of oil tea fruits in situ, and the method comprises the steps: obtaining an RGB image and a hyperspectral image of the oil tea fruits in a natural environment, and constructing a source domain and target domain data set; performing pre-training on the CA-TransUNet + + model by using the source domain, and then introducing a target domain based on a pre-training weight to perform fine tuning so as to obtain a final camellia oleifera fruit semantic segmentation model; and then, performing refined segmentation on a fruit adhesion region in the segmented image through a spectral clustering algorithm, and combining a result with an original hyperspectral image to realize automatic extraction of hyperspectral information of the camellia oleifera fruits. According to the method, through the innovative CA-TransUNet + + image segmentation model, the camellia oleifera fruits can be rapidly and efficiently segmented, the hyperspectral information of the camellia oleifera fruits can be extracted, the linear decision coefficient (R2) of the method and a manual extraction result can reach 99.03%, and the processing time of a single fruit image is about 1-2 seconds. According to the method, the hyperspectral data of the camellia oleifera fruits can be efficiently and accurately acquired in a natural scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of in-situ image segmentation and hyperspectral information of forest fruits, and specifically to a method and system for automatically segmenting and extracting hyperspectral information of oil-tea fruit in-situ based on transfer learning, namely CA-TransUNet++. Background Art

[0002] With the rapid development of intelligent agriculture and forestry, the demand for precise and efficient management in agricultural and forestry production is continuously increasing. Real-time monitoring of the maturity and quality of forest fruits has become the core issue of concern in the industry. In the existing sensor technologies, hyperspectral technology, due to its unique advantage of combining spatial resolution and spectral resolution, can obtain fine spectral information on the surface and inside of objects, and shows great potential in the detection of agricultural and forestry products. Different from traditional imaging technologies, hyperspectral technology can not only identify the appearance features of objects, but also analyze the chemical composition and physiological state of target substances through spectral reflection characteristics. However, at present, the extraction process of hyperspectral information of oil-tea fruit mainly relies on manual operations, such as extracting target spectral features by manually calibrating the region of interest (ROI). This traditional method is not only inefficient, but also may bring significant deviations due to the technical level and subjective judgment of operators, limiting its promotion in large-scale real-time applications in complex environments.

[0003] The core goal of intelligent agriculture and forestry is to achieve precision, intelligence and high efficiency in agriculture and forestry, and the automatic extraction of hyperspectral data is an important technical support for achieving this goal. In the natural environment, there are many technical challenges in accurately extracting the spectral information of target objects from complex background vegetation or crop scenes. For example, in the field scene, the colors and textures of tree canopies, leaves and fruits are often similar, and the illumination changes and shadow effects will further increase the complexity of information extraction. The present invention focuses on oil-tea fruit as the target object, which has a more complex growth environment and a smaller fruit shape compared with common fruits such as apples and pears, and thus becomes a representative and challenging research object. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and system for automatically segmenting and extracting hyperspectral information of oil-tea fruit in-situ in view of the above-mentioned deficiencies of the prior art. The method and system for automatically segmenting and extracting hyperspectral information of oil-tea fruit in-situ can quickly and efficiently segment oil-tea fruit through an innovative semantic segmentation model, namely the image segmentation model of CA-TransUNet++, and extract its hyperspectral information; enabling hyperspectral technology to shift from "relying on manual labor" to "automation" to meet the urgent needs of intelligent agriculture and forestry for real-time and large-scale applications.

[0005] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0006] A method for automatically segmenting and extracting hyperspectral information of oil-tea fruit in-situ based on transfer learning CA-TransUNet++, including:

[0007] Step 1: Collect the RGB images and hyperspectral images of oil-tea fruit respectively, and perform black and white correction on the collected hyperspectral images.

[0008] Step 2: Take the three bands of the hyperspectral image after black and white correction as the R, G, and B bands of the synthetic false color image, and name it the HSI-RGB image; name the collected RGB image of oil-tea fruit as the RAW-RGB image.

[0009] Step 3: Perform semantic segmentation annotation on all HSI-RGB images and RAW-RGB images respectively to obtain a mask image dataset with oil-tea fruit markings.

[0010] Step 4: Take all RAW-RGB images and the corresponding mask images with oil-tea fruit markings as the RAW-RGB dataset, that is, the source domain; take all HSI-RGB images and the corresponding mask images with oil-tea fruit markings as the HSI-RGB dataset, that is, the target domain.

[0011] Step 5: Build an image segmentation model based on CA-TransUNet++, including a Transformer module, an encoder, a decoder, skip connections, and a coordinate attention module.

[0012] Step 6: Use the source domain to pre-train the image segmentation model based on CA-TransUNet++; retain the model weight parameters with the minimum loss function; then, on the basis of maintaining the pre-trained network weights, further train the image segmentation model based on CA-TransUNet++ through the target domain to obtain a trained oil-tea fruit segmentation and recognition model.

[0013] Step 7: Collect the in-situ hyperspectral image of the oil-tea fruit to be detected and perform black and white correction. Take the three bands of the in-situ hyperspectral image of the oil-tea fruit to be detected after black and white correction as the R, G, and B bands of the synthetic false color image to obtain the HSI-RGB image to be detected; use the trained oil-tea fruit segmentation and recognition model to perform preliminary segmentation on the HSI-RGB image to be detected; perform fine segmentation on the fruit adhesion area in the preliminary segmentation result through a spectral clustering method. After the fine segmentation is completed, obtain a mask containing only the oil-tea fruit image.

[0014] Step 8: Multiply the mask after fine segmentation in Step 7 by the hyperspectral image after black and white correction in Step 7, and then retrieve the hyperspectral data of the independent in-situ oil-tea fruit.

[0015] As a further improved technical solution of the present invention, step 1 is specifically as follows:

[0016] Under natural conditions, use a depth camera to capture the RGB image of the oil-tea camellia fruit, use a portable hyperspectral camera to capture the hyperspectral image of the oil-tea camellia fruit under natural conditions, and perform black and white calibration on the collected hyperspectral images:

[0017]

[0018] Among them, I C represents the calibrated hyperspectral image, I r represents the original hyperspectral image of the sample, I s represents the reflectance of the calibrated standard plate, I std represents the actual reflectance of the standard plate.

[0019] As a further improved technical solution of the present invention, step 2 is specifically as follows:

[0020] Use the 638nm, 551nm, and 460nm bands of the black and white calibrated hyperspectral image as the R, G, and B bands for synthesizing the false color image, named HSI-RGB; name the RGB image of the oil-tea camellia fruit captured on-site by the depth camera as RAW-RGB.

[0021] As a further improved technical solution of the present invention, step 3 is specifically as follows:

[0022] Use the intelligent annotation tool X-AnyLabeling to perform semantic segmentation annotation on all the HSI-RGB images and RAW-RGB images in step 2 respectively, to obtain an image dataset with labels, that is, a masked image dataset with fruit markings.

[0023] As a further improved technical solution of the present invention, transfer learning is a commonly used deep learning training strategy, which can transfer the weights of the source domain pre-trained model to the new target task. For this purpose, set a random seed and use the RAW-RGB dataset as the source domain, which is split into a training set and a test set at a ratio of 8:2; at the same time, use the HSI-RGB dataset as the target domain, which is divided into a training set (100 images), a validation set (40 images), and a test set (40 images).

[0024] As a further improved technical solution of the present invention, construct an image segmentation model based on CA-TransUNet++:

[0025] Based on the fine feature extraction of UNet++, the proposed CA-TransUNet++ model in this invention introduces Transformer to capture global spatial information, so as to more efficiently obtain the multi-scale and overall features of outdoor oil-tea fruits. This model consists of four core components: Transformer module, encoder, decoder, skip connection, and coordinate attention module.

[0026] The Transformer module consists of multi-head self-attention mechanism (MSA), multi-layer perceptron (MLP), and layer normalization (LN). The output of the l-th layer of Transformer can be expressed as:

[0027] z′ l =MSA(LN(z l-1 ))+z l-1 ;

[0028] z l =MLP(LN(z′ l ))+z′ l ;

[0029] where LN(·) is the layer normalization operator, and z l is the encoded image representation.

[0030] In the encoder part, this invention adopts a hybrid structure of CNN and Transformer to extract the features of oil-tea fruit images. CNN uses VGG16 as the backbone network, and its relatively shallow number of layers and good generalization ability help to accurately capture the local details of oil-tea fruits. To enhance the acquisition of global information, a Transformer-based module is added in the encoding stage: before the two-dimensional image sequence is input into the Transformer layer, the skip connection structure is improved, and features of different resolutions are fused through UNet++-Block to further improve the feature expression ability of the model.

[0031] U-Net++ introduces convolutional layers similar to the Dense structure on the basis of the direct connection of U-Net and fuses the features of the next-stage convolution. Through means such as deep supervision, nesting, and dense skip connections, U-Net++ can capture features at different levels and effectively improve the flexibility of the model. The output x i,j of the convolutional unit in the UNet++ network structure can be expressed by the following formula:

[0032]

[0033] where the function H(·) represents the convolution operation, the function U(·) represents the upsampling layer, and finally channel addition is used for fusion.

[0034] The decoder is designed to fuse multi-scale information and reconstruct the feature map: the hidden features output by the encoder are gradually transformed into feature maps with higher resolutions, and finally restored to the same size as the original image to form a high-quality segmentation result.

[0035] Coordinate Attention (CA) modules are introduced during the upsampling and downsampling processes, enabling the network to focus more on the effective regions, thereby alleviating the problems of edge blurring and interference of the fruits in complex backgrounds.

[0036] During the training phase, the network parameters are iteratively updated through gradient backpropagation, the Adam optimization algorithm, and automatic learning rate adjustment, and finally a model with excellent performance for oil-tea fruit image segmentation is obtained.

[0037] After training the dataset using the CA-TransUNet++ network, the model can accurately classify each pixel in the image as foreground (oil-tea fruit) or background.

[0038] In the fruit segmentation results, if there is a phenomenon of boundary adhesion, the spectral clustering method is used to make a finer distinction of these adhesion regions.

[0039] The spectral clustering method optimizes by solving the normalized graph cut, regarding the image pixels as nodes in a connected graph and performing spectral decomposition on the cut set. The objective function of the normalized graph cut is:

[0040]

[0041] where A and B are two subsets of the graph, V is the set of nodes containing all pixels, and W(u, v) is the weighted edge value connecting nodes u and v. The pixel intensity values I(u), I(v) and the parameter σ jointly determine:

[0042]

[0043] Spectral clustering performs eigen-decomposition on the graph Laplacian matrix, reduces the high-dimensional features, and then partitions the data.

[0044] When the gray gradient information in the graph is weak, the segmentation result is approximately a Voronoi segmentation. Thus, each independent fruit region can be accurately extracted from the complex and adhered fruit objects and combined with the original hyperspectral image to extract the spectral information of each independent fruit.

[0045] To comprehensively measure the oil-tea fruit segmentation effect, this invention selects three indicators: Mean Intersection over Union (MIoU), Mean Pixel Accuracy (MPA), and Dice Similarity Coefficient. Their formulas are as follows:

[0046]

[0047]

[0048] where n is the number of categories, p ii represents the number of correctly classified pixels of the i-th category in the confusion matrix; p ij represents the number of pixels whose i-th actual value is predicted as the j-th category; p ji represents the number of pixels whose j-th actual value is predicted as the i-th category. Through these metrics, the segmentation performance of the model can be comprehensively and accurately quantified.

[0049] The present invention compares the automatically segmented spectral results with the spectra of manually selected ROIs, and analyzes their correlation and consistency.

[0050] To achieve the above technical objectives, another technical solution adopted by the present invention is:

[0051] A system for automatically segmenting and extracting hyperspectral information of oil-tea fruits in-situ based on transfer learning CA-TransUNet++ includes:

[0052] An image acquisition module for respectively acquiring the RGB image and the hyperspectral image of the oil-tea fruit;

[0053] An image processing module for performing black and white correction on the acquired hyperspectral image; and using the three bands of the black and white corrected hyperspectral image as the R, G, and B bands for synthesizing a false color image, named the HSI-RGB image; naming the acquired RGB image of the oil-tea fruit as the RAW-RGB image;

[0054] A semantic segmentation module for respectively performing semantic segmentation annotation on all the HSI-RGB images and RAW-RGB images to obtain a mask image dataset with oil-tea fruit markings;

[0055] A dataset acquisition module for using all the RAW-RGB images and the corresponding mask images with oil-tea fruit markings as the RAW-RGB dataset, that is, the source domain; and using all the HSI-RGB images and the corresponding mask images with oil-tea fruit markings as the HSI-RGB dataset, that is, the target domain;

[0056] A training module for training the image segmentation model based on CA-TransUNet++ with the source domain and the target domain in sequence to obtain the final oil-tea fruit segmentation and recognition model;

[0057] A boundary adhesion segmentation module for using the trained oil-tea fruit segmentation and recognition model to perform preliminary segmentation on the to-be-detected HSI-RGB image, and performing refined segmentation on the fruit adhesion area by a spectral clustering method;

[0058] The spectral information extraction module is used to multiply the masked result after refined segmentation by the corresponding black-and-white corrected hyperspectral image, and then retrieve the hyperspectral data of the in-situ oil tea fruits independently.

[0059] In addition, the present invention also discloses an information data processing terminal, which can realize the functions of the above-mentioned CA-TransUNet++ in-situ oil tea fruit hyperspectral image segmentation and spectral information extraction network system, and provide feasible ideas and theoretical support for the application on other agricultural products such as citrus, apples and grapes.

[0060] The beneficial effects of the present invention are as follows:

[0061] In a complex field environment, the present invention develops a fully automatic hyperspectral information extraction method, which can quickly and efficiently segment fruits and extract their hyperspectral information. Through an innovative semantic segmentation model, the hyperspectral technology has shifted from "relying on manual labor" to "automation" to meet the urgent needs of smart agriculture and forestry for real-time and large-scale applications. This not only lays a foundation for the further popularization of hyperspectral technology in the agricultural and forestry fields, but also provides technical support for future precision agriculture and sustainable development. Specifically, the present invention uses a depth camera and a portable hyperspectral camera to collect oil tea fruit images in a natural environment, obtains the oil tea fruit mask in the hyperspectral image through a semantic segmentation model (i.e., the oil tea fruit segmentation and recognition model), and combines means such as spectral clustering and image mask decomposition. Finally, the mask is multiplied by the original corrected hyperspectral image to obtain the hyperspectral image of the oil tea fruit after background removal; subsequently, the corresponding hyperspectral information of the oil tea fruit is extracted by means of a spectral extraction algorithm. Compared with manual extraction of hyperspectral information, it has higher efficiency and higher accuracy.

[0062] Compared with the prior art, the present invention proposes a method and system for automatically segmenting and extracting hyperspectral information of in-situ oil tea fruits based on transfer learning. This method can accurately register each channel of the spectral image of oil tea fruits under complex field conditions and achieve foreground segmentation. It only takes 1-2 seconds to extract the spectral information of the fruits of a single tree, while the traditional manual method takes 5-10 minutes. The present invention can efficiently and accurately obtain the hyperspectral data of oil tea fruits in a natural scene.

[0063] The present invention proposes a CA-TransUNet++ model that integrates a coordinate attention module and a Transformer network. Transfer learning is introduced on a small-sample hyperspectral dataset, using VGG16 as the backbone feature extraction network and directly using double upsampling to make the height and width of the final output image equal to those of the input image. The segmentation time of a single-tree fruit image in the present invention is 1.2 s, and the MIoU, MPA, and Dice metrics reach 92.14%, 96.51%, and 95.81% respectively. Further combined with a spectral clustering method, the adhesion fruit area can be refined and segmented, and the fruit spectral information can be automatically extracted from the original hyperspectral image. Compared with the manual ROI extraction result, the spectral consistency between the two reaches R 2 = 99.03%, while significantly shortening the extraction time. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the required drawings. It should be noted that these drawings are only used to exemplarily show some embodiments of the present invention. For those skilled in the art, without creative efforts, other forms of drawings can still be obtained based on these drawings.

[0065] Figure 1 It is a schematic flowchart of a method for automatically segmenting and extracting hyperspectral information of oil-tea camellia fruits in situ based on transfer learning by the CA-TransUNet++ provided in the embodiment of the present invention.

[0066] Figure 2 It is a schematic diagram of oil-tea camellia fruit annotation provided in the embodiment of the present invention.

[0067] Figure 3 It is a model architecture diagram of an oil-tea camellia fruit image segmentation network CA-TransUNet++ provided in the embodiment of the present invention.

[0068] Figure 4 It is an architecture diagram of a coordinate attention module (Coordinate Attention, CA) provided in the embodiment of the present invention.

[0069] Figure 5 It is a comparison diagram of the effects of segmenting oil-tea camellia fruit images by different methods provided in the embodiment of the present invention.

[0070] Figure 6 It is a comparison diagram of the fruit adhesion segmentation edge by the spectral clustering algorithm provided in the embodiment of the present invention.

[0071] Figure 7 It is a comparison diagram of the correlation and consistency between automatic segmentation and manual extraction of spectral information provided in the embodiment of the present invention.

[0072] Figure 8This is the login interface diagram of the hyperspectral information system for in-situ automatic segmentation and extraction of oil-tea fruit based on transfer learning CA-TransUNet++.

[0073] Figure 9 In the hyperspectral information system for in-situ automatic segmentation and extraction of oil-tea fruit based on transfer learning CA-TransUNet++ provided by the embodiments of the present invention, this is the schematic mask diagram of in-situ segmentation of oil-tea fruit.

[0074] Figure 10 In the hyperspectral information system for in-situ automatic segmentation and extraction of oil-tea fruit based on transfer learning CA-TransUNet++ provided by the embodiments of the present invention, this is the curve diagram of oil-tea fruit number and spectral information. Detailed implementation manners

[0075] To more clearly elaborate the purpose, technical solution and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments. It should be noted that the described embodiments are only used to explain the principle of the present invention and do not constitute any limitation to it.

[0076] Aiming at the problems existing in the prior art, the present invention provides a network method and system for in-situ automatic segmentation and extraction of hyperspectral information of oil-tea fruit based on transfer learning CA-TransUNet++, including the following core steps: obtaining RGB images and hyperspectral images of oil-tea fruit in the natural environment, and constructing source domain and target domain data sets; using the source domain data set to perform pre-training on the CA-TransUNet++ model, and then introducing the target domain data set for fine-tuning based on the pre-training weights to obtain the final semantic segmentation model of oil-tea fruit; subsequently, performing refined segmentation on the fruit adhesion area in the segmented image through a spectral clustering algorithm, and combining the result with the original hyperspectral image to realize the automatic extraction of hyperspectral information of oil-tea fruit; the linear determination coefficient (R 2 ) of this method and the manual extraction result can reach 99.03%, and the processing time for a single fruit image is about 1-2 seconds. The present invention can efficiently and accurately obtain the hyperspectral data of oil-tea fruit in the natural scene.

[0077] The following content will describe the implementation scheme of the present invention in detail with reference to the accompanying drawings.

[0078] The present invention installs a portable hyperspectral imager (SOC710vp, Surface Optics Corporation Co., Ltd., USA) on a tripod, and uses Hyper Scanner_2.0.127 to collect hyperspectral images of oil-tea fruit. The distance between the hyperspectral lens and the periphery of the fruit tree is kept between 1 and 1.5 m, and a total of 180 hyperspectral images are collected, covering a variety of lighting conditions.

[0079] The RGB images of oil-tea fruits were captured using an Intel RealSense D435f depth camera. To enhance image diversity and model robustness, during the shooting, the camera was placed 1 - 1.5 m away from the tree crown, and target trees were randomly selected for shooting under different background and lighting scenarios. During the shooting process, due to changes in distance and perspective, the sizes, postures, and occlusion situations of the fruits varied. Eventually, a total of 1040 RGB images of oil-tea fruits were obtained.

[0080] To ensure the diversity of data samples, the image acquisition time covered the entire maturity period of the oil-tea fruits, and the acquisitions were carried out on September 21, 2023 (cloudy), October 8 (sunny), October 22 (cloudy), and November 5 (cloudy) respectively. The acquisition periods were from 9:00 - 12:00 am to 2:00 - 5:00 pm.

[0081] Each hyperspectral image contains a rectangular polytetrafluoroethylene (PTFE) plate (reflectivity 20%). To standardize the lighting conditions and reduce the influence of the instrument's dark current, the following formula was used to perform black-and-white correction on the images:

[0082]

[0083] where I C represents the corrected spectral image, I r represents the acquired hyperspectral image, I s represents the original reflectivity of the standard plate (20%), and I std represents the actual reflectivity obtained from the rectangular PTFE plate in the scanned image during acquisition. The PTFE plate was fixed on a tripod and was in the same scene as the fruit tree.

[0084] Construct a dataset for semantic segmentation:

[0085] From the original hyperspectral images, 638 nm, 551 nm, and 460 nm were selected and mapped to the R, G, and B channels of the synthetic false-color image respectively, thus obtaining 180 RGB images of oil-tea fruits, namely HSI-RGB images. Since the original images obtained by the depth camera had a relatively large resolution, which was not conducive to subsequent model training, a random cropping function written in Python was used to crop them to a size of 640×640, and fruitless or duplicate regions were removed. Eventually, 1000 RAW-RGB images of size 640×640 were obtained.

[0086] The intelligent annotation tool X-AnyLabeling was used to perform semantic segmentation annotation on all the images (including HSI-RGB images and RAW-RGB images), forming a labeled dataset, that is, a mask image dataset with oil-tea fruit markings. Figure 2Shows the annotation example process of oil-tea camellia fruits: (A) RAW-RGB image annotation process; (B) HSI-RGB image annotation process; (a) Select the region of interest; (b) Generate a mask image.

[0087] Manually extract the hyperspectral information of oil-tea camellia fruits: Select the bands with the highest reflectance (885.89 nm) and the lowest reflectance (417.49 nm) for band operation (subtract 417.49 nm from 885.89 nm) to obtain a grayscale image with significant differences between the fruits and the background; Subsequently, segment and construct a mask according to a 0.2 threshold to remove the background and irrelevant pixels, and only retain the region of interest of the oil-tea camellia fruits; Finally, extract the average spectrum of this region of interest as the reflection spectrum of the corresponding sample. The purpose of manual extraction is to compare with the automatic extraction method of the present invention to judge the consistency between the two.

[0088] To improve the performance of the segmentation network under small-sample hyperspectral data, the present invention introduces a transfer learning method:

[0089] Source Domain: The RAW-RGB dataset contains 1000 annotated RGB images of size 640×640 and the corresponding labeled data with semantic segmentation annotations. This dataset is larger in network scale and can provide sufficient training samples to pre-train the initial weights of the segmentation model.

[0090] Target Domain: The HSI-RGB dataset consists of 180 false-color images of oil-tea camellia fruits and the corresponding labeled data with semantic segmentation annotations. The present invention divides it into a training set (100 images), a validation set (40 images), and a test set (40 images).

[0091] First, on the source domain (RAW-RGB dataset), divide the training set and the test set in a ratio of 8:2, and use this larger-scale data to pre-train the CA-TransUNet++ model; Ensure the repeatability of the experiment by fixing the random seed. During network training, the input image resolution is fixed at 512×512, and the real-time saving strategy is used to retain the model weight parameters with the minimum loss function. Then, introduce the target domain (HSI-RGB dataset), and on the basis of maintaining the pre-trained network weights, complete the adaptation of the model on the hyperspectral synthetic image through further training, so as to obtain the final CA-TransUNet++ oil-tea camellia fruit image segmentation network. This method can effectively shorten the training time and significantly improve the segmentation accuracy of the model for oil-tea camellia fruit targets under small-sample conditions.

[0092] See Figure 3, the present invention is improved on the basis of the TransUNet network to form the CA-TransUNet++ network. This network consists of a Transformer module, a CNN encoder, a CNN decoder, skip connections, and a coordinate attention module.

[0093] From left to right, Figure 3 it can be split into three main parts: (1) the input image and feature extraction module (the upper part of the pink box, i.e., the VGG16 Block part); (2) the Transformer global feature processing module (the lower part of the pink box, i.e., the Transformer Layer); (3) the U-Net++ decoding module (the purple box); and finally, the segmentation or detection result is output. The left pink box (VGG16 + Transformer): VGG16 Block: The image is passed through multiple layers of convolution and pooling to gradually reduce the spatial resolution and extract deeper and deeper features. Transformer: On the obtained feature map, through linear projection and rearrangement, the features are converted into a sequence and modules such as multi-head attention (MSA) and feed-forward network (MLP) are used to model the global dependencies. Finally, the output of the Transformer is reshaped back into a 2D feature map (such as (20, 20, 512)) as the input for the subsequent "U-Net++" decoding. The right purple box (U-Net++ decoding): This part includes multi-level upsampling and skip connections, continuously restoring the spatial size of the feature map to the same as the original image (640×640). The "++" concept of U-Net++ is reflected in the multi-level fusion of skip connections, that is, each decoding stage will be concatenated with the corresponding encoded features to help restore details. At the same time, in the upsampling / downsampling process, attention modules (such as CA / Coordinate Attention) are often embedded in the model. The whole process from left to right is roughly: input image → pink box (including CNN and Transformer) → purple box (U-Net++ decoding) → output segmentation result.

[0094] Specific details: The original image is 696*520*3. After image preprocessing, the image is adjusted to 640×640×3 for model training. Enter the convolutional feature extraction module (VGG16 Block): This image will be fed into the first few or all feature extraction layers of a convolutional neural network (such as VGG16). VGG16 consists of several convolutional (Conv) and max pooling (MaxPool) layers: The image size will gradually decrease from 640×640 to 320×320, 160×160, 80×80, 40×40, and even 20×20, while the number of channels will gradually increase (such as 64, 128, 256, 512, etc.), extracting local features such as edges and textures from the shallow layers and more abstract semantic features from the deep layers. Linear mapping Transformer encoding: The final feature map of the CNN will be "unrolled" into a sequence form and input into the Transformer module. The key of the Transformer is to use multi-head self-attention to capture global dependencies, enabling the model to "see" the associations between different positions in the image features. Next, we reshape this "serialized" feature back to a 2D format, such as (20, 20, 512), for subsequent convolutional or decoding operations. Purple box (U-Net++ decoding) stage: The feature output of the above Transformer is usually a high-dimensional but low-spatial-resolution feature map (40, 40, 512), which is the "bottom layer" input of the U-Net++ decoder, that is, Figure 3 the left starting point of the purple box in the middle. U-Net++ multi-level upsampling and skip connections: Figure 3 The right half (purple box) is a decoding process of U-Net++, which is similar to the idea of the classic U-Net: Through successive upsampling (such as from 40×40 to 480×80, then to 320×320, and finally back to 640×640), combined with the features of each layer of the encoder for "skip connections", gradually restoring to the original resolution size. Upsampling (for example, from 40×40 to 480×80, then to 320×320, until back to 640×640). Make skip connections with the corresponding encoding-side features, and perform deeper dense connections under the idea of "++", with multiple intermediate layers for feature fusion, so as to make full use of the details of shallow features (such as textures and edges) and the semantic information of deep features, improving the segmentation accuracy. After each upsampling, fuse the concatenated features through a convolutional layer (such as 3×3 convolution + activation function) and adjust the number of channels as needed. After upsampling to the same size as the original image (640×640), pass through a 1×1 convolution or other forms of output layers to obtain the segmentation result.

[0095] Figure 4It is an architecture diagram of a Coordinate Attention (CA) module provided by an embodiment of the present invention; 1. Input: The input of the entire module is a feature map of (C, H, W) (the left green cube in the figure represents the input feature). Among them, C (Channel) represents the number of channels, H (Height) represents the height of the image or feature map, and W (Width) represents the width of the image or feature map. 2. X Avg Pool / Y Avg Pool: The module will perform average pooling in the "horizontal (X direction)" and "vertical (Y direction)" respectively: X Avg Pool: Pool (C, H, W) along the width direction into (C, H, 1); Y Avg Pool: Pool (C, H, W) along the height direction into (C, 1, W); This is done to aggregate the global information in the X and Y directions respectively to help the network capture coordinate information. 3. Concat&Conv2d: Concatenate (Concat) (C, H, 1) and (C, 1, W) to obtain a feature of (C, 1, H + W); then through one or several convolutional layers (Conv2d + BN + activation function), compress and non-linearly transform these concatenated features. A channel scaling ratio r (8) can be set in this step to reduce the channel dimension and thus reduce the computational amount. 4. Split&Channel Attention: After the convolutional output, it will be split into two branches: one branch is responsible for the attention mapping along the X direction (C×H×1), and the other branch is responsible for the attention mapping along the Y direction (C×1×W). After respective convolutions or activations, usually the weighted coefficients are output through Sigmoid. 5. Re-weight: Broadcast these two attention coefficients back to the corresponding dimensions of the input feature map respectively, and perform weighting on the channels and space: that is, multiply each channel or coordinate position in the feature by the corresponding attention weight, so that the network pays more attention to the useful regions and ignores the useless or noisy regions. The finally obtained output feature (the right green cube in the figure) is the new feature map after "attention weighting".

[0096] Specifically, coordinate attention aims to capture the long-range dependence information of the feature map in the horizontal and vertical directions respectively. Through nn.AdaptiveAvgPool2d for average pooling in the horizontal and vertical directions, followed by convolution and activation functions to complete channel compression and feature fusion, and generate direction-sensitive attention weights, which can greatly improve the focusing degree on the target area after multiplying with the original features.

[0097] Specifically, in U-Net++, different from the traditional U-Net which only contains simple "encoding - decoding" skip connections, U-Net++ adds several fusion modules and Dense-style convolutional operations in the skip connections between adjacent levels, and combines Deep Supervision to achieve dense interaction of features at different semantic levels.

[0098] The core idea is to refine each skip connection layer by layer, and capture finer edge and spatial information through successive concatenation and convolution. The convolutional unit output x of UNet++ i,j can be expressed by the following formula:

[0099]

[0100] where the function H(·) represents the convolutional operation, the function U(·) represents the upsampling layer, and then they are merged through channel addition operation; i represents the depth of the network layer, representing the number of downsampling times or the depth level of the network; j represents the number of convolutional nodes in the same layer (the serial number in the dense connection path); k represents the index of the feature channels (or the channel number dimension); x i,j represents the feature map output by the j-th node in the i-th layer of the network.

[0101] The encoder part uses VGG16 as the backbone network to extract multi-scale image features. Compared with deep networks, VGG16 has a moderate number of layers and good feature generalization ability.

[0102] The present invention adds a nested and dense skip connection structure between the encoder and the decoder, which not only retains delicate spatial information in the decoding stage, but also provides richer multi-scale features for the subsequent Transformer module.

[0103] Different from the traditional TransUNet encoder, the present invention inserts a CA module after the output of each convolutional block to enhance the attention to the key areas of oil-tea fruits.

[0104] In the decoder part, following the nested design of the U-Net++ skip connection, transposed convolution (TransposedConvolution) or upsampling operations are used to gradually restore the spatial resolution, and multi-scale context fusion is achieved through multiple "upsampling + concatenation + convolution + CA module". Finally, a feature map of the same size as the original image is output to obtain a fine segmentation result.

[0105] The present invention incorporates a Transformer module between the encoding and decoding structures to capture long-range dependencies and perform global context modeling on deep features.

[0106] The Transformer module consists of a multi - head self - attention mechanism (MSA), a multi - layer perceptron (MLP), and layer normalization (LN). The output of the l - th layer of the Transformer can be expressed as:

[0107] z′ l = MSA(LN(z l-1 )) + z l-1 ;

[0108] z l = MLP(LN(z′ l )) + z′ l ;

[0109] where MSA represents the multi - head self - attention mechanism, MLP represents the multi - layer perceptron; z l-1 is the output of the previous layer of the Transformer of l, z′ l is the output of the MSA plus the residual connection with z l-1 , LN(·) is the layer normalization operator, and z l is the encoded image representation.

[0110] The ViT class (the dimension cropping and input mapping process of the ViT class exactly corresponds to the processing process in the Transformer Layer module marked in purple in Figure 3 ) further crops the input dimension and maps the high - order feature map to the input size of the Transformer (usually the size of ) after 32 - fold downsampling.

[0111] Obtain the feature map using CNN and then convert it to a new embedding space through a trainable linear projection; subsequently, add the positional embedding (E pos ) to the patch embedding (x p ) to complete the spatial relationship modeling of the feature sequence, and the mathematical expression is as follows.

[0112]

[0113] where is the feature patch extracted by CNN, E is the slice embedding projection, E pos is the positional embedding, N is the patch number; z0 refers to the feature sequence (patch feature embedding + positional embedding) input to the Transformer.

[0114] Model training parameter settings:

[0115] To ensure the fairness of comparison between models, all models in the present invention are trained 100 times with a batch size of 16, and the Adam optimizer (learning rate 0.001, momentum 0.9, weight decay 0.0001) is adopted.

[0116] In the complex growth environment of oil-tea camellia fruits (such as foliage occlusion and fruit overlap), the phenomenon of fruit boundary adhesion often occurs during model prediction. To ensure the accuracy of subsequent spectral information extraction, a spectral clustering method is used to accurately segment the adhesion area.

[0117] The spectral clustering method realizes the precise distinction of fruit boundaries by solving the normalized graph cut problem, modeling the image pixels as a connected graph, and optimizing the cut set using spectral decomposition. The objective function of normalized graph cut is:

[0118]

[0119] where A and B are two subsets of the graph, and A and B refer to two non-overlapping subsets in the same graph (i.e., the image or network structure). These two subsets together constitute the node set V of the entire graph. Ncut(A,B) The objective function value that measures the connection tightness between two subsets A and B, B W(u,v) represents the sum of the weights (similarities) of the edges from node u ∈ A to all nodes v in node set B, A W(u,v) represents the sum of the weights (similarities) of the edges from node u ∈ B to all nodes v in node set A, V W(u,v) represents the sum of the weights (similarities) of the edges from node u ∈ A or u ∈ B to all nodes v in the entire node set V of the graph; W(u, v) is the weighted edge value connecting nodes u and v. The pixel intensity values I(u), I(v) and parameter σ are jointly determined:

[0120]

[0121] Spectral clustering realizes data partitioning by performing eigen decomposition on the Laplacian matrix of the graph to reduce the high-dimensional features.

[0122] The Laplacian matrix of the graph (Laplacian matrix) is constructed according to the similarity matrix (weight matrix) of the graph during spectral clustering. The specific method is as follows:

[0123] (1) First, according to the node similarity or connection weight W(u, v) defined by formula

[0104] :

[0124]

[0125] In this way, the similarity between all nodes in the graph is obtained, forming the adjacency matrix W of the graph, where each element of the matrix represents the weight value or edge strength between the corresponding two nodes.

[0126] (2) Construct the degree matrix:

[0127] Define the degree matrix D, which is a diagonal matrix, and each diagonal element D ii is the sum of the weights of all edges corresponding to node i in the adjacency matrix, that is:

[0128]

[0129] (3) Calculate the Laplacian matrix L:

[0130] The Laplacian matrix of the graph is defined as the difference between the degree matrix D and the adjacency matrix W:

[0131] L = D - W;

[0132] When used for normalized graph cuts, the normalized Laplacian matrix is usually adopted:

[0133]

[0134] When the gray gradient information in the graph is weak, the segmentation result is approximately the Voronoi segmentation. Thus, each independent fruit region can be accurately extracted from complex and adhered fruit objects, and the spectral information of the fruit can be extracted by combining the original hyperspectral image.

[0135] This article mainly uses semantic segmentation to first identify the fruits. After identification, other algorithms will be used to extract the spectral information inside, and finally, a comparison will be made with the manually extracted spectral information to see the degree of difference between the two.

[0136] To comprehensively measure the segmentation effect of the oil-tea fruit image, that is, to evaluate the segmentation performance of the semantic segmentation model (i.e., the oil-tea fruit segmentation and recognition model based on CA-TransUNet++), the present invention selects three commonly used and representative evaluation indicators: mean intersection over union (MIoU), mean pixel accuracy (MPA), and Dice similarity coefficient, and their expressions are as follows:

[0137]

[0138] Where n is the number of categories, and p ii represents the number of correctly classified pixels of the i-th category in the confusion matrix; p ij represents the number of pixels whose actual value of the i-th is predicted as the j-th category; p ji represents the number of pixels whose actual value of the j-th is predicted as the i-th category. Through these indicators, the segmentation performance of the model can be comprehensively and accurately quantified.

[0139] In this embodiment, the segmentation performance of CA-TransUNet++ is compared with that of FCN, DeepLabV3+, UNet, UNet++ and TransUNet. The results show that CA-TransUNet++ achieves the best performance in all metrics, and the MIoU, MPA and Dice reach 92.14%, 96.51% and 95.81% respectively. Table 1 summarizes the performance comparison of different semantic segmentation models on the test set.

[0140] Table 1. Performance comparison of different semantic segmentation models on the test set:

[0141]

[0142]

[0143] In addition, as Figure 5 shown, compared with the existing models, CA-TransUNet++ shows excellent segmentation consistency and accuracy under different lighting and fruit distribution conditions, demonstrating stronger generalization and robustness.

[0144] Due to the dense growth of fruits, fruit boundary adhesion often occurs after segmentation. The spectral clustering algorithm is introduced to finely segment the fruit boundary, Figure 6 intuitively showing the segmentation effect of this method in dealing with overlapping or closely connected fruits.

[0145] To verify the practicality of the predicted segmentation of CA-TransUNet++ in spectral extraction, the present invention compares the automatic segmentation results with the spectra of manually selected ROIs, and analyzes their correlation and consistency.

[0146] Figure 7 The visualization results of the comparison between the two methods are given. Figure 7 In (a) of [], it shows that the overall distribution of the spectral data obtained by automatic segmentation extraction and manual extraction in this embodiment is close to the ideal diagonal line (y = x), and the linear determination coefficient R 2 is as high as 99.03%. Figure 7 (b) is the Bland-Altman analysis of the two methods. Most of the data points fall within the 95% consistency interval, and the deviations are mostly concentrated around 0, indicating that the automatic and manual methods have a high degree of consistency in the quantitative results.

[0147] This embodiment also provides a system for automatically segmenting and extracting hyperspectral information of oil-tea fruit in-situ based on transfer learning CA-TransUNet++, including:

[0148] An image acquisition module for respectively acquiring the RGB image and hyperspectral image of the oil-tea fruit;

[0149] An image processing module, configured to perform black and white correction on the acquired hyperspectral images; and use the three bands of the hyperspectral images after black and white correction as the R, G, and B bands for synthesizing false color images, named HSI-RGB images; name the acquired RGB images of oil-tea camellia fruits as RAW-RGB images.

[0150] A semantic segmentation module, configured to perform semantic segmentation annotation on all HSI-RGB images and RAW-RGB images respectively to obtain a mask image dataset with oil-tea camellia fruit markings.

[0151] A dataset acquisition module, configured to use all RAW-RGB images and the corresponding mask images with oil-tea camellia fruit markings as the RAW-RGB dataset, i.e., the source domain; use all HSI-RGB images and the corresponding mask images with oil-tea camellia fruit markings as the HSI-RGB dataset, i.e., the target domain.

[0152] A training module, configured to train the image segmentation model based on CA-TransUNet++ in sequence with the source domain and the target domain to obtain a final oil-tea camellia fruit segmentation and recognition model.

[0153] A boundary adhesion segmentation module, configured to perform preliminary segmentation on the HSI-RGB images to be detected using the trained oil-tea camellia fruit segmentation and recognition model, and perform refined segmentation on the fruit adhesion area through spectral clustering method.

[0154] A spectral information extraction module, configured to multiply the refined segmentation mask result by the corresponding hyperspectral image after black and white correction, and then retrieve the hyperspectral data of independent in-situ oil-tea camellia fruits.

[0155] In addition, the present invention also discloses an information data processing terminal, which can implement the functions of the above-mentioned CA-TransUNet++ in-situ oil-tea camellia fruit hyperspectral image segmentation and spectral information extraction network system, and provide feasible ideas and theoretical support for the applications on other agricultural products such as citrus, apples, and grapes.

[0156] Furthermore, the above-mentioned system for automatically segmenting and extracting hyperspectral information of in-situ oil-tea camellia fruits based on transfer learning further includes an information display module, which can synchronously output the segmentation results of the in-situ automatic segmentation of oil-tea camellia fruits and the corresponding hyperspectral reflection data to a display device, and can display the segmentation results of oil-tea camellia fruits and spectral analysis diagrams through a visualization interface, such as Figures 8 - 10 shown. Figure 10 The average spectral curve diagram in, spectral curves for all regions are the spectral curves of all regions; the abscissa is the wavelength, and the ordinate is the average spectral value.

[0157] Specifically, the information display module is implemented on a terminal or a computer in a software-driven manner: the mask image output by the CA-TransUNet++ segmentation network, the spectral clustering result of the adhesion area, and the spectral image of the oil-tea fruit after multiplying with the original hyperspectral data are superimposed or juxtaposed for display, so as to visually present the position, shape, and boundary information of the oil-tea fruit in the hyperspectral image on the interface; meanwhile, the information display module performs visualization processing on the extracted spectral information of the oil-tea fruit, and can generate corresponding spectral curves in the graphical user interface (GUI) to realize the rapid viewing and analysis of the reflectance information of the oil-tea fruit in each band.

[0158] The display module highlights the segmented oil-tea fruit area on a computer terminal or a mobile device in a pseudo-color or boundary-labeled manner, forming a distinct contrast with the background (leaves, branches, or other sundries). Specifically, the system can use different colors or semi-transparent masks to identify the target fruit area, enabling users to clearly determine the segmentation accuracy when observing the image.

[0159] Based on the above segmentation visualization map, the information display module automatically detects the fruit area, outputs its corresponding hyperspectral reflection curve to the display window in the form of a chart, and marks the corresponding wavelength range (such as the visible light to near-infrared band) and its reflectance value. For the fruit area with occlusion or adhesion, the system can generate multiple spectral curves respectively after refined segmentation by spectral clustering to support users in viewing the spectral differences of single or multiple fruits.

[0160] This information display module can also save the automatically segmented fruit mask, spectral curve, and related segmentation evaluation indicators (MIoU, MPA, Dice, etc.) in a local or cloud database for subsequent large-scale data analysis and quality traceability. Users can also export the segmentation result map, spectral data table, or report document in the interface, supporting multi-scenario applications such as field management, quality monitoring, and scientific research analysis.

[0161] It should be noted that the various technical features listed in the above embodiments can be combined differently according to requirements. To make the text concise, each combination of technical features in the above embodiments is not described in detail. As long as there is no logical conflict in the combination of such technical features, it should be regarded as the effective record scope of this specification.

[0162] The above embodiments are only the preferred cases of the present invention, and the actual implementation manners of the present invention are not limited to the above description. Any operations such as modification, substitution, improvement, combination, or simplification carried out under the basic concept and technical idea of the present invention are equivalent deformation forms and are also included in the protection scope of the present invention.

[0163] The protection scope of the present invention includes but is not limited to the above embodiments. The protection scope of the present invention shall be subject to the claims, and any substitutions, deformations, and improvements that are easily conceivable by those skilled in the art to this technology shall fall within the protection scope of the present invention.

Claims

1. A method for in-situ automatic segmentation and extraction of hyperspectral information of oil-tea fruits based on transfer learning CA-TransUNet++, characterized in that, Including: Step 1: Respectively collect the RGB image and hyperspectral image of the oil-tea fruit, and perform black and white correction on the collected hyperspectral image; Step 2: Respectively use the three bands of the black and white corrected hyperspectral image as the R, G, and B bands for synthesizing a false color image, named as the HSI-RGB image; name the collected RGB image of the oil-tea fruit as the RAW-RGB image; Step 3: Perform semantic segmentation annotation on all HSI-RGB images and RAW-RGB images respectively to obtain a mask image dataset with oil-tea fruit markings; Step 4: Take all RAW-RGB images and the corresponding mask images with oil-tea fruit markings as the RAW-RGB dataset, that is, the source domain; Take all HSI-RGB images and the corresponding mask images with oil-tea fruit markings as the HSI-RGB dataset, that is, the target domain; Step 5: Construct an image segmentation model based on CA-TransUNet++; Step 6: Use the source domain to pre-train the image segmentation model based on CA-TransUNet++; retain the model weight parameters with the minimum loss function; then, on the basis of maintaining the pre-trained network weights, further train the image segmentation model based on CA-TransUNet++ through the target domain to obtain a trained oil-tea fruit segmentation and recognition model; Step 7: Collect the in-situ hyperspectral image of the oil-tea fruit to be detected and perform black and white correction. Respectively use the three bands of the in-situ hyperspectral image of the oil-tea fruit to be detected after black and white correction as the R, G, and B bands for synthesizing a false color image to obtain the HSI-RGB image to be detected; use the trained oil-tea fruit segmentation and recognition model to perform preliminary segmentation on the HSI-RGB image to be detected; perform refined segmentation on the fruit adhesion area in the preliminary segmentation result through the spectral clustering method. After the refined segmentation is completed, obtain a mask containing only the oil-tea fruit image; Step 8: Multiply the mask refined in Step 7 by the hyperspectral image after black and white correction in Step 7, and then retrieve the hyperspectral data of the independent in-situ oil-tea fruit.

2. The method for in-situ automatic segmentation and extraction of hyperspectral information of oil-tea fruits based on transfer learning CA-TransUNet++ according to claim 1, wherein, The specific content of Step 1 is as follows: Use a depth camera to take the RGB image of the oil-tea fruit in the natural environment, use a portable hyperspectral camera to take the hyperspectral image of the oil-tea fruit in the natural environment, and perform black and white correction on the collected hyperspectral image: Among them, I C represents the calibrated hyperspectral image, I r represents the original hyperspectral image of the sample, I s represents the reflectance of the calibration standard plate after calibration, I std represents the actual reflectance of the standard plate.

3. The method for automatically segmenting and extracting hyperspectral information in-situ of oil-tea fruits based on transfer learning CA-TransUNet++ according to claim 1, characterized in that, The specific content of Step 2 is as follows: Respectively use the 638nm, 551nm, and 460nm bands of the black and white corrected hyperspectral image as the R, G, and B bands for synthesizing a false color image, named as HSI-RGB; name the RGB image of the oil-tea fruit taken by the depth camera on-site as RAW-RGB.

4. The method for in-situ automatic segmentation and extraction of hyperspectral information of oil-tea fruits based on transfer learning CA-TransUNet++ according to claim 1, characterized in that, The specific content of Step 3 is as follows: Use the intelligent annotation tool X-AnyLabeling to perform semantic segmentation annotation on all HSI-RGB images and RAW-RGB images in Step 2 respectively to obtain an image dataset with labels, that is, a mask image dataset with fruit markings.

5. The method for automatically segmenting and extracting hyperspectral information in-situ of oil-tea fruits based on transfer learning CA-TransUNet++ according to claim 1, characterized in that, The image segmentation model based on CA-TransUNet++ in step 5 described above includes a Transformer module, an encoder, a decoder, skip connections, and a coordinate attention module; the Transformer module includes a multi-head self-attention mechanism MSA, a multi-layer perceptron MLP, and layer normalization LN.

6. The method according to claim 1, wherein In step 6 described above, on the source domain, the training set and the test set are divided in a ratio of 8:2; on the target domain, they are divided into a training set, a validation set, and a test set in a ratio of 5:2:

2.

7. A system for automatically segmenting and extracting hyperspectral information of oil-tea fruits in-situ based on transfer learning CA-TransUNet++, characterized in that, It includes: An image acquisition module for respectively acquiring the RGB image and the hyperspectral image of the oil-tea fruit; An image processing module for performing black-and-white correction on the acquired hyperspectral image; and using the three bands of the black-and-white corrected hyperspectral image as the R, G, and B bands for synthesizing a false color image, named the HSI-RGB image; naming the acquired RGB image of the oil-tea fruit as the RAW-RGB image; A semantic segmentation module for respectively performing semantic segmentation annotation on all HSI-RGB images and RAW-RGB images to obtain a mask image dataset with oil-tea fruit markings; A dataset acquisition module for using all RAW-RGB images and the corresponding mask images with oil-tea fruit markings as the RAW-RGB dataset, that is, the source domain; Using all HSI-RGB images and the corresponding mask images with oil-tea fruit markings as the HSI-RGB dataset, that is, the target domain; A training module for training the image segmentation model based on CA-TransUNet++ on the source domain and the target domain in sequence to obtain the final oil-tea fruit segmentation and recognition model; A boundary adhesion segmentation module for using the trained oil-tea fruit segmentation and recognition model to perform preliminary segmentation on the to-be-detected HSI-RGB image, and performing refined segmentation on the fruit adhesion area by a spectral clustering method; A spectral information extraction module for multiplying the refined segmentation mask result by the corresponding black-and-white corrected hyperspectral image, and then retrieving the hyperspectral data of the in-situ oil-tea fruit.

Citation Information

Cited By

  • Document identification method based on spectral analysis

    CN120808127A

  • A document authentication method based on spectral analysis

    CN120808127B