Intelligent tomato maturity identification system and method based on olfaction and vision bimodal fusion

By using an intelligent recognition system that integrates olfactory and visual modalities, combining visual images and olfactory gas sensors, the problems of unstable accuracy and environmental interference in tomato ripeness recognition have been solved, achieving efficient and accurate ripeness recognition that is suitable for large-scale production.

CN121505598AActive Publication Date: 2026-02-10CHINA AGRI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511698339.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-10
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Existing methods for identifying tomato maturity suffer from problems such as high subjectivity, low efficiency, susceptibility to environmental factors, unstable accuracy, and insufficient versatility. Single-modality identification methods are insufficient to meet the needs of high precision and large-scale production.

Method used

An intelligent recognition system employing olfactory and visual dual-modal fusion, through cross-modal spatiotemporal alignment and network-level fusion strategies, combines visual images and olfactory gas sensors to achieve accurate identification of tomato ripeness.

Benefits of technology

It improves the accuracy and robustness of tomato maturity identification, reduces the impact of environmental interference, meets the real-time and stability requirements of large-scale production, and has the capability for large-scale implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505598A_ABST
    Figure CN121505598A_ABST
Patent Text Reader

Abstract

The invention relates to the field of agricultural equipment, and discloses an intelligent tomato maturity identification system and method based on olfaction and vision bimodal fusion. The system comprises an image acquisition module 1, a gas acquisition module 2, a gas distribution visualization module 3, a background tomato deep separation module 4, an RGB image and gas modal fusion feature network 5 and a maturity detection output display 6. The method comprises the following steps: synchronously collecting visual images and surrounding gas of tomatoes; respectively extracting visual and olfactory features; the bimodal features are fused; and identifying the maturity level by using a classification model based on the fusion features. According to the method, appearance and fruit surface gas information are fused, the defects that a single visual mode is observed on one side and is prone to color cheating and environment interference are overcome, more accurate and more robust intelligent identification of tomato maturity is achieved, and the method is suitable for agricultural automatic picking and sorting.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of agricultural equipment, and in particular to a tomato maturity intelligent recognition system and method based on olfaction and vision dual modal fusion. BACKGROUND

[0002] Tomato is an important fruit and vegetable widely planted and consumed globally, and its maturity directly determines the taste, nutritional value, shelf life and market value. Accurate recognition of tomato maturity is a key link to realize precise picking, grading and packaging, cold chain transportation and market circulation of tomatoes, and is of great significance to improve the economic benefits and market competitiveness of the tomato industry.

[0003] Currently, tomato maturity recognition methods mainly include manual recognition and single modal machine recognition. Manual recognition relies on the experience of operators to identify the color, size and smell of tomatoes. However, this method has obvious defects: on the one hand, it is highly subjective, and different operators have different standards, resulting in unstable recognition accuracy; on the other hand, it is inefficient and difficult to meet the needs of large-scale industrial production, and long-term work can easily lead to operator fatigue, further reducing recognition accuracy. Single modal machine recognition methods mainly include visual recognition and olfactory recognition. Visual recognition method collects image information of tomatoes, extracts color, texture, shape and other features, and uses machine learning or deep learning model for maturity classification. This method has the advantages of non-contact and high efficiency, but is greatly affected by environmental factors such as light intensity and shooting angle, which can lead to inaccurate feature extraction and affect recognition accuracy; at the same time, for tomato varieties with little color difference or tomatoes in the transition stage of maturity, it is difficult to achieve accurate recognition relying only on visual features.

[0004] Olfactory recognition method collects volatile gas components (such as ethylene and acetaldehyde) released by tomatoes through gas sensor array, and the types and concentrations of these gas components change with the maturity of tomatoes, and the maturity is recognized by analyzing the gas features. However, single olfactory recognition method also has shortcomings: gas sensors are easily disturbed by other volatile gases in the environment, leading to errors in feature extraction; and different varieties of tomatoes release different volatile gas components, making the model less universal, and the model needs to be retrained for different varieties, increasing the application cost.

[0005] To solve the problems of low accuracy, poor anti-interference ability and insufficient universality of single modal recognition method, the present application proposes a tomato maturity intelligent recognition system and method based on olfaction and vision dual modal fusion, which combines the advantages of the two modalities to achieve accurate and efficient recognition of tomato maturity. SUMMARY

[0006] The application belongs to the technical field of agricultural intelligent detection and equipment, and improves two bottlenecks of tomato maturity recognition in real production environment. First, the single mode method depending on visual images has a significant increase in false detection rate when the light changes, the shooting angle changes and the background is complex, and it is difficult to stably distinguish the mature transition stage samples. Second, the single mode method depending on electronic olfaction has insufficient stability and limited cross-species generalization ability when there is interference and volatile component difference between varieties. The application proposes an intelligent recognition scheme of deep cooperation of olfactory information and visual information, adopts a cross-modal space-time alignment and network-level fusion strategy, and improves the accuracy, robustness and real-time performance in the sorting line and post-harvest grading link.

[0007] The system of the application is composed of six functional units, namely an image acquisition module, a gas acquisition module, a gas distribution visualization module, a background tomato depth separation module, an RGB image and gas modal fusion feature network, and a maturity detection and output display unit. The six modules are time-coordinated in a unified rhythm on the same embedded computing platform to form a top-down closed loop process, and the process sequence is synchronous acquisition, cooperative preprocessing, multi-scale fusion, online determination, result display and data reporting. The system interface, clock source and cache management adopt a unified specification to ensure accurate alignment of cross-modal data within a millisecond time scale.

[0008] The device structure adopts an integrated sealed detection cavity. The cavity material is transparent acrylic, the inner wall is coated with an anti-reflection black film, an Intel D304 depth camera and a ring-shaped constant illumination light are fixedly installed on the top, the vertical distance from the lens to the sample is fixed at 40 cm, and the viewing angle and focal length are kept in a locked state after calibration. An electronic olfactory array composed of 8 metal oxide sensors is installed on the side wall of the cavity, and independent channels are set for ethylene, acetaldehyde, ethanol and maturity-related volatile substances. Temperature and humidity sensors are installed at the bottom of the cavity to perform environmental compensation. The main control unit is an embedded industrial computer responsible for model reasoning, cross-modal fusion and human-computer interaction; the microcontroller is responsible for camera triggering, sensor sampling and lighting control; the touch screen is used to display detection results, task status and historical records, and provides batch export and network reporting functions.

[0009] The gas distribution visualization module maps the 8-channel instantaneous concentration vector into a fixed-size 32x32 pseudo-color heat map. The mapping steps are linear normalization, channel rearrangement, bilinear interpolation and histogram equalization. The heat map is used as the spatial expression input of the downstream network on the olfactory side, so that the convolution operator can process the olfactory features. The background tomato depth separation module sets a single threshold distance limit of 0.35 m and a morphological filling scale of 0.1 m based on the distance map output by the depth camera, and only keeps the maximum connected domain within the target depth range in each frame, and removes the background and near-range debris, and outputs a binary mask to constrain the effective area of the visual branch.

[0010] Data acquisition and time alignment adopt strict synchronization strategy. The vision side acquires RGB frames and obtains depth frames at 640x640 resolution, and the frame rate is set to 30 frames per second. The olfactory side acquires 8-channel time series signals at a sampling rate of 100 Hz, and records temperature and humidity. The unified timestamp is generated and distributed by the master control unit. The synchronization submodule aggregates each frame of RGB image with the olfactory data of 1 s before and after it based on the timestamp, and generates a 3-channel olfactory tensor completely aligned with the image by using sliding window average and polynomial interpolation. The vision side preprocessing includes equal scaling, center cropping, geometric normalization, brightness and contrast enhancement, color temperature and saturation disturbance, mean and variance normalization. The olfactory side preprocessing includes temperature and humidity compensation, third-order Butterworth low-pass filter, drift correction, Z-score normalization and outlier truncation. All preprocessing uses the same work pace between the two modalities to ensure time consistency and statistical consistency.

[0011] The fusion feature network adopts a double-branch isomorphic structure. The vision branch calls the backbone and neck structure of YOLO v11 to output P3, P4 and P5 scale feature maps, and loads visual pre-training weights to accelerate convergence. The olfactory branch adopts the same convolution and CSP and FPN structure as the vision branch, and receives a 3-channel olfactory tensor at the input end. Multi-scale fusion performs shape alignment and channel alignment at each scale, concatenates the same scale features along the channel dimension, and then uses 1x1 convolution, batch normalization and SiLU activation to complete channel re-labeling and compression, followed by SE channel attention and lightweight spatial attention to improve the proportion of effective features and suppress noise response. The fused three-scale features are sent to the YOLOv11 detection head to complete classification and regression, and the post-processing stage uses non-maximum suppression to output maturity level and confidence, and performs spatial consistency verification according to the depth mask to reduce edge misjudgment.

[0012] The training strategy adopts end-to-end joint optimization. The training data consists of image, depth and olfactory triplets collected synchronously, with a batch size of 16, a training round number of 300, a base learning rate of 0.01, and a cosine annealing schedule with 5 warm-up rounds. The weight decay coefficient is set to 0.0005, the gradient clipping threshold is set to 1.0, and the exponential moving average is introduced to improve stability and evaluation consistency. Geometric enhancement is completely synchronized between the vision and olfactory branches, and color enhancement only acts on the vision branch. The loss function is the weighted sum of classification cross entropy and CIOU regression loss, with a weight ratio of 1 to 1. In the verification stage, the same preprocessing, alignment and fusion rules as in the training are followed, and the evaluation indicators include accuracy, recall, average precision, inference delay and throughput.

[0013] The online inference process is fixed to 5 steps. In step 1, the acquisition module starts and generates a synchronization timestamp. In step 2, the depth separation module retains the tomato region within the target depth range and outputs a binary mask. In step 3, the olfactory side updates the 1s sliding buffer, completes compensation, filtering and normalization, and generates a 3-channel olfactory tensor aligned with the current image frame. In step 4, the double-branch network completes feature extraction and multi-scale fusion and outputs the class and position at the detection head. In step 5, non-maximum suppression obtains the final maturity grade and confidence, the touchscreen synchronously displays the results with text labels and color labels, the detection record is written to the local database and reported through the wireless network. The single machine processing capacity reaches 10 samples per second, and the end-to-end delay is controlled within 100 ms.

[0014] The methodological process adopts a fixed timing operation. The operator places the sample in the center of the tray and waits for 60 seconds to stabilize the volatiles, then starts a detection task. The system completes the judgment according to the online inference process, and archives the maturity grade, confidence, corresponding image, olfactory heat map and key intermediate variables. An offline incremental training is performed once a day to update the model parameters, and the update data comes from the new samples and misjudgment samples of the day to maintain long-term accuracy.

[0015] The key parameters and boundary conditions are uniformly set. The depth threshold is 0.35 m, the morphological filling scale is 0.1 m, the RGB resolution is 640x640, the visual frame rate is 30 frames per second, the olfactory sampling rate is 100 Hz, the time aggregation window is 1 s, the olfactory pseudo-image size is 32x32, and the 1x1 convolution output channel at the fusion is 50% of the original splicing channel number. The training batch size is 16, the training round number is 300, the base learning rate is 0.01, the weight decay is 0.0005, the gradient clipping threshold is 1.0, and the loss weighting ratio is 1:1. The above parameters are fixed as default configurations in the delivery version, and are provided as read-only display and audit logs in the maintenance interface.

[0016] In the effect verification link, 5 tomato varieties and 4 maturity grades are selected to build a synchronous sample library, data under different time periods, different light intensities and different environmental temperature and humidity are collected, and cross-validation is used to give statistical results. The average accuracy of the dual-modal fusion model is significantly higher than that of the single visual model and the single olfactory model, the recognition stability of the sample in the transition stage is obviously enhanced, and high accuracy and high recall rate are maintained under low light intensity, shielding and methanol mixed gas interference. Continuous production line test shows that the throughput of the whole machine reaches the level of multiple samples per second, meeting the real-time requirements of batch online grading.

[0017] Compared with the prior art, the present application has three aspects of significant beneficial effects.

[0018] 1. It provides three fixed paths: end-to-end fusion, decision-level fusion, and feature-level fusion, covering three types of engineering scenarios: high precision, speed, and generality, with clear boundaries for deployment and switching.

[0019] 2. By leveraging the complementary advantages of vision and smell, the ability to distinguish between samples with similar colors and subtle morphological differences is enhanced, significantly improving overall accuracy and reducing the error rate during the maturation transition stage.

[0020] 3. Relying on deep separation, environmental compensation, and attention enhancement, it maintains stable output under complex conditions such as low light, textureless surfaces, occlusion, and stray gas interference. It completes automated online processing from acquisition to identification at a fixed pace, and has the ability to be deployed on a large scale and is traceable.

[0021] The system structure, algorithm flow, and key parameters provided in this invention are all fixed settings. Each module can independently complete its intended function and maintain coordination with a unified clock through a fixed interface when the system is cascaded. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a tomato ripeness intelligent recognition system and method based on the fusion of olfactory and visual modalities. Figure 2 This is a schematic diagram of the image acquisition module; Figure 3 This is a schematic diagram of the flow guiding structure of the gas sensor array; Figure 4 This is a front view of the gas sensor array flow guide structure; Figure 5 This is a flowchart illustrating the visual processing workflow. Figure 6 This is a schematic diagram of the olfactory processing flow. Detailed Implementation

[0023] This invention proposes an intelligent tomato ripeness recognition system and method based on the fusion of olfactory and visual modalities. The invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0024] like Figure 1 The embodiment of the present invention shown includes the following main working devices: image acquisition module 1, gas acquisition module 2, gas distribution visualization module 3, background tomato depth separation module 4, RGB image and gas mode fusion feature network 5, and maturity detection output display 6. Each system works together to complete the task.

[0025] (a) Hardware deployment 1. Detection chamber: made of transparent acrylic material, with a black light-absorbing film on the inner wall, a camera and supplementary light installed at the top, a sensor array fixed on the side, and a sample tray at the bottom; 2. Main control unit: NVIDIA Jetson Nano (4GB memory) runs the fusion algorithm and recognition model, and STM32F103 controls the peripherals; 3. Interface configuration: INTEL D304 (camera), I2C (sensor), UART (WiFi / display), HDMI (touchscreen).

[0026] (II) Software Implementation 1. Development environment: Python 3.8, OpenCV 4.5, TensorFlow 2.5; 2. Core code snippet: Feature layer fusion: joint_feature = np.hstack((visual_feature, olfactory_feature)) Decision-making level fusion: weighted_result = 0.89*svm_result + 0.86*rf_result End-to-end network: constructed using PYTORCH, with cross-attention layers achieving feature interaction through matrix multiplication.

[0027] (III) Experimental Verification 1. Experimental sample: 200 samples of each of the 5 tomato varieties (4 grades × 50 samples), for a total of 1000 samples; 2. Test conditions: normal light (5000 lux), strong light (6500 lux), weak light (3500 lux), environment containing 10 ppm methanol impurity gas; 3. Evaluation metrics: accuracy, processing speed, and interference resistance retention rate; 4. Results Analysis: All three fusion strategies are significantly better than single-modal fusion. End-to-end fusion is the best in terms of accuracy and anti-interference, while decision-level fusion is the best in terms of speed.

[0028] This embodiment is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A tomato ripeness intelligent recognition system and method based on olfactory and visual dual-modal fusion, characterized in that, include: Image acquisition module 1 is used to acquire visual images of the tomatoes to be detected; Gas sampling module 2 is used to collect gas samples from the space surrounding the tomato to be tested. The gas distribution visualization module 3, connected to the gas acquisition module, is used to analyze and spatially reconstruct the gas concentration data in the space surrounding the tomato. The background tomato depth separation module 4 is used to extract the target region and separate the background from the acquired tomato image. By introducing a deep learning semantic segmentation algorithm, it can accurately distinguish the tomato subject from the complex background and remove irrelevant interference information. The RGB image and gas modality fusion feature network 5, connected to the background tomato depth separation module and the gas acquisition module, is used to deeply fuse the extracted RGB visual features and gas sensing modality features. The maturity detection output display module 6, connected to the RGB image and gas modality fusion feature network, is used to classify and identify the maturity of the tomato based on the fused feature vector and output the visualized detection results.

2. The system according to claim 1, characterized in that, The dual-modal feature fusion module employs a feature-layer fusion strategy, directly concatenating the visual feature vector and the olfactory feature vector into a joint feature vector. The dual-modal feature fusion module also employs a decision-layer fusion strategy, including: a first classifier (for generating a first classification result based on the visual feature vector), a second classifier (for generating a second classification result based on the olfactory feature vector), and a fusion unit for fusing RGB image features with the distribution map features of the gas of interest converted by olfactory sensing to obtain the final maturity category. The dual-modal feature fusion module uses an end-to-end fusion network based on deep learning, which includes visual and olfactory branches, and performs feature interaction and fusion in the intermediate layers of the network.

3. The system and method according to claim 1, characterized in that, Includes the following steps: The system synchronously or quasi-synchronously acquires visual images of the tomato to be tested and surrounding gas samples; preprocesses the acquired visual images and extracts visual feature vectors (FV); preprocesses the acquired gas sensor data and extracts olfactory feature vectors (FO); fuses the visual feature vectors (FV) and olfactory feature vectors (FO) to obtain a fused feature vector (F_Fused); inputs the fused feature vector (F_Fused) into a pre-trained maturity classification model to obtain the maturity level of the tomato; and outputs the recognition result.

4. The method according to claim 3, characterized in that, The visual features include one or more combinations of color features, texture features, and depth features extracted based on deep learning; the olfactory features include one or more combinations of steady-state response values, transient response curve features, and projection features after feature dimensionality reduction processing of the gas sensor array; the maturity classification model is the DOUBLE BACKBONE YOLO11 network model.

5. The system according to claim 1, characterized in that, The image acquisition module 1 uses an Intel D304 depth camera and a ring-shaped constant brightness fill light. The distance from the lens to the sample tray is fixed at 40 cm, the RGB resolution is 640×640, the frame rate is 30 frames per second, and the intrinsic and extrinsic parameters are locked after one-time calibration at the factory.

6. The system according to claim 1, characterized in that, The image acquisition module 1 sets a distance threshold of 0.35m on the depth map and uses a morphological filling scale of 0.1m to retain only the largest connected region within the target depth range, forming a target binary mask for spatial constraints in subsequent detection.

7. The system according to claim 1, characterized in that, The background tomato depth separation module 4 and the RGB image and gas mode fusion feature network 5 are connected in series with 1×1 convolution, batch normalization and SiLU activation before the detection head, and SE channel attention is superimposed to improve the proportion of effective features and suppress noise response.

8. A method for visually recognizing tomato ripeness based on the system described in any one of claims 1 to 6, characterized in that, Including steps S1 to S5: S1 acquires RGB images at a resolution of 640×640 and 30 frames per second, and simultaneously acquires depth maps. S2, visual preprocessing is completed in the order of proportional scaling, center cropping, brightness enhancement, contrast enhancement, color temperature perturbation, saturation perturbation and mean-variance normalization. S3, set a distance threshold of 0.35 m on the depth map and use a morphological filling scale of 0.1 m to obtain a binary mask that only contains the target connected domain; S4 uses YOLO v11 trunk and neck to extract P3, P4 and P5 three-scale features and completes classification and regression at the detection head; S5 uses non-maximum suppression to output 4 levels of maturity and confidence, and uses the binary mask to perform spatial consistency verification, display, archive, and report.

9. The method according to claim 4, characterized in that, During the training phase, end-to-end joint optimization was adopted with a batch size of 16, 300 training rounds, a base learning rate of 0.01, cosine annealing scheduling with 5 warm-up rounds, a weight decay of 0.0005, a gradient clipping threshold of 1.0, and a loss function that is a weighted sum of classification cross-entropy and CIOU regression loss with a weight ratio of 1:

1. Exponential moving average was also enabled to stabilize the evaluation.

10. The method according to claim 4, characterized in that, The online inference single-machine processing capacity reaches 10 samples per second, with an end-to-end latency of no more than 100 ms. The output includes text labels and color labels and is archived together with RGB images, target binary masks, maturity levels, and timestamps.

Citation Information

Patent Citations

  • Method for identifying real-time position of inspection robot and evaluating tomato counting yield

    CN116681964A

  • Fruit detection method and device fusing multi-modal information

    CN119048887A

  • System and method for predicting fluid type and thermal maturity

    US20230054795A1

  • Cigar tobacco leaf harvesting maturity identification method and system based on integrated learning

    US20240013380A1