Banana spatiotemporal multi-modal quality prediction system and training method thereof
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]本发明缓解了现有香蕉时品质预测技术存在难以全面反映香蕉品质演化的内在机理,长时序预测过程中易产生误差累积,预测精度低和香蕉表面衰腐的局部差异大的问题
[0023]本发明所述的一种香蕉时空多模态品质预测系统及其训练方法是基于深度学习实现的,有效缓解了现有香蕉时品质预测技术存在难以全面反映香蕉品质演化的内在机理,长时序预测过程中易产生误差累积,预测精度低和香蕉表面衰腐的局部差异大的问题。具体有益效果包括:
Smart Images

Figure CN122548696A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent detection and preservation information technology for agricultural products, specifically to the field of intelligent detection and preservation information technology for bananas. Background Technology
[0002] With the continued growth of global banana consumption and the extension of cold chain logistics, the banana industry faces the challenge of rapid quality deterioration and high spoilage rates during post-harvest storage and transportation. Banana quality forecasting is a crucial process in blueberry preservation; it allows for adjustments to temperature, humidity, or atmospheric conditions before spoilage occurs, effectively reducing spoilage rates and extending shelf life. Therefore, achieving long-term, high-precision banana quality forecasting is essential for minimizing post-harvest losses, optimizing inventory turnover, and improving the overall efficiency of the supply chain.
[0003] Currently, banana quality prediction technologies mainly include: The first type is a quality assessment method based on single-modal images, which uses RGB images to extract features such as color and texture, and combines convolutional neural networks (CNN) or traditional machine learning models to classify or regress banana quality. This method can only capture the visible spectral information of the banana surface and cannot reflect the changes in internal maturity and biochemical indicators, making it difficult to fully reflect the internal mechanism of banana quality evolution.
[0004] The second category is image prediction methods based on time-series modeling. By introducing models such as ConvLSTM, 3D-CNN and Transformer, image sequences of bananas at different time points are modeled to achieve short-term decay trend prediction. However, error accumulation is prone to occur in long-term time-series prediction, leading to a decrease in prediction accuracy.
[0005] The third category integrates physical and chemical indicators such as weight and hardness, but it mostly stays at the level of simple feature splicing or shallow fusion, making it difficult to depict the local differences in the decay of banana surface. This results in significant deficiencies in the prediction results in terms of detailed expression and structural consistency, and insufficient spatial feature modeling capabilities.
[0006] In summary, existing banana quality prediction technologies suffer from several problems: difficulty in fully reflecting the intrinsic mechanisms of banana quality evolution, susceptibility to error accumulation during long-term prediction, low prediction accuracy, and significant local differences in banana surface decay. Summary of the Invention
[0007] This invention alleviates the problems of existing banana quality prediction technologies, such as difficulty in fully reflecting the intrinsic mechanism of banana quality evolution, easy accumulation of errors during long-term prediction, low prediction accuracy, and large local differences in banana surface decay.
[0008] This invention provides the following solution: Option 1: A spatiotemporal multimodal quality prediction system for bananas, the system comprising: The image data acquisition module is used to preprocess and extract spatial features from the time series images of the bananas to be detected to obtain visual spatial features. The physiological data acquisition module is used to preprocess and nonlinearly map the time series image of the banana to be detected according to the physiological data corresponding to the timestamp in sequence to obtain physiological characteristics. A multimodal data fusion module is used to fuse the visual spatial features and the physiological features through multimodal data to obtain global temporal features; The multimodal quality prediction module is used to obtain the banana prediction image and corresponding prediction physiological data by multi-task collaborative decoding of the global temporal features; The predicted physiological data include weight and stiffness.
[0009] Furthermore, in one embodiment of the present invention, the multimodal data fusion module includes: The multimodal spatiotemporal feature acquisition unit is used to splice the visual spatial features and the physiological features to obtain spliced features; it is also used to perform feature recalibration on the spliced features to obtain multimodal spatiotemporal features. The global temporal feature acquisition unit is used to dynamically evolve the multimodal spatiotemporal features through a long short-term memory network to obtain global temporal features.
[0010] Furthermore, in one embodiment of the present invention, the obtained multimodal spatiotemporal features for
[0011] in, for The adaptive weight tensor at time step, for The splicing features mentioned at the time, superscript Indicates the baseline state.
[0012] Furthermore, in one embodiment of the present invention, the dynamic evolution is as follows:
[0013] in, For activation function, For the current input weight matrix, The historical state weight matrix, This is the bias term vector.
[0014] Furthermore, in one embodiment of the present invention, in the multimodal quality prediction module, the multi-task collaborative decoding is... ,
[0015] Subscript Indicates the current prediction time. For predicting continuous quality indicators, For the predicted physical characteristic parameters, For the predicted hardness grade, This is the decoder weight matrix. It is a hidden state. This is the decoder bias vector. For the regression task weight matrix, This is the bias vector for the regression task.
[0016] Option 2: A training method for a spatiotemporal multimodal quality prediction system for bananas, comprising the following steps: Step S1: Obtain the training set, which includes the time series of banana images and the physiological data corresponding to the timestamps; Step S2: Based on the training set, train the banana spatiotemporal multimodal quality prediction system using gradient descent and multi-task joint loss function; The banana spatiotemporal multimodal quality prediction system is the banana spatiotemporal multimodal quality prediction system described in Scheme 1.
[0017] Furthermore, in one embodiment of the present invention, the multi-task joint loss function is:
[0018] in, The total number of training samples. For the first The true continuous quality index of each sample For the first Continuous quality indicators predicted for each sample For the first The true physical characteristic parameters of each sample For the first The physical characteristic parameters predicted for each sample. For the first The true hardness grade of each sample For the first The predicted hardness level for each sample. These are the weighting coefficients for the physical feature loss. These are the weighting coefficients for the classification loss.
[0019] Furthermore, in one embodiment of the present invention, the gradient descent method is implemented based on the chain rule, wherein the chain rule is...
[0020] in, The set of learnable parameters for the model. To predict the target time, Let be the hidden state at time step t.
[0021] Option 3: An electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes a program stored in memory, it implements the method described in Scheme 2.
[0022] Option 4: A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method described in Option 2.
[0023] The present invention discloses a spatiotemporal multimodal banana quality prediction system and its training method, which is based on deep learning. This system effectively alleviates the problems of existing banana quality prediction technologies, such as difficulty in comprehensively reflecting the intrinsic mechanism of banana quality evolution, easy error accumulation during long-term prediction, low prediction accuracy, and large local differences in banana surface decay. Specific beneficial effects include: 1. The spatiotemporal multimodal quality prediction system for bananas described in this invention, through a multimodal spatial-temporal joint modeling mechanism, can simultaneously constrain the decay and evolution process of bananas at the pixel level, feature level, and physiological index level. It not only achieves high-precision prediction of the spread of decayed areas, but also accurately depicts the dynamic changes in weight loss and hardness decay, thereby significantly improving the reliability and interpretability of the model in actual fruit and vegetable preservation prediction tasks.
[0024] 2. The present invention provides a spatiotemporal multimodal quality prediction system for bananas, addressing the temporal continuity, spatial nonuniformity, and strong nonlinear degradation issues associated with banana storage and decay. Building upon traditional time-series image prediction, it introduces physiological indicators such as weight and firmness to construct a deep time-series prediction model with multimodal coupling, achieving deep fusion of image and physiological data. This fusion is not a simple feature stitching; its core challenges lie in the spatiotemporal asynchrony of heterogeneous data, the complexity of cross-modal nonlinear mapping, and the amplification of long-term time-series errors across modalities. Existing technologies mostly remain at a shallow stitching level and are limited by the scarcity of paired agricultural data, making deep coupling difficult. This invention accurately matches visual and physiological features through cross-modal attention, combines adaptive recalibration to dynamically balance bimodal weights, and incorporates physical consistency constraints to force both to conform to metabolic patterns, thus avoiding "fusion without integration" at the architectural level.
[0025] 3. The spatiotemporal multimodal quality prediction system for bananas described in this invention introduces an enhanced time series modeling mechanism to effectively capture nonlinear dynamic changes over a long period of time and reduce error accumulation.
[0026] 4. The spatiotemporal multimodal quality prediction system for bananas described in this invention improves the ability to characterize local decay regions through multi-scale spatial feature extraction and reconstruction strategies, thereby significantly improving the clarity, structural consistency and overall stability of the predicted image, and achieving high-precision prediction of the banana quality evolution process.
[0027] The method described in this invention is applicable to interdisciplinary fields such as computer vision, time series prediction modeling, multimodal data fusion, and non-destructive testing of agricultural product quality. Attached Figure Description
[0028] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a structural diagram of the banana spatiotemporal multimodal quality prediction system described in Implementation Method 1; Figure 2 It is the loss function curve described in Implementation Method Nine; Figure 3 This is the 24-hour PSNR result graph described in Implementation Method Nine; Figure 4 This is the result graph for 24 hours as described in Implementation Method Nine; Figure 5 This is the PSNR result graph for 48 hours as described in Implementation Method Nine; Figure 6 This is the result graph of 48 hours as described in Implementation Method Nine; Figure 7 This is the PSNR result diagram for 72 hours as described in Implementation Method Nine; Figure 8 This is the result graph after 72 hours as described in Implementation Method Nine; Figure 9 This is the PSNR result graph for 96 hours as described in Implementation Method Nine; Figure 10 This is the result graph after 96 hours as described in Implementation Method Nine. Detailed Implementation
[0029] Various embodiments of the present invention will now be clearly and completely described with reference to the accompanying drawings. The embodiments described with reference to the drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0030] Implementation Method 1: The spatiotemporal multimodal quality prediction system for bananas described in this implementation method, such as... Figure 1As shown, the banana spatiotemporal multimodal quality prediction system includes: The image data acquisition module is used to preprocess and extract spatial features from the time series images of the bananas to be detected to obtain visual spatial features. The physiological data acquisition module is used to preprocess and nonlinearly map the time series image of the banana to be detected according to the physiological data corresponding to the timestamp in sequence to obtain physiological characteristics. A multimodal data fusion module is used to fuse the visual spatial features and the physiological features through multimodal data to obtain global temporal features; The multimodal quality prediction module is used to obtain the banana prediction image and corresponding prediction physiological data by multi-task collaborative decoding of the global temporal features; The predicted physiological data include weight and stiffness.
[0031] In this embodiment, the image data acquisition module includes preprocessing such as size normalization, numerical standardization, and noise reduction filtering. The preprocessing is used to eliminate interference from illumination changes and imaging noise.
[0032] In this embodiment, the image data acquisition module is further used to segment the image time series of the banana to be detected based on the HSV color space model using computer vision methods to obtain normal areas and browned areas; it is also used to obtain the browned area ratio based on the normal areas and browned areas; the browned area ratio is used to quantitatively evaluate the influence of external intervention factors such as blue laser irradiation on the browning process of bananas.
[0033] In particular, due to the color characteristics of bananas in the HSV color space, which exhibit "hue shift towards brown, saturation reduction, and brightness decrease," the HSV color space is not only used to calculate the browning area ratio, but also serves as a key pre-optimization method for visual feature extraction. By decoupling color and brightness information, it effectively eliminates the interference of storage environment light fluctuations on image quality and significantly improves the segmentation accuracy of browning areas. At the same time, the quantized values of hue (H) and saturation (S) channels can be directly used as explicit features reflecting maturity and decay, supplementing the semantic interpretability of high-dimensional features of convolutional neural networks, and providing a physical consistency verification basis for the "visual browning rate - physiological decay rate" for subsequent multimodal fusion modules, ensuring that the prediction results conform to the actual decay evolution law of bananas.
[0034] In this embodiment, in the image data acquisition module, the spatial feature extraction is performed by using a convolutional neural network (CNN) encoder to encode the preprocessed image sequence frame by frame, mapping the original pixels to a high-dimensional feature space, so as to fully characterize the microscopic spatial information such as the color evolution, texture degradation and rot spread of the banana surface.
[0035] In this embodiment, the preprocessing in the physiological data acquisition module is Z-score standardization, which is used to eliminate dimensional differences.
[0036] In this embodiment, the nonlinear mapping in the physiological data acquisition module is a nonlinear mapping of standardized weight and firmness data through a fully connected embedding module. It is not a simple mathematical transformation, but an implicit modeling of the physical process of banana "water loss-softening": the fully connected layer can automatically capture the nonlinear coupling relationship between weight loss and firmness decay through multi-layer weight adjustment (e.g., the nonlinear feature of slow firmness change in the early stage of fruit water loss and sudden firmness drop near the decay stage). After mapping scalar physical quantities into high-dimensional feature vectors, each dimension can correspond to a specific physiological state (such as water migration rate, cell wall degradation degree, etc.), so that the model can learn the actual decay law of "how much firmness decreases for every 1% decrease in weight" from the data level. The extracted deep semantic features can be directly aligned with visual features in spatial dimension, providing physiological priors that conform to agronomic logic for subsequent multimodal fusion, and avoiding the physical contradiction of "weight decreases but firmness remains unchanged" in the prediction results.
[0037] In this embodiment, the observation data of bananas on the discrete time axis in the image data acquisition module and the physiological data acquisition module are represented as follows:
[0038] in Indicates the first t The image of the banana to be detected is captured at all times. and These represent the weight and hardness at the corresponding time steps, respectively. t For time step.
[0039] The prediction task can be formalized as learning a nonlinear mapping from a high-dimensional spatiotemporal joint space to a future state space, given a historical multimodal sequence:
[0040] In this embodiment, the spatial feature extraction in the image data acquisition module includes multi-layer convolution operators and non-linear activation functions.
[0041] The multi-layer convolution operator is used to map the time series image of the banana to be detected to a high-dimensional feature space.
[0042] In the aforementioned multi-layer convolution operator, for any time step t , No. The convolution operation can be represented as:
[0043] in Indicates the first Layer output feature map, and These represent the kernel weights and bias terms, respectively.
[0044] From a linear algebra perspective, this process can be equivalently represented as a linear transformation of the vectorized input:
[0045] in It is the sparse block Toeplitz matrix formed by expanding the convolution kernel.
[0046] The nonlinear activation function σ( ):
[0047] To enhance the model's ability to characterize nonlinear degradation patterns such as banana surface color degradation, rot spot diffusion, and texture blurring.
[0048] In this embodiment, the physiological data acquisition module projects the physiological data into a feature space by embedding a nonlinear mapping:
[0049] This process enables the vectorized expression of intrinsic quality changes such as fruit dehydration and tissue softening.
[0050] The banana spatiotemporal multimodal quality prediction system described in this embodiment includes an image data acquisition module, a physiological data acquisition module, a multimodal data fusion module, and a multimodal quality prediction module.
[0051] The image data acquisition module inputs a sequence of banana surface images acquired over continuous time. It performs size normalization and numerical standardization on the images at each time step to reduce interference from illumination variations and imaging noise. A convolutional neural network (CNN) is then used to encode the preprocessed images, mapping the original pixels to a high-dimensional feature space to fully characterize key spatial information such as color changes, texture degradation, and rot spread on the banana surface.
[0052] In addition, the image was segmented based on the HSV color space model, and computer vision methods were used to identify and separate the normal and browned areas in the image. The browning index was calculated to quantitatively assess the specific impact of external factors such as blue laser irradiation on the browning process of bananas.
[0053] The physiological data acquisition module uses weight and firmness data at corresponding time steps as auxiliary inputs for synchronous processing and eliminates dimensional differences through standardization. A fully connected embedding module performs feature mapping on the standardized weight and firmness data to extract intrinsic quality characteristics reflecting fruit dehydration and tissue softening patterns.
[0054] The multimodal quality prediction module generates banana decay prediction images for future time steps through image branching, achieving pixel-level prediction of appearance changes; and outputs weight and hardness prediction values for the corresponding time steps through regression branching, achieving quantitative prediction of internal quality changes.
[0055] The multimodal quality prediction module implements multi-task collaborative decoding and end-to-end optimization. It synchronously outputs banana decay prediction images at future moments and quantitative predictions of weight and hardness at the corresponding time steps, achieving unified joint prediction of appearance deterioration and internal quality.
[0056] Implementation Method Two: This implementation method further defines the spatiotemporal multimodal quality prediction system for bananas described in Implementation Method One. In this implementation method, the multimodal data fusion module includes: The multimodal spatiotemporal feature acquisition unit is used to splice the visual spatial features and the physiological features to obtain spliced features; it is also used to perform feature recalibration on the spliced features to obtain multimodal spatiotemporal features. The global temporal feature acquisition unit is used to dynamically evolve the multimodal spatiotemporal features through a long short-term memory network to obtain global temporal features.
[0057] In this embodiment, the splicing is
[0058] This implementation further defines the multimodal data fusion module and explains the multimodal data fusion method. This method employs a cross-modal attention fusion mechanism to deeply couple visual spatial features with physiological features, constructing a unified multimodal spatiotemporal feature representation that includes both appearance and internal quality information. This significantly enhances the model's ability to represent the intrinsic mechanisms of banana quality changes. Based on this, a time-series modeling mechanism (such as LSTM or Transformer) is introduced to recursively model the fused features in chronological order. This allows the banana spatiotemporal multimodal quality prediction system to gradually accumulate historical information, thereby profoundly depicting the long-term dynamic evolution of banana decay and quality changes.
[0059] Implementation Method 3: This implementation method further defines the spatiotemporal multimodal quality prediction system for bananas described in Implementation Method 2. In this implementation method, the obtained multimodal spatiotemporal features... for
[0060] in, for The adaptive weight tensor at time step, for The splicing features mentioned at the time, superscript Indicates the baseline state.
[0061] This embodiment further defines multimodal data fusion and provides an example of feature recalibration. This adaptive recalibration method introduces anisotropic scaling in the feature subspace to highlight decay areas and features related to quality changes.
[0062] Implementation Method Four: This implementation method further defines the spatiotemporal multimodal quality prediction system for bananas described in Implementation Method Two. In this implementation method, the dynamic evolution is...
[0063] in, For activation function, For the current input weight matrix, The historical state weight matrix, This is the bias term vector.
[0064] In this embodiment, from the perspective of continuous dynamics, the discrete process can be regarded as a numerical approximation of the following differential equation:
[0065] This provides the model with a stronger ability to explain time evolution.
[0066] This implementation further defines multimodal data fusion and explains the dynamic evolution. The method constructs a high-dimensional nonlinear dynamic system to characterize the time dependence and cumulative effects during the decay process. Based on historical context information, this method can recursively infer the hidden states of multiple consecutive future time steps, thereby effectively suppressing error accumulation in long-sequence prediction while ensuring temporal continuity. This implementation is not an abstract mathematical operation, but strictly follows the physical laws of banana postharvest physiology. Apparent characteristics of the real-time response at time t (such as abrupt changes in browning area). The item then recalls historical physiological inertia (such as the lag in the decrease of muscle firmness); activation function By introducing nonlinear saturation characteristics, the model accurately simulates the physical process of banana decay, which is "slow in the early stages and exponentially accelerated in the later stages," avoiding the counterintuitive result of "uniform decay" predicted by linear models. Through this recursive mechanism, the model mathematically binds scalar physiological data such as weight and hardness with the diffusion trend at the image pixel level in a physically meaningful way, thus enabling the hidden state to be... Each update is equivalent to a numerical simulation of the "dehydration-softening-browning" chain reaction of bananas, thus ensuring that the prediction results not only conform to the evolutionary logic of high-dimensional features, but also strictly follow the physical objective laws of post-harvest deterioration of agricultural products, and are practically feasible.
[0067] Implementation Method 5: This implementation method further defines the spatiotemporal multimodal quality prediction system for bananas described in Implementation Method 1. In this implementation method, the multi-task collaborative decoding in the multimodal quality prediction module is... ,
[0068] Subscript Indicates the current prediction time. For predicting continuous quality indicators, For the predicted physical characteristic parameters, For the predicted hardness grade, This is the decoder weight matrix. It is a hidden state. This is the decoder bias vector. For the regression task weight matrix, This is the bias vector for the regression task.
[0069] In this embodiment, the current prediction time T is in the range of [24h, 48h, 72h, 96h].
[0070] In this embodiment, the multi-task collaborative decoding is achieved through a linear mapping constructed by deconvolution or upsampling convolution.
[0071] This embodiment further defines the multimodal quality prediction module and describes the multi-task collaborative decoding. This method maps the latent features of bananas from the high-dimensional semantic space back to the pixel space through linear mapping, and realizes the joint prediction of images and physiological indicators through the multi-task decoding structure.
[0072] Implementation method six: The training method of the spatiotemporal multimodal quality prediction system for bananas described in this implementation method includes the following steps: Step S1: Obtain the training set, which includes the time series of banana images and the physiological data corresponding to the timestamps; Step S2: Based on the training set, the banana spatiotemporal multimodal quality prediction system is trained using gradient descent and a multi-task joint loss function. The training ends when the value of the multi-task joint loss function on the validation set does not decrease significantly for 20 consecutive rounds within the preset maximum training rounds, i.e., the change in the loss function is less than the set threshold of 0.005.
[0073] The banana spatiotemporal multimodal quality prediction system is any one of the banana spatiotemporal multimodal quality prediction systems described in Embodiments 1 to 5.
[0074] In this embodiment, the gradient descent method is as follows:
[0075] in, For the first The set of model parameters at the nth iteration represents the model's parameters at the nth iteration. k The current state after the end of the training round. For the first The set of model parameters after the next iteration represents the next state that the model will be updated to after this gradient calculation and backpropagation. The rate of change of the loss function at the current parameter position, The learning rate controls the step size for parameter updates. k This represents the current number of training iterations.
[0076] The training method described in this embodiment constructs a multi-task joint optimization strategy, designs a composite loss function, and utilizes a backpropagation mechanism to perform end-to-end collaborative updates of the banana spatiotemporal multimodal quality prediction system. This ensures both the realism of appearance predictions and the accuracy of intrinsic quality predictions, continuously improving the model's comprehensive predictive ability for banana decay trends. Consequently, it enhances the clarity and structural consistency of predicted images while maintaining high accuracy in intrinsic quality predictions.
[0077] Implementation Method Seven: This implementation method further defines the training method described in Implementation Method Six. In this implementation method, the multi-task joint loss function is...
[0078] in, The total number of training samples. For the first True continuous quality index of each sample For the first Continuous quality indicators predicted for each sample For the first The true physical characteristic parameters of each sample For the first The physical characteristic parameters predicted for each sample. For the first The true hardness grade of each sample For the first The predicted hardness level for each sample. These are the weighting coefficients for the physical feature loss. These are the weighting coefficients for the classification loss.
[0079] Implementation Method Eight: This implementation method further defines the training method described in Implementation Method Six. In this implementation method, the gradient descent method is implemented based on the chain rule, and the chain rule is...
[0080] in, The set of learnable parameters for the model. To predict the target time, For the first The hidden state at each time step.
[0081] This embodiment further refines the training method. Addressing the non-convex nature of banana quality prediction in a high-dimensional parameter space (i.e., the existence of numerous local optima, easily leading to the model falling into "pseudo-optima" and failing to accurately capture decay inflection points), a chain rule is employed to achieve temporal backpropagation of gradients. In the "high-noise, strongly temporally coupled" agricultural multimodal scenario involved in this invention, direct application encounters obstacles such as "gradient vanishing" and "physical feature distortion"—that is, as the prediction time increases, the gradient contribution of early key physiological characteristics (such as the initial water loss rate) is diluted by subsequent environmental noise, causing the model to fail to establish the causal logic of "physiological changes preceding apparent manifestations." Therefore, this invention provides a targeted adaptation of the standard chain rule to the physiological characteristics of bananas: a dynamic gradient scaling factor based on the banana respiration hopping characteristics is introduced during the chain multiplication process, forcibly and precisely back-allocating the final prediction error to each historical node according to time weights. This technique not only breaks the local optimum trap of high-dimensional non-convex space, enabling the model to quickly converge to the global optimum parameters even with small samples; more importantly, it endows the model with a temporal reasoning ability like a "metabolic clock," allowing the model to "remember" early physiological abnormalities like a real banana, thereby achieving early warning of sudden cold damage or mold growth.
[0082] Implementation Method Nine: The banana spatiotemporal multimodal quality prediction system used in this implementation method is based on the banana spatiotemporal multimodal quality prediction system described in Implementation Method One, combined with the banana spatiotemporal multimodal quality prediction systems optimized in Implementation Methods Two to Five.
[0083] The training method used in this embodiment is based on the training method of the banana spatiotemporal multimodal quality prediction system described in Embodiment 6, combined with the optimized training methods of Embodiments 7 and 8.
[0084] This implementation selects several different banana quality prediction systems from existing similar technologies, namely 3D-CNN, ConvLSTM, PredRNN, ST-ResNet, and Transformer. For example... Figure 2 As shown, different banana quality prediction systems exhibit significant differences in convergence speed and final loss values during training. 3D-CNN and ST-ResNet have relatively high initial losses, especially 3D-CNN, which fluctuates considerably in the early training stages, eventually converging to a loss of approximately 0.06–0.09, indicating limitations in modeling time-series dependencies. ConvLSTM initially drops rapidly before oscillating between 0.01 and 0.02, demonstrating some ability to capture long-term dependencies, but still being limited in spatial feature representation. The Transformer model starts with a low loss, but exhibits significant fluctuations during training, ultimately converging to a value between 0.044 and 0.065, showing relative stability in capturing long-range dependencies, but insufficient ability to model local spatial features. PredRNN shows a rapid decline in loss, eventually converging to the 0.03–0.05 range, but still shows significant fluctuations in some training stages. In contrast, this implementation exhibits a robust downward trend from the early stages of training and maintains smooth convergence throughout the training process, ultimately stabilizing the loss value at around 0.663. Compared to other models, it not only demonstrates lower volatility but also maintains consistency in multimodal information while capturing spatiotemporal features. Especially in the mid-to-late training stages, the loss curve of this implementation is smoother than that of 3D-CNN, ConvLSTM, ST-ResNet, and Transformer, indicating that the model has superior stability and robustness in balancing image spatial features and temporal dynamic evolution.
[0085] The Peak Signal-to-Noise Ratio (PSNR), an objective quantitative metric for measuring image reconstruction quality, was used to quantitatively analyze the 24-hour reconstruction performance of different banana quality prediction systems. This implementation consistently maintained a significant lead throughout the entire prediction sequence. In contrast, traditional spatiotemporal modeling methods such as 3D-CNN and ConvLSTM stabilized at approximately 19 dB and 19 dB, respectively, while PredRNN and ST-ResNet converged to approximately 17 dB and 22–23 dB, respectively, and the Transformer model remained around 18.5 dB. It can be seen that although ST-ResNet showed some improvement in later stages, its peak value (approximately 23 dB) was still significantly lower than the performance gain of approximately 15 dB or more achieved by this implementation. From a dynamic evolution perspective, this implementation exhibited faster convergence speed and higher reconstruction accuracy in the early stages of training, indicating that the constructed temporal modeling mechanism can more effectively capture the complex nonlinear evolutionary characteristics and spatial texture degradation patterns during banana decay. Meanwhile, in the later stages, the PSNR of this implementation fluctuates less, showing stronger stability and generalization ability, while other methods generally suffer from slow convergence or performance oscillation.
[0086] In the 24-hour spatiotemporal multimodal quality prediction of bananas, such as Figure 4 As shown, from the perspective of image visual quality, the predicted image output by this embodiment is closer to the real sample in terms of color, texture, and overall shape (True). Compared with comparative models such as 3DCNN, ConvLSTM, PredRNN, STResNet, and Transformer, the generated banana image has clearer edges, and the changes in spots and colors are more consistent with the natural decay process. In terms of weight and hardness, the weight (38.0) and hardness (1.3) given by this embodiment are not only closer to the true values (33.4 / 0.6), but also more accurate in terms of numerical values than other models, especially significantly better than ConvLSTM and Transformer models (weight deviation is as high as 2-3 times). Combined with the performance of the proposed method at up to 38.5 dB in the aforementioned PSNR data, it can be further confirmed that this embodiment can effectively capture the spatiotemporal coupling physical-visual evolution law of bananas during decay, and achieve high-precision attribute inference while suppressing information distortion. Thus, it shows excellent effectiveness and practical value in both image prediction and physical index regression, and has the dual advantages of high-fidelity reconstruction and physical attribute consistency.
[0087] Table 1 shows the performance differences of different models in the 24-hour banana rotting image prediction task. From the core metric PSNR (Peak Signal-to-Noise Ratio, a higher value indicates a higher pixel-level match between the reconstructed image and the real sample), this implementation significantly outperforms all the comparison models (such as ConvLSTM's 22.626dB and STRESNET's 19.477dB) with a score of 37.741dB, demonstrating its absolute advantage in restoring visual details such as banana browning diffusion and peel wrinkles. Among the error metrics, ConvLSTM had the lowest values for MAE (0.019) and RMSE (0.069) (small absolute error), but its MAPE (mean absolute percentage error) was as high as 2.160, indicating that the relative fluctuation of the prediction results was extremely large. This could be due to sample differences (such as the initial weight difference of different batches of bananas), leading to a serious imbalance between the predicted and actual values. In contrast, the MAE (0.026) and RMSE (0.073) of this implementation were slightly higher than ConvLSTM, but its MAPE was only 0.837 (the lowest), indicating that its relative error of the prediction results was more stable and its anti-interference ability was stronger. In terms of structural similarity (SSIM), ConvLSTM ranked first with 0.837, while this implementation was 0.750, still able to preserve image structural information well. Overall, this implementation, with its highest PSNR (best pixel-level accuracy) and lowest MAPE (most stable relative error), performed outstandingly in terms of multi-metric balance and practical application reliability, and was particularly suitable for agricultural scenarios that require accurate tracking of decay trends and guidance for preservation decisions.
[0088] Table 1. Experimental results of banana rotting image prediction over 24 hours.
[0089] The PSNR was used to quantitatively analyze the 48-hour reconstruction performance of different banana quality prediction systems, such as... Figure 5As shown, this implementation method further outperforms other models in continuous time steps. Specifically, in the prediction of the second day, the PSNR of this implementation method remains at a high level, increasing from approximately 35dB to 37.7dB, demonstrating higher pixel reconstruction accuracy compared to 3D-CNN, ConvLSTM, PredRNN, ST-ResNet, and Transformer. Furthermore, the PSNR curve of this implementation method shows a stable and slow upward trend, indicating that it can still stably capture the degradation features of the fruit surface even on the second day when the decay process accelerates. Other methods, such as ConvLSTM and 3D-CNN, exhibit lower PSNR starting points and slower growth, suggesting their susceptibility to noise interference under complex decay patterns. This implementation method maintains high accuracy and robustness in continuous time prediction, fully validating its effectiveness in banana decay image modeling and temporal prediction, providing a reliable visual feature basis for further weight and hardness prediction.
[0090] In the 48-hour spatiotemporal multimodal quality prediction of bananas, such as Figure 6 As shown, this implementation demonstrates significant effectiveness in both visual representation and physical property inference dimensions. From the visualization results... Figure 6 As can be seen, the predicted image output by this embodiment closely approximates the real sample in terms of shape contour, skin color, and local browning gradient, with clear edges and natural texture. In contrast, the predicted images from comparative models such as 3DCNN, ConvLSTM, PredRNN, STResNet, and Transformer generally suffer from problems such as blurred edges, texture collapse, or color deviation, failing to accurately reproduce the visual evolution of bananas during long-term decay. Regarding physical property prediction, the predicted values for weight and hardness from this embodiment deviate very little from the actual values. For example, the predicted weights are 38.0, 40.2, and 33.2, and the predicted hardness is 1.3, 0.5, and 0.5, which closely match the actual values (33.4 / 0.6, 33.2 / 0.5). In contrast, the weight prediction errors of the comparative models generally exceed 10%, with ConvLSTM and Transformer approaching 20%. The hardness prediction also deviates from the actual trend due to insufficient spatiotemporal correlation modeling. Especially in the 48-hour long-term prediction scenario, this implementation method effectively captures the nonlinear spatiotemporal coupling relationship of "weight-hardness-image texture", significantly suppresses information decay and error accumulation, maintains high fidelity and attribute accuracy, and demonstrates excellent robustness and generalization ability, fully verifying its practical value and theoretical superiority in long-term dynamic process prediction.
[0091] Table 2 shows the performance differences of different models in the 48-hour banana decay image prediction task. From the core metric PSNR (Peak Signal-to-Noise Ratio, a higher value indicates a higher pixel-level match between the reconstructed image and the real sample), this implementation (THIS PAPER) significantly outperforms all the comparison models (such as ConvLSTM's 22.397dB and ST-ResNet's 19.1dB) with a PSNR of 37.792dB, indicating its absolute advantage in restoring visual details such as banana browning diffusion and epidermal wrinkling under complex decay patterns in the mid-stage. Among error-related metrics, ConvLSTM has the lowest MAE (0.021) (smallest absolute error), but its MAPE (mean absolute percentage error) is as high as 2.073 (highest), indicating that the relative fluctuation of the prediction results is extremely large—for example, for banana samples with large initial weight differences, the ratio of the predicted value to the actual value may be severely unbalanced, resulting in insufficient reliability. In contrast, the MAE (0.0251) and RMSE (0.069) of this implementation are both at a low level, and the MAPE is only 0.856 (low), indicating that its prediction results have small absolute errors and strong relative stability, and outstanding resistance to sample difference interference. In terms of structural similarity (SSIM), ConvLSTM ranks first with 0.800, while this implementation is second best at 0.735, still able to preserve the overall structural information of the image well. Overall, this implementation, with its highest PSNR (best pixel-level accuracy) and stable MAPE (controllable relative error), performs outstandingly in terms of multi-metric balance and practical application reliability; although ConvLSTM has a low MAE, its excessively high MAPE makes it unsuitable for scenarios requiring stable ratio judgment.
[0092] Table 2. Experimental results of banana rotting image prediction after 48 hours.
[0093] The PSNR was used to quantitatively analyze the 72-hour reconstruction performance of different banana quality prediction systems, such as... Figure 7As shown, this implementation achieves higher PSNR values than 3D-CNN (approximately 11.91–19.10 dB), CONVLSTM (approximately 14.08–19.60 dB), PREDRNN (approximately 10.89–17.08 dB), ST-ResNet (approximately 9.32–23.26 dB), and Transformer (approximately 11.24–18.62 dB) overall. This trend indicates that this implementation can continuously capture the spatial texture and structural information of banana decay in long-term decay prediction, especially in complex decay regions, where the image reconstruction quality is significantly better than traditional temporal modeling methods and convolutional networks. Further observation of the PSNR curve reveals that this implementation has a smooth curve with a steady upward trend, indicating high stability and robustness in continuous temporal prediction. Moreover, compared with PREDRNN, this implementation shows a more significant performance improvement in the mid-to-late decay stages, with a PSNR difference exceeding 20 dB, demonstrating a stronger ability to model long-term nonlinear decay evolution patterns. Based on the experimental results combining weight and hardness measurement data, it can be inferred that this implementation method can effectively correlate physical indicators with image features, achieve multimodal information fusion, and thus improve the accuracy and reliability of decay prediction.
[0094] In the 72-hour spatiotemporal multimodal quality prediction of bananas, such as Figure 8 As shown, this embodiment compares the visual prediction results with the predicted values of physical properties (weight, hardness) of other mainstream models (3DCNN, ConvLSTM, PredRNN, STResNet, Transformer). In terms of visual quality, the predicted image generated by this embodiment is closest to the real sample (True) in terms of shape contour, skin color, and local browning distribution, with clear edges and natural texture. In contrast, the comparison models generally suffer from image blurring, loss of detail, or structural distortion (such as block artifacts in PredRNN, overall blurring in ConvLSTM and Transformer, and local deformation in STResNet). In terms of physical property prediction, the predicted values of weight and hardness in this embodiment (e.g., weight 38.0 vs. actual 39.5, hardness 0.4 vs. 0.5 in the first row; weight 40.2 vs. actual 33.2, hardness 0.5 vs. 0.5 in the second row) deviate significantly from the actual values compared to other models (e.g., ConvLSTM weight 32.39 / 32.60, Transformer 30.68 / 31.89, hardness 0.35 / 0.43, 0.57 / 0.38), demonstrating stronger physical consistency. Especially in 48-hour long-term prediction, this embodiment effectively suppresses information decay and error accumulation. Its prediction results are superior to the comparison models in terms of visual fidelity and attribute accuracy, fully verifying its effectiveness and robustness in modeling the spatiotemporal evolution of banana decay, and making it suitable for high-precision prediction tasks of long-term dynamic processes.
[0095] Table 3 shows the performance differences of different models in the 72-hour banana decay image prediction task. From the core metric PSNR (Peak Signal-to-Noise Ratio, a higher value indicates a higher pixel-level match between the reconstructed image and the real sample), this implementation (THIS PAPER) significantly outperforms all the comparison models (such as ConvLSTM's 22.8101dB and ST-ResNet's 19.543dB) with a PSNR of 38.348dB, indicating its absolute advantage in reconstructing complex visual details such as browning diffusion and epidermal wrinkling of bananas during the 72-hour long-term decay stage. Among the error metrics, this implementation has the lowest MAE (0.025) (better than ConvLSTM's 0.026), a low RMSE (0.071), and a much lower MAPE (0.919) than ConvLSTM's 2.304 (the highest), indicating that its prediction results have the smallest absolute error and the strongest relative stability—even with differences in the initial state of different batches of bananas, the ratio fluctuation between the predicted and actual values is controlled within a reasonable range. In terms of structural similarity (SSIM), ConvLSTM leads with 0.787, while this implementation achieves 0.729 (second best), still preserving the overall structural information of the image relatively well. In contrast, although PREDRNN is mentioned as "excellent," its actual PSNR is only 16.996dB (lowest of all), and its MAE (0.032) and MAPE (1.244) are both higher than this implementation, indicating a significant decrease in accuracy during long-term prediction. ST-ResNet, Transformer, and 3D-CNN all have PSNRs below 20dB and SSIMs below 0.7, making it difficult to capture the highly nonlinear texture changes of long-term decay. Overall, this implementation, with its highest PSNR (best pixel-level accuracy), lowest MAE (smallest absolute error), and controllable MAPE (small relative fluctuation), excels in multi-indicator balance and long-term prediction reliability, making it particularly suitable for practical agricultural scenarios that require accurate tracking of 72-hour decay trends and guidance for dynamic adjustments in storage.
[0096] Table 3. Experimental results of banana rot image prediction after 72 hours.
[0097] The PSNR was used to quantitatively analyze the 96-hour reconstruction performance of different banana quality prediction systems, such as... Figure 9As shown, this implementation maintains a continuously rising PSNR trend throughout the prediction sequence, eventually reaching approximately 37.45 dB, which is higher than 3D-CNN (approximately 19.25 dB), ConvLSTM (approximately 19.11 dB), STResNet (approximately 18.53 dB), Transformer (approximately 17.96 dB), and PredRNN (approximately 16.45 dB), demonstrating a significant improvement in pixel-level recovery capability.
[0098] Analysis of the curve features shows that this implementation method can consistently improve detail fidelity and structural similarity in the prediction of decayed images on the fourth day. The continuous increase in PSNR indicates that it can effectively capture the complex nonlinear spatial texture evolution during banana decay over long time. In contrast, traditional temporal prediction models such as CONVLSTM and PREDRNN have relatively flat PSNR curves with limited growth, indicating that they have bottlenecks in detail recovery and structural preservation when dealing with long-term prediction tasks involving high degradation and multimodal feature accumulation. While 3D-CNN, Transformer, and ST-ResNet show slight improvements at some time steps, their overall performance is significantly lower than the method presented in this paper, making it difficult to fully reflect the true evolution of the decayed region.
[0099] In the 96-hour spatiotemporal multimodal quality prediction of bananas, such as Figure 10As shown, this implementation demonstrates excellent effectiveness and robustness in long-term spatiotemporal evolution modeling. From a visual representation perspective, the predicted images generated by this implementation closely approximate the real sample True in terms of morphological contour, peel color, and local browning distribution. The banana's curved shape is natural, and the gradient transition between the golden yellow color and brown spots on the peel conforms to the long-term decay pattern. In contrast, comparative models such as 3DCNN, ConvLSTM, PredRNN, STResNet, and Transformer generally exhibit visual distortion. 3DCNN images are too dark and have blurred details, ConvLSTM and Transformer images are generally blurred and have lost edges, PredRNN shows blocky artifacts, and STResNet shows severe local deformation. None of these can accurately reproduce the natural decay characteristics of the banana after 96 hours. In terms of physical property prediction, the deviations between the predicted and actual values of weight and hardness in this embodiment are significantly smaller than those of the comparative models: In weight prediction, the errors of the first group (32.8) and the actual value (39.5) of the algorithm in this paper are about 16.9% and about 17.9% respectively, while the errors of the second group (33.33) and the actual value (40.6) are affected by long-term time-series uncertainties. However, the errors of the comparative models, such as 3DCNN (weights of 33.94 and 32.79) and Transformer (weights of 30.91 and 30.94), are significantly greater, exceeding 20%. In hardness prediction, the deviations of the first group (0.3 and 0.4) of this embodiment from the actual values (0.5 and 0.5) are smaller than those of 3DCNN (0.42 and 0.37) and ConvLSTM (0.49 and 0.44), and the overall trend is more in line with the physical law that hardness continuously decreases with decay. Especially in the 96-hour ultra-long-term prediction scenario, this implementation method effectively captures the nonlinear spatiotemporal coupling relationship between weight, hardness and image texture, significantly suppressing information decay and error accumulation. It outperforms the comparison model in both visual fidelity and attribute accuracy, fully verifying its practical value and theoretical superiority in long-term dynamic process prediction.
[0100] Table 4 shows the performance differences of different models in the 96-hour banana decay image prediction task. This implementation (THIS PAPER) leads in all core metrics: PSNR reaches 37.454dB (highest in the field, far exceeding PREDRNN's 16.447dB), SSIM is 0.743 (highest in the field), MAE is only 0.023 (lowest in the field), MAPE is 0.868 (lowest in the field), and RMSE is 0.067 (lowest in the field). This indicates that this method can not only accurately restore the visual details of banana decay after 96 hours (such as browning areas and epidermal wrinkles) in long-term prediction, but also control the prediction error to a minimum—both the absolute error (MAE / RMSE) and relative fluctuation (MAPE) are significantly better than other models. In particular, compared with PREDRNN (PSNR only 16.447dB, MAPE as high as 0.981), the pixel-level accuracy and structural consistency advantages of this implementation are extremely prominent.
[0101] Among the other models, ST-ResNet performed the worst (PSNR 19.008dB, SSIM 0.273), indicating that traditional residual convolution is difficult to capture the complex evolution of long-term decay. Although CONVLSTM has a low MAE (0.018), its MAPE is as high as 1.68 (the highest in the field), reflecting that its prediction results are highly volatile and its reliability is insufficient. 3D-CNN and Transformer performed moderately, but all indicators were inferior to this implementation.
[0102] Table 4. Experimental results of banana rot image prediction after 96 hours.
Claims
1. A banana spatio-temporal multi-modal quality prediction system, characterized in that, The banana spatiotemporal multimodal quality prediction system includes: The image data acquisition module is used to preprocess and extract spatial features from the time series images of the bananas to be detected to obtain visual spatial features. The physiological data acquisition module is used to preprocess and nonlinearly map the time series image of the banana to be detected according to the physiological data corresponding to the timestamp in sequence to obtain physiological characteristics. A multimodal data fusion module is used to fuse the visual spatial features and the physiological features through multimodal data to obtain global temporal features; The multimodal quality prediction module is used to obtain the banana prediction image and corresponding prediction physiological data by multi-task collaborative decoding of the global temporal features; The predicted physiological data include weight and stiffness.
2. The banana spatiotemporal multimodal quality prediction system according to claim 1, characterized in that, The multimodal data fusion module includes: The multimodal spatiotemporal feature acquisition unit is used to splice the visual spatial features and the physiological features to obtain spliced features; it is also used to perform feature recalibration on the spliced features to obtain multimodal spatiotemporal features. The global temporal feature acquisition unit is used to dynamically evolve the multimodal spatiotemporal features through a long short-term memory network to obtain global temporal features.
3. The banana spatiotemporal multimodal quality prediction system according to claim 2, characterized in that, The obtained multimodal spatiotemporal features for in, for The adaptive weight tensor at time step, for The splicing features mentioned at the time, superscript Indicates the baseline state.
4. The banana spatiotemporal multimodal quality prediction system according to claim 2, characterized in that, The dynamic evolution is as follows in, For activation function, For the current input weight matrix, The historical state weight matrix, This is the bias term vector.
5. The banana spatiotemporal multimodal quality prediction system according to claim 1, characterized in that, In the multimodal quality prediction module, the multi-task collaborative decoding is... , Subscript Indicates the current prediction time. For predicting continuous quality indicators, For the predicted physical characteristic parameters, For the predicted hardness grade, This is the decoder weight matrix. It is a hidden state. This is the decoder bias vector. For the regression task weight matrix, This is the bias vector for the regression task.
6. A training method for a spatiotemporal multimodal quality prediction system for bananas, characterized in that, Includes the following steps: Step S1: Obtain the training set, which includes the time series of banana images and the physiological data corresponding to the timestamps; Step S2: Based on the training set, train the banana spatiotemporal multimodal quality prediction system using gradient descent and multi-task joint loss function; The banana spatiotemporal multimodal quality prediction system is any one of the banana spatiotemporal multimodal quality prediction systems described in claims 1 to 5.
7. The training method for the banana spatiotemporal multimodal quality prediction system according to claim 6, characterized in that, The multi-task joint loss function is: in, The total number of training samples. For the first The true continuous quality index of each sample For the first Continuous quality indicators predicted for each sample For the first The true physical characteristic parameters of each sample For the first The physical characteristic parameters predicted for each sample. For the first The true hardness grade of each sample For the first The predicted hardness level for each sample. These are the weighting coefficients for the physical feature loss. These are the weighting coefficients for the classification loss.
8. The training method for the banana spatiotemporal multimodal quality prediction system according to claim 6, characterized in that, The gradient descent method is implemented based on the chain rule, which is: in, The set of learnable parameters for the model. To predict the target time, Let be the hidden state at time step t.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 6-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 6-8.