A Multimodal AI Intelligent Vegetable and Fruit Weighing Method and System

By collecting multimodal information of vegetables, vegetables and fruits by multiple sensor clusters and dynamically aligning and fusion, combining ICP algorithm and corruption integral equation, the problem of inaccurate shelf life prediction in the existing technology is solved, and high-precision shelf life prediction of vegetables, vegetables and fruits is achieved.

CN120030503BActive Publication Date: 2025-07-22四川参盘供应链科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510518038.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the prior art, it is difficult to fully obtain multimodal information of vegetables, vegetables and fruits, such as surface point cloud depth, weight, gas and temperature and humidity information, resulting in low accuracy in shelf life prediction and difficulty in effectively fusion of multimodal data.

Method used

Couple multiple sensors into a sensor cluster, collect multimodal information, and generate a unified timestamp sequence through interpolation synchronization, perform dynamic alignment and fusion, use ICP algorithm to generate vegetable and vegetable models, and build corruption integral equations for shelf life prediction.

Benefits of technology

High-precision identification and volume calculation of vegetable and vegetable varieties is achieved, and the accuracy and robustness of shelf life prediction are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030503B_ABST
    Figure CN120030503B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal AI intelligent weighing method and system for vegetables and fruits, belonging to the field of multi-modal data processing. The method includes: coupling multiple sensors into a sensor cluster to collect surface point cloud depth information, weight information, gas information, and temperature and humidity information of vegetables and fruits; interpolating and synchronizing the non-uniform sampling data in the sensor cluster to generate a unified timestamp sequence and generate joint representation data; generating a vegetable and fruit model through the ICP algorithm based on the data collected by the visual sensor array, identifying the type of the weighed vegetables and fruits through the generated model, and calculating the volume of the weighed item; constructing a spoilage integral equation according to the multi-modal features in the joint representation data, and realizing the prediction of the shelf life through this equation. The present invention can realize the prediction of the shelf life of the weighed vegetables and fruits, and improve the robustness and reliability of the prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal data processing, and particularly to a multimodal AI intelligent weighing method and system for vegetables and fruits. Background Art

[0002] In the prior art, a single sensor is usually used for detection, such as a vision sensor, a weight sensor, a gas sensor, etc. It is impossible to comprehensively obtain multimodal information of vegetables and fruits, such as surface point cloud depth information, weight information, gas information, and temperature and humidity information, resulting in low prediction accuracy. At the same time, it is difficult for the prior art to effectively fuse multimodal data and accurately predict the shelf life of vegetables and fruits. Summary of the Invention

[0003] One of the purposes of the present invention is to provide a multimodal AI intelligent weighing method for vegetables and fruits to solve the problem in the prior art that it is difficult to effectively fuse multimodal data and accurately predict the shelf life of vegetables and fruits.

[0004] The present invention is realized by the following technical solutions. A multimodal AI intelligent weighing method for vegetables and fruits includes the following steps: S100: Coupling multiple sensors into a sensor cluster to collect surface point cloud depth information, weight information, gas information, and temperature and humidity information of vegetables and fruits; S200: Interpolating and synchronizing the non-uniform sampling data in the sensor cluster to generate a unified timestamp sequence. The sensor data with the same unified timestamp sequence is used as an asynchronous multimodal data set, and the asynchronous multimodal data set is dynamically aligned and fused in the feature space to generate joint representation data; S300: Generating a 3D model of vegetables and fruits through the ICP algorithm based on the data collected by the vision sensor array, identifying the types of the weighed vegetables and fruits through the generated model, and calculating the volume of the weighed items; S400: Constructing a spoilage integral equation according to the multimodal features in the joint representation data, and predicting the shelf life through this equation.

[0005] Further, the weighing method for vegetables and fruits further includes step S500: Displaying the 3D model of the weighed vegetables and fruits through a visualization interface, and at the same time showing the volume included in the 3D model, the spoilage degree change curve drawn by the spoilage integral equation, and marking the final shelf life.

[0006] Further, coupling multiple sensors into a sensor cluster includes: After the weight sensor detects a weight change, other sensors are awakened; The vision sensor array is triggered by weight for scanning, and when the weight is detected to be stable, the gas sensor array is started; When the gas sensor array detects a sudden change in the ambient VOCs baseline, a secondary scan verification is triggered.

[0007] Furthermore, the sensor cluster includes: a visual sensor array, a gas sensor array, an environmental sensor array, and a weight sensor.

[0008] Furthermore, the visual sensor array is composed of ToF sensors or structured light sensors. ToF sensors are suitable for dynamic scenarios, and structured light sensors are suitable for static scenarios.

[0009] Furthermore, the visual sensor array may also include an RGB camera and an infrared spectral sensor for assisting in correcting the influence of light on point cloud color information.

[0010] Furthermore, the gas sensor array is composed of MOS gas sensors and VOC gas sensors, which are used for detecting the odor of fruits and vegetables and detecting metabolic gases such as ethylene, CO2, and volatile organic compounds (VOCs).

[0011] Furthermore, the environmental sensor array includes temperature and humidity sensors for compensating the temperature and humidity of the gas sensor data.

[0012] Furthermore, the joint characterization data is generated through the following sub-steps: S210. Perform dynamic time alignment. Assume that the sensor data is a multi-modal time series containing K sensor modalities, where the data of each modality is independently and asynchronously collected, and all data carry a unified global timestamp sequence. The original data representation of modality k is:

[0013] ,

[0014] , where is the set of original data of modality k, is the data point of modality k collected at timestamp t j R is the set of original data, d k is the vector dimension, is the set of valid timestamps of modality k, indicating that data is collected at these time points for this modality. T is the global timestamp, and t with subscript numbers and letters (t 1, t 2, t N ) represents a specific timestamp; S220. Construct a cross-modal time window, fill in the missing values, and generate a unified aligned feature tensor. The feature of modality k after alignment is:

[0015]

[0016] where is the modality feature after the global timestamp pair, is the length of the dynamic time window, is the weight of the weighted average, The value filled by the nearest neighbor, otherwise the nearest neighbor filling value; S230. Define a linear projection matrix for each modality k, and map the global timestamp to the feature space for the modality features after it through the following formula:

[0017] , where is the projection of modality k in the feature space, is the linear projection matrix, , d h is the vector dimension of the feature space, b k is the bias term; S240. Calculate the attention score of each modality k at time t i . Perform weighted summation on the features of each modality through the attention score, and calculate the weighted summation according to the following formula:

[0018] , where is the joint representation at time t i , is the attention score.

[0019] Furthermore, the weight of the weighted average in step S220 is calculated through the following formula:

[0020] , where exp is the exponential function, is the decay factor, used to control the influence of time proximity, and the closer the time, the higher the weight.

[0021] The attention score is calculated through the following formula:

[0022] , where Softmax is the function, is the transpose operation of vector v, is the hyperbolic tangent activation function, is the attention transformation matrix of modality k, is the context encoding at time t i , and V is the context transformation matrix.

[0023] Furthermore, the context encoding is calculated according to the following formula:

[0024] , where is the joint representation of the previous moment, and LSTM is the long short-term memory network, used to extract context features.

[0025] Furthermore, the volume of the weighed item is calculated through the following formula:

[0026] ,

[0027] wherein, the mass is calculated according to the data of the weight sensor, is the estimated volume, is the reference density.

[0028] Furthermore, step S300 further includes the multimodal constraint of the ICP algorithm objective function, and the multimodal constraint is represented by the following formula:

[0029] , wherein, is the objective function of the ICP algorithm, is the i-th point in the depth information of the surface point cloud, is the j-th point in the target cloud corresponding to , is the weight coefficient used to balance the influence of geometric error and multimodal similarity, is the similarity metric based on joint representation, which measures the similarity between the multimodal features of the corresponding points in the depth information of the surface point cloud and the depth information of the target surface point cloud, and are the joint feature vectors corresponding to and .

[0030] Furthermore, the corruption integral equation is constructed according to the following steps:

[0031] S410. Determine the influence factors according to the type and volume of the fruits and vegetables, and determine which multimodal features in the currently weighed type of fruits and vegetables will affect the corruption process, and assign weights to each feature;

[0032] S420. Describe the decay of the influence of the selected features over time through an exponential decay factor, and construct the integral equation from the initial time t0 to the current time t as shown in the following formula:

[0033] ,

[0034] ,

[0035] wherein, is the degree of corruption, is the weighted coefficient of the k-th modal feature, is the feature vector of the k-th mode at time t'; is the exponential decay factor, which is used to represent that as the time t - t' increases, the influence of past features on the current degree of corruption gradually weakens, e is the natural constant, and μ is the decay parameter; is the final shelf life, D threshold is the critical value of the degree of corruption.

[0036] On the other hand, the present invention provides a multi-modal AI intelligent vegetable and fruit weighing system, which includes a processor and a memory. A computer program is stored in the memory. When the computer program is executed by the processor, the above-mentioned multi-modal AI intelligent vegetable and fruit weighing method is realized.

[0037] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0038] 1. By coupling multiple sensors into a sensor cluster, the present invention comprehensively obtains multi-modal information of vegetables and fruits. At the same time, by interpolating and synchronizing non-uniformly sampled data collected by multiple sensors, a unified timestamp sequence is generated, and asynchronous multi-modal data sets are dynamically aligned and fused in the feature space to generate joint representation data, effectively solving the problem that it is difficult to fuse multi-modal data in the prior art.

[0039] 2. Based on the data collected by the visual sensor array and combined with the ICP algorithm, the present invention generates a high-precision and complete model, identifies the types of the weighed vegetables and fruits, and accurately calculates the volume of the weighed items, providing an important basis for shelf-life prediction.

[0040] 3. According to the multi-modal features in the joint representation data, the present invention constructs a spoilage integral equation, and realizes the prediction of the shelf life through this equation, improving the robustness and reliability of the prediction. Description of the Drawings

[0041] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0042] Figure 1 It is the flowchart of the method provided in Embodiment 1 of the present invention.

[0043] Figure 2 It is the flowchart of the multi-sensor coupling operation provided in Embodiment 1 of the present invention.

[0044] Figure 3 It is the timing diagram of the multi-sensor coupling cooperation sequence provided in Embodiment 1 of the present invention.

[0045] Figure 4 It is the algorithm timing diagram of the joint representation data generation provided in Embodiment 1 of the present invention.

[0046] Figure 5 It is the timing diagram of the optimization process of the ICP algorithm provided in Embodiment 1 of the present invention.

[0047] Figure 6 It is the flowchart of the shelf-life prediction of Step 4 provided in Embodiment 1 of the present invention.

[0048] Figure 7 This is the predicted timing diagram of the shelf life for Step 4 provided in Embodiment 1 of the present invention. Detailed implementation manners

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated herein generally may be arranged and designed in a variety of different configurations.

[0050] Embodiment 1

[0051] This embodiment discloses a multi-modal AI intelligent weighing method for vegetables and fruits. In the prior art, a single sensor is usually used to weigh and detect vegetables and fruits, and it is impossible to comprehensively obtain multi-modal information of vegetables and fruits, such as surface point cloud depth information, weight information, gas information, and temperature and humidity information, resulting in low prediction accuracy. And it is difficult for the prior art to effectively fuse multi-modal data and accurately predict the shelf life of vegetables and fruits. At the same time, when the prior art obtains image data and changes in temperature, humidity, and weight, it is impossible to consider introducing more advanced sensors and image acquisition devices to improve the accuracy and real-time performance of the data.

[0052] To solve the above problems, the multi-modal AI intelligent weighing method for vegetables and fruits disclosed in this embodiment couples multiple sensors into a sensor cluster for collecting multi-modal information of vegetables and fruits, interpolates and synchronizes the non-uniform sampling data collected by the multiple sensors to generate a unified time stamp sequence, and uses the sensor data with the same unified time stamp sequence as an asynchronous multi-modal data set, dynamically aligns and fuses the asynchronous multi-modal data set in the feature space to generate joint representation data. Then, a high-precision and complete model is generated through the ICP algorithm based on the data collected by the visual sensor array to identify the types of the weighed vegetables and fruits, and the volume of the weighed items is calculated. Finally, the system constructs a spoilage integral equation based on the multi-modal features in the joint representation data, and predicts the shelf life through this equation. Thus, the accuracy and robustness of the shelf life prediction of vegetables and fruits are improved, and the problem that it is impossible to effectively fuse multi-modal data and difficult to accurately predict the shelf life of vegetables and fruits in the prior art is solved.

[0053] It should be noted that the 'non-uniform sampling data' in this embodiment refers to the phenomenon of unequal intervals or irregularities between the data collected at different time points. Specifically, the data collection of sensors usually follows a certain time interval. However, if due to factors such as the response time of the sensors, differences in sensor performance, and external environmental impacts, the sampling times of multiple sensors are not completely synchronized, the time intervals (i.e., sampling intervals) between the data collected by these sensors will be inconsistent. Such data is called 'non-uniform sampling data'. In this embodiment, multiple sensors are coupled into a sensor cluster for collecting different types of information. These sensors may sample data at different time points due to various reasons, so the sampling data they generate may be non-uniform (i.e., the data is not collected at fixed intervals in time). In order to compare and fuse these data within the same time frame, it is usually necessary to interpolate and synchronize these non-uniform sampling data so that the data of all sensors can be mapped to a unified timestamp sequence. This is to ensure that in subsequent analysis and processing, all data can be aligned on the same time basis, so as to perform effective dynamic alignment and fusion to generate comprehensive joint characterization data. Simply put, 'non-uniform sampling data' refers to the situation where the time intervals for sensor data collection are uneven or irregular. In this case, interpolation or other processing is usually required to synchronize these data in time for subsequent analysis and fusion work.

[0054] Figure 1 The method flow chart of the multi-modal AI intelligent vegetable and fruit weighing method disclosed in this embodiment is shown. It can be seen from the figure that this embodiment includes the following steps:

[0055] Step 1: Couple multiple sensors into a sensor cluster for collecting the surface point cloud depth information, weight information, gas information, and temperature and humidity information of vegetables and fruits.

[0056] Specifically, the sensor cluster in this embodiment includes the following sensors:

[0057] An array of vision sensors, composed of ToF sensors or structured light sensors. ToF sensors are suitable for dynamic scenarios (such as scanning fruits and vegetables on a conveyor belt), and structured light sensors are suitable for static scenarios (such as manually placed fruits and vegetables). The vision sensors can also include RGB cameras and infrared spectrum sensors to assist with point cloud color information, correct the influence of light, and perform image segmentation to separate fruits and vegetables from the background.

[0058] The weight sensor obtains mass data for calculating density with volume and verifying the accuracy of the point cloud algorithm.

[0059] Gas sensor array, including MOS gas sensors and VOC gas sensors, is used for odor detection of fruits and vegetables, detecting metabolic gases such as ethylene, CO2, volatile organic compounds (VOCs), etc.

[0060] Environmental sensor array, including temperature and humidity sensors, is used for temperature and humidity compensation of gas sensor data.

[0061] Specifically, Figure 2 The multi-sensor coupling working flow chart is shown. It can be seen from the figure that the multi-sensor coupling working process is as follows:

[0062] After the weight sensor detects a weight change (fruits and vegetables are put in), other sensors are awakened.

[0063] The vision sensor array is triggered by weight for scanning. When the weight is detected to be stable, the gas sensor array is started. When the gas sensor array detects a baseline mutation of environmental VOCs (such as the generation of strange smells due to fruit and vegetable decay), a secondary scan verification is triggered. Figure 3 The timing diagram of the multi-sensor coupling cooperation sequence is shown.

[0064] Step 2: Interpolate and synchronize the non-uniform sampling data collected by multiple sensors to generate a unified timestamp sequence. The sensor data with the same unified timestamp sequence is used as an asynchronous multi-modal data set, and the asynchronous multi-modal data set is dynamically aligned and fused in the feature space to generate joint representation data.

[0065] Specifically, Figure 4 The algorithm timing diagram of the joint representation data generation in this embodiment is shown. It can be seen from the figure that the joint representation can be generated through the following steps:

[0066] 1) First, perform dynamic time alignment. Assume that the sensor data is a multi-modal time series, including K sensor modalities (such as ToF, weighing, gas), where the data of each modality is independently and asynchronously collected, but all data carry a unified global timestamp sequence. The original data representation of modality k is:

[0067] ,

[0068] ,

[0069] Among them, is the original data set of modality k, is the data point of modality k collected at timestamp t j R is the original data set, d k is the vector dimension, is the set of valid timestamps for modality k, indicating that data is collected for this modality at these time points. T is the global timestamp, and t with a subscript number and letter (t 1, t 2, t N ) represents a specific timestamp.

[0070] Then, construct a cross-modal time window, fill in the missing values, and generate a uniformly aligned feature tensor. The aligned feature of modality k is:

[0071]

[0072] Among them, is the modality feature after the global timestamp pair, is the dynamic time window length, is the weight of the weighted average, is the value filled by the nearest neighbor, and the nearest valid value can be read from the historical data.

[0073] Specifically, in this embodiment, the weight of the weighted average can be calculated by the following formula:

[0074] ,

[0075] Among them, exp is the exponential function, is the decay factor, which is used to control the influence of time proximity. The closer the time, the higher the weight.

[0076] 2) Map the features of each modality to a unified space to eliminate the dimension difference. Define a learnable linear projection matrix for each modality k, and map the modality feature after the global timestamp pair to the feature space through the following formula:

[0077] ,

[0078] Among them, is the projection of modality k in the feature space, is the linear projection matrix, , d h is the vector dimension of the feature space, b k is the bias term, indicating the feature projection bias of modality k.

[0079] 3) Calculate the attention score of each modality k at time t i , and perform weighted summation on the features of each modality through the attention score. Specifically, the weighted summation can be calculated according to the following formula:

[0080] ,

[0081] Among them, is at time ti Joint representation is the attention score.

[0082] In this embodiment, the attention score can be calculated by the following formula:

[0083] ,

[0084] where Softmax is a function, is the transpose operation of vector v, is the hyperbolic tangent activation function for non - linear transformation, is the attention transformation matrix of modality k, is at time t i context encoding. V is the context transformation matrix, , is the joint representation of the previous moment, and LSTM is the long short - term memory network for extracting context features.

[0085] Step 3: Generate a high - precision and complete model through the ICP algorithm based on the data collected by the visual sensor array, and calculate the volume of the weighed item through the sensor data.

[0086] At the same time, combine the RGB data (color histogram) in the joint representation data and the weight to estimate the volume, and the volume estimation is calculated by the following formula:

[0087] ,

[0088] where the mass is calculated according to the data of the weight sensor, is the estimated volume, is the reference density. There is a reference density table of common fruits and vegetables in the system, and the corresponding data can be obtained by looking up the table. Through the estimated volume, the point cloud is scaled and corrected using the volume ratio to accelerate the ICP convergence and reduce the convergence time of the ICP algorithm.

[0089] In order to enable the ICP algorithm to achieve point cloud modeling faster, a multi - modal constraint is introduced into the objective function of the ICP algorithm, so that the registered point cloud is not only geometrically aligned, but also consistent in multi - modal features (such as color, texture, etc.), thereby improving the robustness of the registration. In this embodiment, the multi - modal constraint is represented by the following formula:

[0090] ,

[0091] where, is the objective function of the ICP algorithm, is the i - th point in the source point cloud, is the corresponding point in the target cloud to The corresponding j-th point, is the weight coefficient used to balance the influence of geometric error and multimodal similarity, is the similarity measure based on the joint representation, which measures the similarity between the multimodal features of the corresponding points in the source point cloud and the target point cloud, and is and the corresponding joint feature vector.

[0092] It should be noted that Figure 5 shows the timing diagram of the optimization process of the ICP algorithm in this embodiment. It can be seen from the figure that the multimodal constraint in this embodiment combines the geometric error and multimodal similarity, so that the objective function not only considers the geometric alignment of the point cloud, but also considers the consistency of multimodal features. This method makes the point cloud registration more robust and accurate, especially in the application scenarios of fruits and vegetables with complex features. In addition, by adjusting the weight coefficient, the influence between geometric error and multimodal similarity can be flexibly weighed to optimize the registration effect.

[0093] Step 4: The system constructs a spoilage integral equation based on the multimodal features in the joint representation data, and realizes the prediction of the shelf life through this equation. Figure 6 shows the prediction flowchart of the shelf life in Step 4 of this embodiment. Figure 7 shows the prediction timing diagram of the shelf life in Step 4 of this embodiment. It can be seen from the figure that the spoilage integral equation can be constructed according to the following steps:

[0094] 1) First, determine the influencing factors according to the types and volumes of fruits and vegetables, and determine which multimodal features in the currently weighed fruit and vegetable types will affect the spoilage process, and assign weights to each feature.

[0095] 2) Describe the decay of the influence of the selected features over time through an exponential decay factor, and construct an integral equation from the initial time t0 to the current time t to accumulate the historical influence of all features. Through this integral equation, the end shelf life can be predicted.

[0096] In this embodiment, the integral equation is shown as follows:

[0097] ,

[0098] where is the degree of spoilage, is the weighted coefficient of the k-th modal feature, indicating the importance of this modal feature in the spoilage process. is the eigenvector of the k-th mode at time t'. Different modes can represent different environmental parameters or influencing factors, such as temperature, humidity, light, etc.

[0099] is the exponential decay factor, which is used to represent that the influence of past features on the current corruption degree gradually weakens as time t - t' increases. e is the natural constant, and μ is the decay parameter used to control the decay rate.

[0100] The end shelf life is expressed by the following formula:

[0101] ,

[0102] where, is the end shelf life, D threshold is the critical value of the corruption degree. When D(t) reaches or exceeds this value, it is considered to be corrupted and the end of the shelf life is reached. By this definition, it means to find the first time point T threshold at which D(t) reaches or exceeds the threshold D end , that is, the end of the shelf life of the product.

[0103] Step 5: Display the weighed vegetables and fruits through a real-time visualization interface, the 3D model reconstructed from the point cloud data, which includes the volume details of the model, the curve of the corruption degree changing with time drawn based on the corruption integral equation, and mark the end shelf life.

[0104] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-modal AI intelligent weighing method for vegetables and fruits, characterized in that, The vegetable and fruit weighing method includes: S100. Couple multiple sensors into a sensor cluster, collect the surface point cloud depth information, weight information, gas information, and temperature and humidity information of vegetables and fruits, and use the collected information as non-uniform sampling data; S200. Interpolate and synchronize the non-uniform sampling data to generate a unified timestamp sequence. Sensor data with the same unified timestamp sequence is used as an asynchronous multi-modal data set. Dynamically align and fuse the asynchronous multi-modal data set in the feature space to generate joint representation data; S300. Generate a 3D model of the vegetables and fruits through the ICP algorithm based on the data collected by the visual sensor array, identify the types of the weighed vegetables and fruits through the generated model, and calculate the volume of the weighed items; S400. Construct a spoilage integral equation based on the multi-modal features in the joint representation data, and predict the shelf life through this equation; The spoilage integral equation is constructed according to the following steps: S410. Determine which multi-modal features in the currently weighed vegetable and fruit type will affect the spoilage process according to the type and volume of the vegetables and fruits, and assign weights to each feature; S420. Describe the decay of the selected feature influence over time through an exponential decay factor, and construct an integral equation from the initial time t0 to the current time t as shown in the following formula: , , Among them, is the degree of corruption, is the weighting coefficient of the k-th modal feature, is the eigenvector of the k-th mode at time t'; is the exponential decay factor, which is used to represent that as the time t - t' increases, the influence of past features on the current corruption degree gradually weakens. e is the natural constant, and μ is the decay parameter; is the final shelf life, D threshold is the critical value of the degree of spoilage.

2. The multi-modal AI intelligent vegetable and fruit weighing method according to claim 1, wherein The vegetable and fruit weighing method further includes step S500: Display the 3D model of the weighed vegetables and fruits through a visualization interface, and at the same time display the volume included in the 3D model, the spoilage degree change curve drawn based on the spoilage integral equation, and mark the final shelf life.

3. The multimodal AI intelligent vegetable and fruit weighing method according to claim 1, wherein The coupling of multiple sensors into a sensor cluster includes: After the weight sensor detects a weight change, start the wake-up step; The visual sensor array is triggered by weight for scanning, and when the weight is detected to be stable, start the gas sensor array; When the gas sensor array detects a sudden change in the ambient VOCs baseline, trigger a secondary scan for verification.

4. The multimodal AI intelligent vegetable and fruit weighing method according to claim 1, wherein The joint representation data is generated through the following sub-steps: S210. Perform dynamic time alignment. Assume that the sensor data is a multi-modal time series, including K sensor modalities, where the data of each modality is independently and asynchronously collected, and all data carry a unified global timestamp sequence. The original data representation of modality k is: , , Among them, is the original data set of mode k, is the data point of mode k collected at timestamp t j , R is the original data set, d k is the vector dimension, is the set of valid timestamps of mode k, indicating that data is collected for this mode at these time points. T is the global timestamp, and t with subscript numbers and letters (t 1, t 2, t N ) represents a specific timestamp; S220. Construct a cross-modal time window, fill in the missing values and generate a unified aligned feature tensor. The aligned feature of modality k is: , Among them, is the modal feature after it at the global timestamp, is the dynamic time window length, is the weight of the weighted average, is the value filled by the nearest neighbor, otherwise it is the value filled by the nearest neighbor; S230. Define a linear projection matrix for each modality k, and map the global timestamp to the subsequent modality features in the feature space through the following formula: , wherein, is the projection of mode k in the feature space, is the linear projection matrix, , d h is the vector dimension of the feature space, b k is the bias term; S240. Calculate the attention score of each modality k at time t i and perform weighted summation on the features of each modality based on the attention score. The weighted summation is calculated according to the following formula: , Among them, is the joint representation at time t i , and is the attention score.

5. The multi-modal AI intelligent vegetable and fruit weighing method according to claim 4, characterized in that, The weights of the weighted average in step S220 are calculated through the following formula: , where exp is the exponential function, is the decay factor, which is used to control the influence of time proximity, and the weight is higher when the time is closer.

6. The multimodal AI intelligent vegetable and fruit weighing method according to claim 4, wherein The attention score is calculated through the following formula: , where Softmax is a function, is the transpose operation of vector v, is the hyperbolic tangent activation function, is the attention transformation matrix of modality k, is the context encoding at time t i and V is the context transformation matrix.

7. The multi-modal AI intelligent vegetable and fruit weighing method according to claim 1, characterized in that, The volume of the weighed item is calculated through the following formula: , wherein the mass is calculated based on the data of the weight sensor, is the estimated volume, is the reference density.

8. The multi-modal AI intelligent vegetable and fruit weighing method according to claim 1, characterized in that, Step S300 further includes multi-modal constraints on the ICP algorithm objective function. The multi-modal constraints are represented by the following formula: , Among them, is the objective function of the ICP algorithm, is the i-th point in the source point cloud, is the j-th point in the target point cloud corresponding to , is the weight coefficient used to balance the influence of geometric error and multimodal similarity, is the similarity measure based on joint representation, which measures the similarity between the multimodal features of corresponding points in the source point cloud and the target point cloud, and are the joint feature vectors corresponding to and .

9. A multimodal AI intelligent vegetable and fruit weighing system, characterized in that, The vegetable and fruit weighing system includes: A processor; A memory stores a computer program which, when executed by a processor, implements the multimodal AI intelligent vegetable and fruit weighing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal fusion detection method for freshness of pork tenderloin based on deep learning

    CN119691509A

  • Shelf-life monitoring sensor-transponder system

    US20050248455A1