Multi-mode AI intelligent vegetable and fruit weighing method and system

By using multiple sensor clusters for multimodal data acquisition and fusion, and using ICP algorithm to generate 3D models of vegetables, vegetables and fruits, the problem of difficult to fusion of multimodal data in the prior art is solved, and the accurate prediction of the shelf life of vegetables and fruits is achieved.

CN120030503AActive Publication Date: 2025-05-23四川参盘供应链科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510518038.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-23
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate multimodal data and cannot accurately predict the shelf life of vegetables and fruits.

Method used

By coupling multiple sensors into a sensor cluster, multimodal information is collected, and a unified timestamp sequence is generated through interpolation synchronization, dynamic alignment and fusion are performed in the feature space to generate joint characterization data. Then, a 3D model of vegetables, vegetables and fruits is generated based on the data of the visual sensor array through the ICP algorithm, and the species is identified and the volume is calculated. Finally, the corruption integral equation is constructed based on the joint characterization data to achieve the prediction of shelf life.

Benefits of technology

It realizes the effective integration of multimodal data of vegetables, vegetables and fruits, and improves the accuracy and robustness of shelf life prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030503A_ABST
    Figure CN120030503A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode AI intelligent vegetable and fruit weighing method and system, and belongs to the field of multi-mode data processing, and the method comprises the steps: enabling a plurality of sensors to be coupled into a sensor cluster, and collecting the surface point cloud depth information, weight information, gas information and temperature and humidity information of vegetables and fruits; performing interpolation synchronization on the non-uniform sampling data in the sensor cluster, generating a unified timestamp sequence, and generating joint representation data; a vegetable and fruit model is generated through an ICP algorithm according to data collected by the visual sensor array, the type of the weighed vegetable and fruit is recognized through the generated model, and the volume of the weighed object is obtained through calculation; and according to the multi-modal characteristics in the joint representation data, constructing a corruption integral equation, and realizing the prediction of the shelf life through the equation. The method can achieve the prediction of the shelf life of the weighed vegetables and fruits, and improves the prediction robustness and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal data processing, and in particular to a multimodal AI intelligent vegetable and fruit weighing method and system. Background Art

[0002] The existing technology usually uses a single sensor for detection, such as visual sensors, weight sensors, gas sensors, etc., which cannot fully obtain multimodal information of vegetables and fruits, such as surface point cloud depth information, weight information, gas information, temperature and humidity information, etc., resulting in low prediction accuracy. At the same time, the existing technology is difficult to effectively integrate multimodal data and cannot accurately predict the shelf life of vegetables and fruits. Summary of the invention

[0003] One of the purposes of the present invention is to provide a multimodal AI intelligent vegetable and fruit weighing method to solve the problem in the prior art that it is difficult to effectively integrate multimodal data and cannot accurately predict the shelf life of vegetables and fruits.

[0004] The present invention is implemented by the following technical scheme, a multimodal AI intelligent vegetable and fruit weighing method, comprising the following steps: S100, coupling multiple sensors into a sensor cluster, collecting surface point cloud depth information, weight information, gas information, and temperature and humidity information of vegetables and fruits; S200, interpolating and synchronizing the non-uniform sampling data in the sensor cluster to generate a unified timestamp sequence, and the sensor data with the same unified timestamp sequence is used as an asynchronous multimodal data set, and the asynchronous multimodal data set is dynamically aligned and fused in the feature space to generate joint representation data; S300, generating a 3D model of vegetables and fruits according to the data collected by the visual sensor array through the ICP algorithm, identifying the type of vegetables and fruits to be weighed through the generated model, and calculating the volume of the weighed items; S400, constructing a corruption integral equation according to the multimodal features in the joint representation data, and realizing the prediction of the shelf life through the equation.

[0005] Furthermore, the vegetable and fruit weighing method also includes step S500, displaying a 3D model of the weighed vegetables and fruits through a visual interface, and displaying the volume contained in the 3D model, a corruption degree change curve drawn by a corruption integral equation, and marking the final shelf life.

[0006] Furthermore, coupling multiple sensors into a sensor cluster includes: the weight sensor wakes up other sensors after detecting a weight change; the visual sensor array triggers scanning based on weight, and starts the gas sensor array when the weight is detected to be stable; when the gas sensor array detects a sudden change in the environmental VOCs baseline, a secondary scan verification is triggered.

[0007] Furthermore, the sensor cluster includes: a visual sensor array, a gas sensor array, an environmental sensor array and a weight sensor.

[0008] Furthermore, the visual sensor array is composed of ToF sensors or structured light sensors. The ToF sensor is suitable for dynamic scenes, and the structured light sensor is suitable for static scenes.

[0009] Furthermore, the visual sensor array may also include an RGB camera and an infrared spectrum sensor for assisting the point cloud color information to correct the light effects.

[0010] Furthermore, the gas sensor array consists of MOS gas sensors and VOC gas sensors, which are used for odor detection of fruits and vegetables, detection of ethylene, CO 2 , volatile organic compounds (VOCs) and other metabolic gases.

[0011] Furthermore, the environmental sensor array includes temperature and humidity sensors for performing temperature and humidity compensation on the gas sensor data.

[0012] Furthermore, the joint representation data is generated through the following sub-steps: S210, dynamic time alignment is performed, assuming that the sensor data is a multimodal time series, including K sensor modalities, where each modality data is independently and asynchronously collected, and all data carry a unified global timestamp sequence. The original data representation of modality k is:

[0013] ,

[0014] ,in, is the original data set of mode k, At timestamp t j The collected data points of mode k, R is the original data set, d k is the vector dimension, is the valid timestamp set of mode k, indicating that the mode has data collection at these time points, T is the global timestamp, t plus subscript numbers and letters (t 1, t 2, t N ) represents a specific timestamp; S220, construct a cross-modal time window, fill in missing values ​​and generate a uniformly aligned feature tensor. The features of the aligned modality k are:

[0015]

[0016] in, To align the modal features after the global timestamp, is the dynamic time window length, is the weight of the weighted average, is the value filled by the nearest neighbor, otherwise is the value filled by the nearest neighbor; S230, define a linear projection matrix for each modality k, and map the modality features after the global timestamp to the feature space through the following formula:

[0017] ,in, is the projection of mode k in the feature space, is the linear projection matrix, , d h is the vector dimension of the feature space, b k is a bias term; S240, for each mode k, calculate its i The attention score of each modality is weighted and summed according to the following formula:

[0018] ,in, For time t i The joint representation of Score for attention.

[0019] Furthermore, the weight of the weighted average in step S220 is calculated by the following formula:

[0020] , where exp is the exponential function, It is a decay factor used to control the influence of time proximity. The closer the time, the higher the weight.

[0021] The attention score is calculated as follows:

[0022] , where Softmax is a function, is the transpose operation of vector v, is the hyperbolic tangent activation function, is the attention transformation matrix of modality k, For time t i Context encoding, V is the context transformation matrix.

[0023] Furthermore, the context encoding is calculated according to the following formula:

[0024] ,in, It is the joint representation of the previous moment, and LSTM is a long short-term memory network, which is used to extract contextual features.

[0025] Furthermore, the volume of the weighed object is calculated by the following formula:

[0026] ,

[0027] Among them, the mass is calculated based on the data of the weight sensor, To estimate the volume, is the reference density.

[0028] Furthermore, step S300 also includes a multimodal constraint of the objective function of the ICP algorithm, and the multimodal constraint is expressed by the following formula:

[0029] ,in, is the objective function of the ICP algorithm, is the i-th point in the surface point cloud depth information, For the target cloud The corresponding j-th point is is the weight coefficient used to balance the influence of geometric error and multimodal similarity, It is a similarity metric based on joint representation, which measures the similarity between the multimodal features of the corresponding points in the surface point cloud depth information and the target surface point cloud depth information. and For and The corresponding joint eigenvector.

[0030] Furthermore, the corruption integral equation is constructed according to the following steps:

[0031] S410, determining the influencing factors according to the types and volumes of vegetables and fruits, determining which multimodal features of the currently weighed vegetables and fruits will affect the spoilage process, and assigning a weight to each feature;

[0032] S420, describing the decay of the selected feature influence over time by an exponential decay factor, constructing a 0 The integral equation up to the current time t is as follows:

[0033] ,

[0034] ,

[0035] in, For the degree of corruption, is the weighted coefficient of the kth modal feature, is the eigenvector of the kth mode at time t'; is an exponential decay factor, which is used to indicate that as time t-t' increases, the influence of past characteristics on the current corruption level gradually weakens, e is a natural constant, and μ is a decay parameter; The final shelf life, D threshold is the critical value of corruption level.

[0036] On the other hand, the present invention provides a multimodal AI intelligent vegetable and fruit weighing system, which includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the multimodal AI intelligent vegetable and fruit weighing method as described above is implemented.

[0037] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0038] 1. The present invention couples multiple sensors into a sensor cluster to comprehensively acquire multimodal information of vegetables and fruits. At the same time, it interpolates and synchronizes the non-uniform sampling data collected by multiple sensors to generate a unified timestamp sequence, and dynamically aligns and fuses the asynchronous multimodal data set in the feature space to generate joint representation data, which effectively solves the problem that it is difficult to fuse multimodal data in the prior art.

[0039] 2. The present invention generates a high-precision and complete model based on the data collected by the visual sensor array and the ICP algorithm, identifies the types of the weighed vegetables and fruits, and accurately calculates the volume of the weighed items, providing an important basis for shelf life prediction.

[0040] 3. The present invention constructs a corruption integral equation based on the multimodal features in the joint characterization data, and uses the equation to predict the shelf life, thereby improving the robustness and reliability of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:

[0042] Figure 1 This is a flow chart of the method provided in Example 1 of the present invention.

[0043] Figure 2 This is a multi-sensor coupling workflow diagram provided in Example 1 of the present invention.

[0044] Figure 3 A timing diagram of the multi-sensor coupling cooperation sequence provided in Example 1 of the present invention.

[0045] Figure 4 This is a timing diagram of the algorithm for generating joint characterization data provided in Example 1 of the present invention.

[0046] Figure 5 This is a timing diagram of the optimization process of the ICP algorithm provided in Example 1 of the present invention.

[0047] Figure 6 This is a flowchart for predicting the shelf life of step 4 provided in Example 1 of the present invention.

[0048] Figure 7 This is a timing diagram for predicting the shelf life of step 4 provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0050] Example 1

[0051] This embodiment discloses a multimodal AI intelligent vegetable and fruit weighing method. The prior art usually uses a single sensor to weigh and detect vegetables and fruits, and is unable to fully obtain multimodal information of vegetables and fruits, such as surface point cloud depth information, weight information, gas information, and temperature and humidity information, resulting in low prediction accuracy. In addition, the prior art is difficult to effectively integrate multimodal data and cannot accurately predict the shelf life of vegetables and fruits. At the same time, when obtaining image data and temperature, humidity, and weight changes, the prior art cannot consider introducing more advanced sensors and image acquisition equipment to improve data accuracy and real-time performance.

[0052] In order to solve the above problems, the multimodal AI intelligent vegetable and fruit weighing method disclosed in this embodiment couples multiple sensors into a sensor cluster for collecting multimodal information of vegetables and fruits, and interpolates and synchronizes the non-uniform sampling data collected by multiple sensors to generate a unified timestamp sequence, and regards the sensor data with the same unified timestamp sequence as an asynchronous multimodal data set, and dynamically aligns and fuses the asynchronous multimodal data set in the feature space to generate joint representation data. Then, a high-precision and complete model is generated by the ICP algorithm based on the data collected by the visual sensor array, the type of vegetables and fruits weighed is identified, and the volume of the weighed items is calculated. Finally, the system constructs a corruption integral equation based on the multimodal features in the joint representation data, and the shelf life is predicted by the equation. Thereby improving the accuracy and robustness of the prediction of the shelf life of vegetables and fruits, and solving the problem that the prior art cannot effectively fuse multimodal data and is difficult to accurately predict the shelf life of vegetables and fruits.

[0053] It should be noted that the "non-uniform sampling data" in this embodiment refers to the phenomenon that there are unequal intervals or irregularities between the data collected at different time points. Specifically, the data collection of the sensor is usually carried out at a certain time interval, but if the sampling time of multiple sensors is not completely synchronized due to factors such as the response time of the sensor, the difference in sensor performance, the influence of the external environment, etc., the time intervals (that is, the sampling intervals) between the data collected by these sensors will be inconsistent, and such data is called "non-uniform sampling data". In this embodiment, multiple sensors are coupled into a sensor cluster for collecting different types of information. These sensors may sample data at different time points for various reasons, so the sampled data they generate may be non-uniform (that is, the data is not collected at fixed intervals in time). In order to enable these data to be compared and fused in the same time frame, it is usually necessary to interpolate and synchronize these non-uniform sampling data so that the data of all sensors can correspond to a unified timestamp sequence. This is to ensure that in the subsequent analysis and processing process, all data can be aligned on the same time basis, so as to perform effective dynamic alignment and fusion, and generate comprehensive joint characterization data. In short, "non-uniform sampling data" refers to the situation where the time intervals of sensor data collection are uneven or irregular. In this case, interpolation or other processing is usually required to synchronize the data in time for subsequent analysis and fusion.

[0054] Figure 1 The flowchart of the multimodal AI intelligent vegetable and fruit weighing method disclosed in this embodiment is shown. It can be seen from the figure that this embodiment includes the following steps:

[0055] Step 1: Couple multiple sensors into a sensor cluster to collect surface point cloud depth information, weight information, gas information, and temperature and humidity information of vegetables and fruits.

[0056] Specifically, the sensor cluster in this implementation includes the following sensors:

[0057] The visual sensor array is composed of ToF sensors or structured light sensors. ToF sensors are suitable for dynamic scenes (such as scanning fruits and vegetables on a conveyor belt), and structured light sensors are suitable for static scenes (such as manually placed fruits and vegetables). The visual sensor can also include RGB cameras and infrared spectrum sensors to assist in point cloud color information, correct light effects, and perform image segmentation to separate fruits and vegetables from their backgrounds.

[0058] The weight sensor obtains mass data, which is used to calculate density with volume and verify the accuracy of the point cloud algorithm.

[0059] Gas sensor array, including MOS gas sensor and VOC gas sensor, used for odor detection of fruits and vegetables, detection of ethylene, CO 2 , volatile organic compounds (VOCs) and other metabolic gases.

[0060] Environmental sensor array, including temperature and humidity sensors, is used to perform temperature and humidity compensation on gas sensor data.

[0061] Specifically, Figure 2 The multi-sensor coupling workflow diagram is shown. From the diagram, it can be seen that the multi-sensor coupling workflow is:

[0062] After the weight sensor detects a change in weight (fruits and vegetables are put in), it wakes up other sensors.

[0063] The visual sensor array triggers scanning based on weight. When the weight is detected to be stable, the gas sensor array is activated. When the gas sensor array detects a sudden change in the environmental VOCs baseline (such as rotting vegetables producing odor), a secondary scan verification is triggered. Figure 3 A timing diagram showing the cooperative sequence of multi-sensor coupling.

[0064] Step 2: Interpolate and synchronize the non-uniform sampling data collected by multiple sensors to generate a unified timestamp sequence. Treat the sensor data with the same unified timestamp sequence as an asynchronous multimodal data set, and dynamically align and fuse the asynchronous multimodal data set in the feature space to generate joint representation data.

[0065] Specifically, Figure 4 The algorithm timing diagram for generating joint representation data in this embodiment is shown. It can be seen from the diagram that the joint representation can be generated through the following steps:

[0066] 1) First, dynamic time alignment is performed. Assume that the sensor data is a multimodal time series, including K sensor modes (such as ToF, weighing, gas), where each modality data is collected independently and asynchronously, but all data carry a unified global timestamp sequence. The original data of mode k is represented as:

[0067] ,

[0068] ,

[0069] in, is the original data set of mode k, At timestamp t j The collected data points of mode k, R is the original data set, d k is the vector dimension, is the valid timestamp set of mode k, indicating that the mode has data collection at these time points, T is the global timestamp, t plus subscript numbers and letters (t 1, t 2, t N ) indicates a specific timestamp.

[0070] Then, we construct a cross-modal time window, fill in missing values ​​and generate a uniformly aligned feature tensor. The features of the aligned modality k are:

[0071]

[0072] in, To align the modal features after the global timestamp, is the length of the dynamic time window, is the weight of the weighted average, To fill in the nearest neighbor value, the most recent valid value can be read from the historical data.

[0073] Specifically, in this embodiment, the weight of the weighted average can be calculated by the following formula:

[0074] ,

[0075] Among them, exp is the exponential function, It is a decay factor used to control the influence of time proximity. The closer the time, the higher the weight.

[0076] 2) Map the features of each modality to a unified space to eliminate dimensional differences, define a learnable linear projection matrix for each modality k, and map the modal features after the global timestamp to the feature space using the following formula:

[0077] ,

[0078] in, is the projection of mode k in the feature space, is the linear projection matrix, , d h is the vector dimension of the feature space, b k is the bias term, which represents the feature projection bias of mode k.

[0079] 3) For each mode k, calculate its i The attention score of each modality is weighted and summed by the attention score. Specifically, the weighted summation can be calculated according to the following formula:

[0080] ,

[0081] in, For time ti The joint representation of Score for attention.

[0082] In this embodiment, the attention score can be calculated by the following formula:

[0083] ,

[0084] Among them, Softmax is a function, is the transpose operation of vector v, is the hyperbolic tangent activation function, used for nonlinear transformation, is the attention transformation matrix of modality k, For time t i Context encoding. V is the context transformation matrix, , It is the joint representation of the previous moment, and LSTM is a long short-term memory network, which is used to extract contextual features.

[0085] Step 3: Generate a high-precision and complete model through the ICP algorithm based on the data collected by the visual sensor array, and calculate the volume of the weighed object through the sensor data.

[0086] At the same time, the RGB data (color histogram) in the joint representation data and the weight are combined to estimate the volume. The volume estimation is calculated by the following formula:

[0087] ,

[0088] Among them, the mass is calculated based on the data of the weight sensor, To estimate the volume, For reference density, the system contains a reference density table of common fruits and vegetables, and the corresponding data can be obtained by looking up the table. By estimating the volume, the point cloud is scaled and corrected using the volume ratio to accelerate the ICP convergence and reduce the convergence time of the ICP algorithm.

[0089] In order to enable the ICP algorithm to achieve point cloud modeling faster, multimodal constraints are introduced into the objective function of the ICP algorithm so that the registered point cloud is not only geometrically aligned, but also maintains consistency in multimodal features (such as color, texture, etc.), thereby improving the robustness of the registration. In this embodiment, the multimodal constraint is expressed by the following formula:

[0090] ,

[0091] in, is the objective function of the ICP algorithm, is the i-th point in the source point cloud, For the target cloud The corresponding j-th point is is the weight coefficient used to balance the influence of geometric error and multimodal similarity, It is a similarity measure based on joint representation, which measures the similarity between the multimodal features of corresponding points in the source point cloud and the target point cloud. and For and The corresponding joint eigenvector.

[0092] It should be noted that Figure 5 The timing diagram of the optimization process of the ICP algorithm in this embodiment is shown. It can be seen from the figure that the multimodal constraint in this embodiment combines geometric error and multimodal similarity together, so that the objective function not only considers the geometric alignment of the point cloud, but also considers the consistency of multimodal features. This method makes the point cloud registration more robust and accurate, especially in the application scenarios of vegetables and fruits with complex features. In addition, by adjusting the weight coefficient, the influence between geometric error and multimodal similarity can be flexibly weighed to optimize the registration effect.

[0093] Step 4: The system constructs a corruption integral equation based on the multimodal features in the joint representation data, and uses this equation to predict the shelf life. Figure 6 The flowchart of predicting the shelf life in step 4 of this embodiment is shown. Figure 7 The predicted time sequence diagram of the shelf life in step 4 of this embodiment is shown. As can be seen from the figure, the corruption integral equation can be constructed according to the following steps:

[0094] 1) First, determine the influencing factors based on the type and volume of vegetables and fruits, determine which multimodal features of the currently weighed vegetables and fruits will affect the spoilage process, and assign weights to each feature.

[0095] 2) Use the exponential decay factor to describe the decay of the selected feature influence over time, and construct 0 The integral equation up to the current time t is used to accumulate the historical impact of all features, and the terminal shelf life can be predicted through this integral equation.

[0096] In this embodiment, the integral equation is as follows:

[0097] ,

[0098] in, For the degree of corruption, is the weighted coefficient of the kth modal feature, indicating the importance of this modal feature in the corruption process. is the eigenvector of the kth mode at time t'. Different modes can represent different environmental parameters or influencing factors, such as temperature, humidity, light, etc.

[0099] is an exponential decay factor, which is used to indicate that as time t-t' increases, the influence of past characteristics on the current corruption level gradually weakens. e is a natural constant, and μ is a decay parameter used to control the decay speed.

[0100] The end shelf life is expressed by the following formula:

[0101] ,

[0102] in, is the end shelf life, D threshold is the critical value of the degree of corruption. When D(t) reaches or exceeds this value, it is considered to be corrupt and has reached the end of its shelf life. This definition means finding the first time that D(t) reaches or exceeds the threshold D threshold Time point T end , which is the end of the product's shelf life.

[0103] Step 5: The weighed vegetables and fruits are displayed through a real-time visual interface, and the 3D model reconstructed through the point cloud data includes the volume details of the model, and a curve of the degree of corruption over time drawn based on the corruption integral equation, and the terminal shelf life is marked.

[0104] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multimodal AI intelligent vegetable and fruit weighing method, characterized in that: The vegetable and fruit weighing method comprises: S100, coupling multiple sensors into a sensor cluster, collecting surface point cloud depth information, weight information, gas information, and temperature and humidity information of vegetables and fruits, and using the collected information as non-uniform sampling data; S200, performing interpolation synchronization on the non-uniform sampling data to generate a unified timestamp sequence, The sensor data with the same unified timestamp sequence is regarded as an asynchronous multimodal data set. Dynamically align and fuse asynchronous multimodal data sets in feature space to generate joint representation data; S300, generating a 3D model of vegetables and fruits through an ICP algorithm according to the data collected by the visual sensor array, identifying the type of vegetables and fruits weighed through the generated model, and calculating the volume of the weighed items; S400, constructing a corruption integral equation based on the multimodal features in the joint representation data, and using the equation to predict the shelf life.

2. The multimodal AI intelligent vegetable and fruit weighing method according to claim 1 is characterized in that: The vegetable and fruit weighing method further comprises step S500: The 3D model of the weighed vegetables and fruits is displayed through a visual interface, and the volume contained in the 3D model, the corruption degree change curve drawn by the corruption integral equation, and the final shelf life are also marked.

3. The multimodal AI intelligent vegetable and fruit weighing method according to claim 1, characterized in that: The coupling of the plurality of sensors into a sensor cluster comprises: After the weight sensor detects the weight change, the wake-up step is started; The visual sensor array triggers scanning based on weight, and activates the gas sensor array when the weight is detected to be stable; When the gas sensor array detects a sudden change in the ambient VOCs baseline, a secondary scan verification is triggered.

4. The multimodal AI intelligent vegetable and fruit weighing method according to claim 1, characterized in that: The joint characterization data is generated by the following sub-steps: S210, performing dynamic time alignment, assuming that the sensor data is a multimodal time series, including K sensor modes, where each modality data is collected independently and asynchronously, and all data carry a unified global timestamp sequence, The original data of mode k is represented as: , , in, is the original data set of mode k, At timestamp t j The collected data points of mode k, R is the original data set, d k is the vector dimension, is the valid timestamp set of mode k, indicating that the mode has data collection at these time points, T is the global timestamp, t plus subscript numbers and letters (t 1, t 2, t N ) indicates a specific timestamp; S220, construct a cross-modal time window, fill in missing values ​​and generate a uniformly aligned feature tensor. The features of the aligned modality k are: , in, To align the modal features after the global timestamp, is the length of the dynamic time window, is the weight of the weighted average, is the value filled by the nearest neighbor, otherwise is the value filled by the nearest neighbor; S230, define a linear projection matrix for each modality k, and map the modal features after the global timestamp to the feature space by the following formula: , in, is the projection of mode k in the feature space, is the linear projection matrix, , d h is the vector dimension of the feature space, b k is the bias term; S240, calculate the velocity of each mode k at time t i The attention score of each modality is weighted and summed according to the following formula: , in, For time t i The joint representation of Score for attention.

5. The multimodal AI intelligent vegetable and fruit weighing method according to claim 4 is characterized in that: The weight of the weighted average in step S220 is calculated by the following formula: , Among them, exp is the exponential function, It is a decay factor used to control the influence of time proximity. The closer the time, the higher the weight.

6. The multimodal AI intelligent vegetable and fruit weighing method according to claim 4 is characterized in that: The attention score is calculated by the following formula: , Among them, Softmax is a function, is the transpose operation of vector v, is the hyperbolic tangent activation function, is the attention transformation matrix of modality k, For time t i Context encoding, V is the context transformation matrix.

7. The multimodal AI intelligent vegetable and fruit weighing method according to claim 1, characterized in that: The volume of the weighed object is calculated by the following formula: , Among them, the mass is calculated based on the data of the weight sensor, To estimate the volume, is the reference density.

8. The multimodal AI intelligent vegetable and fruit weighing method according to claim 1, characterized in that: The step S300 also includes multi-modal constraints of the ICP algorithm objective function, The multimodal constraint is expressed as follows: , in, is the objective function of the ICP algorithm, is the i-th point in the surface point cloud depth information, For the target cloud The corresponding j-th point is is the weight coefficient used to balance the influence of geometric error and multimodal similarity, It is a similarity metric based on joint representation, which measures the similarity between the multimodal features of the corresponding points in the surface point cloud depth information and the target surface point cloud depth information. and For and The corresponding joint eigenvector.

9. The multimodal AI intelligent vegetable and fruit weighing method according to claim 1, characterized in that: The corruption integral equation is constructed according to the following steps: S410, determining the influencing factors according to the types and volumes of vegetables and fruits, determining which multimodal features of the currently weighed vegetables and fruits will affect the spoilage process, and assigning a weight to each feature; S420, using an exponential decay factor to describe the decay of the selected feature influence over time, constructing an integral equation from the initial time t0 to the current time t as shown in the following formula: , , in, For the degree of corruption, is the weighted coefficient of the kth modal feature, is the eigenvector of the kth mode at time t'; is an exponential decay factor, which is used to indicate that the influence of past characteristics on the current corruption level gradually weakens as time t-t' increases, e is a natural constant, and μ is a decay parameter; The final shelf life, D threshold is the critical value of corruption level.

10. A multi-modal AI intelligent vegetable and fruit weighing system, characterized in that: The vegetable and fruit weighing method system comprises: processor; A memory storing a computer program, which, when executed by a processor, implements the multimodal AI intelligent vegetable and fruit weighing method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Winter jujube preservation control method and system based on IoT and AI

    CN119128451A

  • Rabbit meat freshness detection method based on electronic nose technology

    CN119643646A

  • Multi-modal fusion detection method for freshness of pork tenderloin based on deep learning

    CN119691509A

  • Shelf-life monitoring sensor-transponder system

    US20050248455A1