Picking robot fruit ripeness evaluation method fused with atlas reasoning
By integrating graph reasoning and adaptive correction, the problem of recognition instability in fruit ripeness assessment under environmental changes was solved, achieving high-precision and high-stability fruit ripeness assessment and improving the recognition accuracy and reliability of the harvesting robot.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGXIA ZHIAN AGRI & FORESTRY TECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-28
AI Technical Summary
Existing fruit ripeness assessment methods are sensitive to ambient light and shading conditions, resulting in unstable identification results. Furthermore, they lack effective fusion of multi-source information and analysis of model uncertainties, leading to insufficient assessment accuracy and low reliability.
A method combining fusion graph reasoning and adaptive correction is adopted. By synchronizing the time, registering the space, and filtering the noise of multimodal sensing data, a fusion graph of fruit and environmental features is constructed. Graph reasoning is used for information propagation and feature updating. Evidence deep learning algorithm is introduced for adaptive correction to calculate the confidence and uncertainty of fruit ripeness level.
It achieves high-precision and high-stability fruit ripeness assessment in complex agricultural environments, significantly improving the recognition accuracy and reliability of the harvesting robot and providing a stable basis for harvesting decisions.
Smart Images

Figure CN121935657A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of evaluation and analysis technology, specifically relating to a method for evaluating the ripeness of fruit in a harvesting robot that integrates graph reasoning. Background Technology
[0002] Currently, fruit ripeness assessment largely relies on machine vision recognition technology, which analyzes single visual features such as fruit color, texture, and shape to determine ripeness. However, these methods are extremely sensitive to ambient lighting, shooting angle, and occlusion conditions, easily leading to unstable recognition results. Some studies have attempted to combine environmental parameters such as temperature, humidity, and light for multimodal fusion, but most only use simple feature splicing or weighted averaging, failing to effectively establish semantic relationships between fruit features and environmental features, resulting in insufficient adaptability of the fusion results to complex scenarios.
[0003] Meanwhile, existing deep learning models mostly employ static feature weights or single-classification outputs, lacking analysis of model uncertainties and evidence conflicts, making it difficult to maintain output stability when there are biases or noise interference from multiple sources. Especially in the operation scenario of harvesting robots, lighting changes and environmental fluctuations are frequent. Traditional vision models cannot dynamically adjust feature weights, nor can they achieve adaptive correction of highly conflicting features, resulting in insufficient accuracy, low reliability, and a lack of interpretability in fruit ripeness assessment.
[0004] Therefore, how to provide a fruit ripeness assessment method for harvesting robots that integrates graph reasoning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a fruit ripeness assessment method for a harvesting robot that integrates graph reasoning. This invention integrates graph reasoning and adaptive correction to achieve intelligent fruit ripeness assessment with high accuracy and high stability.
[0006] A method for assessing fruit ripeness using a harvesting robot based on fusion graph reasoning, according to an embodiment of the present invention, includes the following steps: Image data and environmental data of the target fruit and its surrounding environment are collected, and time synchronization, spatial registration and noise filtering are performed to obtain multimodal sensing data; Preprocessing of image data in multimodal sensing data; The preprocessed image data is input into a deep feature extraction network to extract depth features. At the same time, environmental features are extracted from the multimodal perception data, and all features are integrated to form a multimodal feature vector. Constructing a fusion map of fruit features and environmental features based on multimodal feature vectors; The fused graph is input into the graph reasoning process to perform information propagation and feature updates between nodes. The node embedding vector is calculated through recursive propagation to generate fruit ripeness reasoning embedding features. An improved evidence deep learning algorithm is developed by embedding graph reasoning into feature input. An evidence dynamic adjustment coefficient is introduced to adaptively correct the Dirichlet distribution parameters corresponding to the input features. The confidence and uncertainty of fruit ripeness grade are calculated based on maximizing the evidence likelihood function. Evidence support and evidence conflict are calculated based on confidence and uncertainty. Confidence correction is performed based on evidence support and evidence conflict. When the evidence conflict exceeds a preset threshold, the weight parameters of the feature evidence are iteratively updated to obtain the fruit ripeness assessment result and output the fruit ripeness level.
[0007] Optionally, the generation of the multimodal sensing data includes: Collect image and environmental data of the target fruit and its surrounding environment; The image data and environmental data are uniformly time-stamped to form a multi-source data set corresponding to time. Spatial registration is performed between the synchronized image data and the environmental data; Noise filtering is performed on the image and environmental data that have been synchronized and spatially registered. Environmental data are smoothed using a weighted moving average method. By time synchronization, spatial registration, and noise filtering, multimodal sensing data that is temporally corresponding, spatially aligned, and has low noise interference can be obtained.
[0008] Optionally, the preprocessing includes: Noise removal is performed on image data in multimodal sensing data by using a smoothing algorithm based on pixel neighborhood relationships to reduce high-frequency noise components. Color correction is performed on the image data after noise removal, and the color distribution of the image is normalized and corrected based on the color channel ratio adjustment model. Image enhancement processing is performed on the color-corrected image data, and a local contrast enhancement method is used to strengthen the brightness difference between the fruit area and the background.
[0009] Optionally, the formation of the multimodal feature vector includes: The preprocessed fruit image data is input into a deep feature extraction network; The deep feature extraction network is used to extract fruit color features, fruit surface texture features, and fruit morphology features from fruit image data. Temperature, humidity, and light intensity features are extracted from environmental data. These environmental features are then normalized to unify the numerical range of environmental features of different dimensions between zero and one. By fusing fruit image features with environmental features, a multimodal feature vector is constructed. In the feature fusion process, a feature fusion loss function is defined, and the weight parameters and bias parameters are iteratively optimized based on the loss function to obtain a multimodal feature vector that comprehensively reflects the relationship between fruit appearance features and environmental features.
[0010] Optionally, the generation of the fusion map of fruit features and environmental features includes: The fruit color feature, fruit surface texture feature, fruit shape feature, ambient temperature feature, ambient humidity feature, and ambient light intensity feature in the multimodal feature vector are defined as nodes respectively. Each node stores the statistical attributes of the corresponding feature, forming a node feature set. Establish connections between nodes; Pearson correlation coefficient is used to quantitatively describe the semantic correlation between nodes; Based on the correlation weights between nodes, the connecting edges between nodes are assigned corresponding weights to construct a weighted feature association graph. By combining the weighted feature association graph with the node feature set, a fusion map of fruit features and environmental features is generated.
[0011] Optionally, the generation of the fruit ripeness embedding feature includes: The fusion map of fruit features and environmental features is input into the map reasoning process; The node features are updated layer by layer through a multi-layer recursive propagation mechanism. Each layer of propagation integrates the structural information and feature semantics of the adjacent nodes. After completing multi-layer propagation and feature update, a weighted average aggregation method is used to fuse the embedding vectors of all nodes to generate fruit ripeness inference embedding features.
[0012] Optionally, the generation of the confidence level and uncertainty includes: Using fruit ripeness inference embedding features as input features, an evidence modeling structure based on Dirichlet distribution is established, with different fruit ripeness levels as output nodes; In the process of evidence modeling, a dynamic adjustment coefficient for evidence is introduced. The distribution parameters of each maturity level are adaptively corrected according to the complexity of the input features, the noise level of the samples, and the confidence of the samples. By adjusting the ratio of high-confidence features to low-confidence features, the set of corrected distribution parameters is obtained. An evidence likelihood function is constructed based on the modified set of distribution parameters. The distribution parameters are iteratively optimized by taking the logarithm of the evidence likelihood function and minimizing its negative value, and the optimized distribution parameters are obtained. The confidence level and uncertainty of fruit ripeness grade are calculated based on the optimized distribution parameters.
[0013] Optionally, the generation of the fruit ripeness grade includes: Calculate the degree of support and conflict of evidence based on the degree of trust and uncertainty. The stability of the fruit ripeness assessment result is determined by comparing the degree of evidence conflict with the preset threshold. When the degree of evidence conflict is lower than the preset threshold, the current feature evidence weight parameter remains unchanged. When the degree of evidence conflict exceeds the preset threshold, the confidence correction process is initiated. During the confidence correction process, the correction coefficient is calculated based on the degree of conflict. The feature evidence weight parameters are adaptively and iteratively updated based on the correction coefficient. When the evidence conflict degree of a certain fruit feature exceeds the preset threshold, its evidence weight is reduced. When the evidence support degree of a certain fruit feature is higher than the average level, its evidence weight is increased. Through multiple iterations, the overall evidence distribution gradually becomes more consistent. When the evidence conflict level after iterative update is reduced to within a preset threshold range, the confidence correction process is terminated, the final feature evidence set is used as the optimized feature evidence set, the fruit ripeness assessment result is generated based on the optimized feature evidence set, and the fruit ripeness level is output.
[0014] The beneficial effects of this invention are: This invention establishes an intelligent fruit ripeness assessment method for harvesting robots by integrating a deep learning framework combining graph reasoning and evidence theory, achieving end-to-end optimization from multimodal perception to confidence correction. Compared with traditional schemes relying on single visual features or shallow fusion models, this invention utilizes temporal synchronization and spatial registration of image data and environmental data to ensure consistency of multi-source data across the spatiotemporal dimensions. By constructing a fusion graph model from fruit color, texture, and morphological features along with environmental temperature, humidity, and light characteristics, the system can capture the deep semantic relationships between fruit physiological characteristics and environmental factors, achieving structured modeling and accurate representation of the fruit ripening process in complex operational scenarios, thereby significantly improving the stability and generalization ability of ripeness assessment.
[0015] This invention further introduces an evidence modeling method based on Dirichlet distribution, simultaneously generating the confidence and uncertainty of fruit ripeness levels at the model output layer. This enables the evaluation results to possess not only classification capabilities but also confidence inference capabilities. A dynamic balance is achieved by combining an adaptive correction mechanism across three dimensions: feature complexity, sample noise level, and sample confidence. Complexity adjustment weights ensure full utilization of high-information features, noise suppression factors effectively weaken the influence of anomalous features, and confidence adjustment coefficients increase the weight of reliable samples in the modeling process. These three factors are normalized using the Softmax function to calculate a comprehensive adjustment coefficient, ensuring the model maintains adaptive optimization capabilities under different feature quality conditions. Through a confidence correction process based on evidence conflict, this invention can dynamically detect evidence conflicts between different ripeness levels, suppress high-conflict features with weights, strengthen high-consistency features, gradually reduce model uncertainty, and achieve stable and reliable ripeness assessment output.
[0016] In summary, this invention overcomes the technical shortcomings of existing fruit ripeness assessment methods, such as unstable identification under changing environmental conditions, insufficient feature fusion, and low result reliability. It establishes an intelligent assessment system with semantic reasoning capabilities, a confidence adjustment mechanism, and adaptive correction characteristics. This method can stably output high-confidence fruit ripeness grades in complex agricultural environments, providing accurate and reliable harvesting decisions for harvesting robots and significantly improving the intelligence and reliability of fruit detection and harvesting operations. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0018] Figure 1 This is a schematic diagram of the overall process of a fruit ripeness assessment method for a harvesting robot that integrates graph reasoning, as proposed in this invention. Figure 2 This is a schematic diagram of the adaptive correction and confidence correction process based on evidence modeling for a fruit ripeness assessment method for a harvesting robot that integrates graph reasoning, as proposed in this invention. Detailed Implementation
[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0020] refer to Figure 1 and Figure 2 A method for fruit ripeness assessment using a fruit-harvesting robot that integrates graph reasoning includes the following steps: Image data and environmental data of the target fruit and its surrounding environment are collected, and time synchronization, spatial registration and noise filtering are performed to obtain multimodal sensing data; Preprocessing of image data in multimodal sensing data; The preprocessed image data is input into a deep feature extraction network to extract depth features. At the same time, environmental features are extracted from the multimodal perception data, and all features are integrated to form a multimodal feature vector. Constructing a fusion map of fruit features and environmental features based on multimodal feature vectors; The fused graph is input into the graph reasoning process to perform information propagation and feature updates between nodes. The node embedding vector is calculated through recursive propagation to generate fruit ripeness reasoning embedding features. An improved evidence deep learning algorithm is developed by embedding graph reasoning into feature input. An evidence dynamic adjustment coefficient is introduced to adaptively correct the Dirichlet distribution parameters corresponding to the input features. The confidence and uncertainty of fruit ripeness grade are calculated based on maximizing the evidence likelihood function. Evidence support and evidence conflict are calculated based on confidence and uncertainty. Confidence correction is performed based on evidence support and evidence conflict. When the evidence conflict exceeds a preset threshold, the weight parameters of the feature evidence are iteratively updated to obtain the fruit ripeness assessment result and output the fruit ripeness level.
[0021] In this embodiment, the generation of the multimodal sensing data includes: Image data and environmental data of the target fruit and its surrounding environment are collected, including temperature data, humidity data, light intensity data and sensor detection data; The image data and environmental data are uniformly time-stamped so that data from different sources are synchronously acquired within the same sampling period, forming a multi-source data set corresponding to time. Spatial registration is performed on the synchronized image data and environmental data. The spatial registration is based on spatial mapping using an external parameter matrix, which consists of rotation and displacement parameters. This external parameter matrix is used to map pixel coordinates in the image coordinate system to a unified world coordinate system, thereby realizing the correspondence between image data feature points and environmental data sampling points in the same spatial reference system and ensuring the spatial consistency between image data and environmental data. The image data and environmental data that have completed time synchronization and spatial registration are subjected to noise filtering. The noise filtering includes Gaussian smoothing of the image data, with the square of the spatial distance between the pixel and the filter center as the weight factor. The larger the distance, the smaller the weight, thereby reducing the influence of background noise. The environmental data is smoothed using a weighted moving average method. By setting weights for environmental parameters in continuous time series and calculating a weighted average, random fluctuations are reduced and data stability is improved. By synchronizing time, registering space, and filtering noise, multimodal sensing data that is temporally corresponding, spatially aligned, and has low noise interference is obtained, providing a spatiotemporally consistent input basis for fruit ripeness assessment.
[0022] In this embodiment, the preprocessing includes: Noise removal processing is performed on the image data in the multimodal sensing data. A smoothing algorithm based on pixel neighborhood relationship is used to weaken high-frequency noise components, making the edge features of the fruit area more continuous. Color correction processing is performed on the image data after noise removal. The color distribution of the image is normalized and corrected based on the color channel ratio adjustment model to reduce the impact of ambient light differences on the color characteristics of the fruit. Image enhancement processing is performed on the color-corrected image data. The local contrast enhancement method is used to strengthen the brightness difference between the fruit area and the background, and to improve the salience of the fruit area edges. By removing noise, correcting color, and enhancing images, the distinguishability and stability of the target area of the fruit are improved.
[0023] In this embodiment, the formation of the multimodal feature vector includes: The preprocessed fruit image data is input into a deep feature extraction network, which consists of multiple convolutional layers, pooling layers, and fully connected layers. The convolutional layers are used to calculate the feature response of local regions of the fruit image to extract local texture information and edge structure features of the fruit image. The pooling layers are used to downsample the features output by the convolutional layers to retain the main feature distribution and reduce the data dimension. The fully connected layers are used to map the extracted features to a high-dimensional feature space to form a global feature representation. The deep feature extraction network extracts fruit color features, fruit surface texture features, and fruit morphology features from fruit image data. Fruit color features are used to characterize the color distribution of the fruit skin, fruit surface texture features are used to describe the detailed structure and roughness of the fruit surface, and fruit morphology features are used to reflect the outer contour and geometric shape of the fruit. Temperature, humidity, and light intensity features are extracted from environmental data. These environmental features are then normalized to unify the numerical range of environmental features of different dimensions between zero and one. Fruit image features and environmental features are fused to construct a multimodal feature vector. The fruit image feature vector consists of color, texture, and morphology features, while the environmental feature vector consists of temperature, humidity, and light intensity features. The fruit image feature vector and environmental feature vector are fused by weighted summation. During the fusion process, weight parameters and bias parameters are assigned to the fruit image features and environmental features respectively. The fusion weight parameters are continuously adjusted during training using gradient descent to minimize the error between the fused multimodal feature vector and the target ripeness label. In the feature fusion process, a feature fusion loss function is defined, which consists of a fusion consistency loss term and a ripeness prediction loss term. The value of the fusion consistency loss term is equal to the sum of the squares of the differences between the corresponding positions of visual features and environmental features in all samples, divided by the total number of samples. The value of the ripeness prediction loss term is equal to the average of the cross-entropy between the true ripeness distribution and the predicted ripeness distribution of all samples. The fusion consistency loss term is used to constrain the consistency between visual features and environmental features in the embedding space, and the ripeness prediction loss term is used to measure the difference between the fused feature prediction result and the true ripeness label. The total loss value of the feature fusion loss function is a weighted sum of two terms. The network weight parameters and bias parameters are iteratively optimized through the backpropagation algorithm until the loss function converges to the minimum value. This loss function is used to measure the difference between the fused features and the target ripeness result, and the weight parameters and bias parameters are iteratively optimized according to the loss function, so that the result of multimodal feature fusion more accurately represents the relationship between fruit visual features and environmental features, and obtains a multimodal feature vector that comprehensively reflects the relationship between fruit appearance features and environmental features.
[0024] In this embodiment, the generation of the fusion map of fruit features and environmental features includes: The fruit color feature, fruit surface texture feature, fruit shape feature, ambient temperature feature, ambient humidity feature, and ambient light intensity feature in the multimodal feature vector are defined as nodes. Each node is used to represent a single feature dimension. Each node stores the statistical attributes of the corresponding feature, including the feature mean, feature standard deviation, and normalized feature value sequence, forming a node feature set. Establish connections between nodes, where the edges of the connections are used to characterize the semantic correlation between different nodes. The existence of an edge indicates that there is a certain feature association between nodes, and the semantic correlation is used to reflect the strength of the statistical association between the features of different nodes. The establishment of the connection relationship specifically includes: The initial connection is determined based on feature correlation. The correlation coefficient between the feature values of any two nodes is calculated. When the absolute value of the correlation coefficient is greater than a set threshold, a connection edge is established between the corresponding nodes to reflect the correlation strength between features. Based on spatial similarity, neighboring connections are established. When feature sampling points are adjacent in spatial location, connections are established between the corresponding nodes to ensure the spatial continuity of the graph structure. Directed connections are determined based on physical dependencies. When environmental features have an influence relationship on fruit features, the direction of the edges is changed from environmental feature nodes to fruit feature nodes, forming a directed connection structure. The edge weights are calculated based on the feature values of the established connections. The edge weights are determined according to the Pearson correlation coefficient. The Pearson correlation coefficient is calculated as the ratio of the product of the covariance and standard deviation of the node feature values. The result is normalized to the [0,1] interval to characterize the semantic association strength between features. The Pearson correlation coefficient is used to quantitatively describe the semantic correlation between nodes. The Pearson correlation coefficient method is used to calculate the normalized sequence of each node's features. The correlation value between nodes is obtained by calculating the ratio of the product of the covariance of the feature samples and their respective standard deviations. The correlation value is used as the weight of the connection edge between nodes, and the weight value ranges from negative one to positive one. Based on the correlation weight results between each node, the connecting edges between nodes are assigned corresponding weights to construct a weighted feature association graph. The weighted feature association graph uses nodes as vertices and weighted edges as connections to reflect the quantitative relationship between fruit features and environmental features. By combining the weighted feature association graph with the node feature set, a fusion graph of fruit features and environmental features is generated. The fusion graph structurally includes a node set and a weighted adjacency matrix, which is used to represent the overall correlation between the visual features of the fruit and the environmental features in a structured way, providing an input basis for graph inference.
[0025] In this embodiment, the generation of the fruit ripeness embedding feature includes: The fusion graph of fruit features and environmental features is input into the graph inference process. The graph inference is based on a graph convolutional network structure to perform information propagation and feature updates between nodes. In each layer of graph convolution calculation, nodes are weighted and aggregated according to the feature vectors of their connected nodes and the weights of the connecting edges. The aggregated results are linearly mapped and nonlinearly activated through a trainable weight matrix to generate updated feature representations of the nodes. The node features are updated layer by layer through a multi-layer recursive propagation mechanism. Each layer of propagation fuses the structural information and feature semantics of the adjacent nodes, so that the node embedding vector can obtain the structural feature expression of the global fusion graph while maintaining the local feature information. After completing multi-layer propagation and feature update, a weighted average aggregation method is used to fuse all node embedding vectors to generate fruit ripeness inference embedding features. These fruit ripeness inference embedding features comprehensively represent the global semantic association between fruit features and environmental features.
[0026] In this embodiment, the generation of the confidence level and uncertainty includes: Using fruit ripeness inference embedding features as input features, an evidence modeling structure based on Dirichlet distribution is established, with different fruit ripeness levels as output nodes. Each ripeness level corresponds to a distribution parameter, which is used to characterize the evidence strength of that level. The distribution parameter is used to describe the evidence strength distribution corresponding to different fruit ripeness levels. Each fruit ripeness level corresponds to a distribution parameter, and the sum of all distribution parameter sets is used to describe the overall evidence quantity. When the overall evidence quantity is large, the output stability of the model is improved; when the overall evidence quantity is small, the uncertainty of the model increases. In the process of evidence modeling, a dynamic adjustment coefficient for evidence is introduced. The distribution parameters of each maturity level are adaptively corrected according to the complexity of the input features, the noise level of the samples, and the confidence of the samples. By adjusting the ratio of high-confidence features to low-confidence features, the set of corrected distribution parameters is obtained. The adaptive correction specifically includes: Initial adjustment weights are calculated based on the complexity of the input features. The complexity value of the feature is determined by a weighted sum of the number of feature dimensions and the feature information entropy. The number of feature dimensions represents the number of independent features in the sample, and the feature information entropy represents the randomness and uncertainty of the distribution of sample feature values. During calculation, the weight ratio of the number of feature dimensions and the feature information entropy is determined so that the sum of their weights is one. The number of feature dimensions is multiplied by its corresponding weight, and the feature information entropy is multiplied by its corresponding weight. The two results are then added together to obtain the feature complexity value. When a sample has more feature dimensions or higher information entropy, the calculated feature complexity value increases. The calculation result is normalized, and its value is limited to between zero and one, reflecting the overall information richness and complexity of the sample features. The initial adjustment weights are calculated based on the complexity value of the input features, and their value is determined through a proportional mapping relationship. During calculation, the feature complexity value is taken as a ratio to a constant, which reflects the relative level of feature complexity.
[0027] The noise level of the samples is calculated and a noise suppression factor is generated. The noise level is the ratio of the sample feature variance to the feature mean. The noise threshold is determined by the 95% confidence interval of the noise level of the training samples. When the sample noise level exceeds the noise threshold, the noise suppression factor is set to 1. When the sample noise level is higher than the noise threshold, the noise suppression factor is the ratio of the noise threshold to the sample noise level. This reduces the weight of high-noise samples in the update of distribution parameters, thereby suppressing the influence of noise features on model training. The confidence level of a sample is calculated and a confidence adjustment coefficient is generated. The confidence adjustment coefficient is calculated with the sample confidence level as input and its value is obtained through a linear mapping relationship. During the calculation, the sample's label consistency score and collection reliability score are weighted and averaged to obtain the overall confidence level of the sample. The label consistency score is used to measure the degree of matching between the sample's label and the average label of similar samples. The collection reliability score is used to reflect the stability of the collection device when acquiring the sample data. The two scores are multiplied by their respective weight ratios and then summed to obtain the sample confidence level, which is limited to the range between zero and one. The confidence adjustment coefficient is obtained by adding the sample confidence level and a constant to the base of one. The confidence adjustment coefficient is equal to one plus a constant multiplied by the sample confidence level. The higher the sample confidence level, the larger the confidence adjustment coefficient is, indicating that the sample has a higher adjustment range when the distribution parameter is corrected. In order to prevent the influence of high confidence samples from being over-amplified, a maximum upper limit value is set to one plus a constant. When the value exceeds the upper limit, the upper limit is taken as the final confidence adjustment coefficient. The complexity adjustment weights, noise suppression factors, and confidence adjustment coefficients are normalized using the Softmax function to ensure a weighted sum of 1. A comprehensive adjustment coefficient is then calculated, and based on this coefficient, the distribution parameters of all fruit ripeness grades are proportionally adjusted to obtain the adaptively adjusted set of distribution parameters. The complexity adjustment weights are calculated by dividing the feature complexity value by a fixed proportional constant, which describes the relative position of the sample in the complexity distribution. When the feature complexity value is large, this ratio increases, indicating that the sample has high information density and feature diversity. When the feature complexity value is small or close to zero, the ratio decreases, indicating limited feature information. To make the weight distribution smoother, this ratio is divided by... The result obtained by adding a constant to itself is used as the output value, so that the weight increases with the feature complexity. The comprehensive adjustment coefficient is calculated by the complexity adjustment weight, noise suppression factor and confidence adjustment coefficient. First, the complexity adjustment weight, noise suppression factor and confidence adjustment coefficient are normalized so that the weighted sum of the three is equal to one. The normalization adopts the Softmax function, which takes the exponent of the original value of each adjustment factor and compares it with the sum of the exponents of all factors to obtain the corresponding normalized result. This process ensures that the three are mutually constrained in value and the overall weight remains balanced. After the normalization is completed, the normalized results of the complexity adjustment weight, noise suppression factor and confidence adjustment coefficient are added according to the preset ratio to obtain the comprehensive adjustment coefficient. An evidence likelihood function is constructed based on the modified set of distribution parameters. For each maturity level, its modified distribution parameters are normalized to the sum of the distribution parameters of all maturity levels to obtain the predicted probability of that level. The predicted probabilities of all maturity levels are combined with the corresponding true maturity labels, and the likelihood function is formed by multiplying the predicted probabilities of all levels. This is used to measure the degree of matching between the model output and the true maturity label. The distribution parameters are iteratively optimized by taking the logarithm of the evidence likelihood function and minimizing its negative value to obtain the optimized distribution parameters. The confidence level and uncertainty of the fruit ripeness grade are calculated based on the optimized distribution parameters. The confidence level is calculated as the ratio of the distribution parameter of the ripeness grade to the sum of all distribution parameters, and the uncertainty is calculated as the ratio of the total number of fruit ripeness grades to the sum of all distribution parameters. The obtained confidence level and uncertainty are used as input data for the confidence correction step to further improve the accuracy and stability of the fruit ripeness assessment results.
[0028] In this embodiment, the generation of the fruit ripeness grade includes: Evidence support and evidence conflict are calculated based on confidence and uncertainty. Evidence support is calculated based on the average consistency between confidence levels of different fruit ripeness levels, and is used to reflect the consistency strength of evidence corresponding to each fruit ripeness level. Evidence conflict is calculated based on the degree of difference between confidence levels of different fruit ripeness levels, and is used to characterize the inconsistency strength of the model in the multi-evidence reasoning results. The stability of the fruit ripeness assessment result is determined by comparing the degree of evidence conflict with a preset threshold. When the degree of evidence conflict is lower than the preset threshold, the current feature evidence weight parameter remains unchanged. When the degree of evidence conflict exceeds the preset threshold, a confidence correction process is initiated. During the confidence correction process, a correction coefficient is calculated based on the degree of conflict. The correction coefficient is used to control the adjustment range of the evidence weight of each fruit feature. The correction coefficient is equal to the difference between the degree of evidence conflict of each feature and the threshold divided by the threshold. If the difference is negative, it is zero. When the degree of evidence conflict is lower than the threshold, the correction coefficient is small, and the adjustment range of the feature evidence weight is small or remains unchanged. When the degree of evidence conflict is higher than the threshold, the correction coefficient increases with the degree of conflict exceeding the threshold, thereby increasing the adjustment range of the feature evidence weight, realizing the reduction of the weight of high-conflict features and the maintenance or improvement of the weight of low-conflict features. The feature evidence weight parameters are adaptively and iteratively updated based on the correction coefficient. When the evidence conflict of a certain fruit feature exceeds the preset threshold, its evidence weight is reduced. When the evidence support of a certain fruit feature is higher than the average level, its evidence weight is increased. Through multiple iterations, the overall evidence distribution gradually becomes more consistent, thereby reducing the adverse impact of high-conflict features on the evaluation results and reducing the evaluation bias caused by uncertainty. When the evidence conflict level after iterative updates decreases to within a preset threshold range, the confidence correction process is terminated. The final set of feature evidence is used as the optimized set of feature evidence. Based on the optimized set of feature evidence, a fruit ripeness assessment result is generated, and the fruit ripeness level is output. The fruit ripeness assessment result has high confidence and stability, providing an assessment basis for the final determination of the fruit ripeness level.
[0029] Example 1: To verify the feasibility of this invention in practice, it was applied to the fruit ripeness recognition system of an automated harvesting robot in a southern orchard for an agricultural equipment company. This orchard is located in a hilly area with a warm and humid climate, densely distributed fruit trees, and frequent changes in ambient light. Fruit colors vary significantly but are easily affected by shadows, light reflection, and ambient temperature. Traditional fruit recognition methods based on single visual features suffer from unstable ripeness recognition, significant environmental interference, and large fluctuations in assessment accuracy in this scenario, failing to provide a reliable basis for harvesting robot decisions. To address this issue, the fruit ripeness assessment method based on fusion graph reasoning provided by this invention was integrated into the robot system to verify its stability and reliability in complex natural environments.
[0030] In practical applications, when the harvesting robot operates in the orchard, it captures images of the fruit using a high-resolution camera mounted at the end of the robotic arm. Simultaneously, it utilizes integrated multimodal sensors to acquire environmental information such as ambient temperature, humidity, and light intensity in real time. The system first performs time synchronization on the collected multi-source data to ensure that the image acquisition time is consistent with the environmental sensor sampling time. Subsequently, it performs spatial registration to make the visual image coordinates correspond to the environmental sampling coordinates under a unified spatial reference frame, thereby obtaining multimodal perception data that is consistent in time and space. To reduce visual noise caused by light fluctuations, dust, and wind blowing fruit branches, the system performs image noise filtering and data smoothing after registration, making the input data more stable and providing a reliable foundation for subsequent feature analysis.
[0031] After data preprocessing, the system inputs the fruit image into a deep feature extraction network. This network has learned multi-level information such as fruit color distribution, surface texture changes, and geometric shape during the training phase. Local texture features are extracted through convolutional layers, feature distribution is aggregated through pooling layers, and fully connected layers are mapped to a high-dimensional semantic space, thus forming a deep feature representation of the fruit. At the same time, the system processes environmental data, extracts temperature, humidity, and light intensity features, and normalizes them to unify their numerical distribution within the same dimension range so that they can be fused with visual features. Subsequently, the system combines the fruit visual features and environmental features in a weighted fusion manner to construct a multimodal feature vector. This feature vector comprehensively reflects the correspondence between the fruit's appearance and its growing environment.
[0032] To further capture the semantic dependencies between different features, the system constructs a fusion graph of fruit and environmental features based on the fused features. Fruit color, surface texture, morphology, temperature, humidity, and light intensity are defined as nodes, and semantically relevant edges are established between nodes. The connection relationship is determined based on feature relevance and spatial proximity, and the edge weight reflects the correlation strength between features. By constructing a weighted directed graph, the system can clearly represent the direction and intensity of the influence of environmental features on the fruit ripening process. Subsequently, the fusion graph is input into the graph inference process, and a graph convolution structure is used for information propagation and feature updating between nodes. This allows each node to retain its own feature information while fusing semantic information from other relevant nodes. After multiple recursive propagations, the system generates a fruit ripeness inference embedding feature, which can comprehensively express the global semantic relationship between fruit and environmental factors.
[0033] After generating the embedded features for inference, the system enters the evidence modeling stage. The model maps each maturity level to a distribution parameter and models the evidence strength at different maturity levels using a Dirichlet distribution. To improve the model's adaptability to complex sample conditions, an adaptive correction mechanism is introduced. This mechanism dynamically adjusts the evidence distribution parameters based on three factors: the complexity of the input features, the sample noise level, and the sample credibility. Feature complexity is obtained by weighted summation of the number of feature dimensions and information entropy, reflecting the richness of sample feature information. Sample noise level is measured by the ratio of variance to mean; samples with high noise receive a smaller noise suppression factor, thus reducing their impact on evidence modeling. Sample credibility is obtained by weighted summation of label consistency score and data acquisition device reliability score, used to determine the range of credibility of the sample data.
[0034] The system normalizes the three factors and uses the Softmax function to calculate the comprehensive adjustment coefficient, ensuring that the sum of the three factors remains constant. In this way, samples with high complexity, low noise, and high confidence obtain a larger comprehensive adjustment coefficient during the modeling process, thereby enhancing their influence on the correction of distribution parameters. The system proportionally corrects the distribution parameters of each maturity level based on the comprehensive adjustment coefficient, resulting in an adaptively adjusted parameter set. This process effectively balances the contributions of high-confidence samples and low-confidence samples, reduces the uncertainty caused by outliers, and provides a more stable probability distribution for subsequent maturity inference.
[0035] When calculating confidence and uncertainty, the system normalizes the distribution parameter set according to the correction, calculates the confidence corresponding to each maturity level, and calculates the overall uncertainty based on the sum of all parameters. Confidence reflects the credibility of the model's judgment on a certain maturity level, while uncertainty characterizes the stability of the current assessment. When the system detects that the confidence is low or the uncertainty is high, it will further calculate the evidence support and evidence conflict to analyze the consistency and conflict between evidence at different maturity levels. If the evidence conflict exceeds a set threshold, the system will initiate a confidence correction process, iteratively adjusting the weights of each feature evidence by calculating correction coefficients, so that the weights of high-conflict features are reduced, while the weights of low-conflict features are maintained or increased, thereby achieving self-balancing of evidence within the model.
[0036] During continuous operation in the orchard, the system automatically completes multiple rounds of confidence correction. In each round, the weight distribution is dynamically adjusted based on changes in the conflict level. Finally, when the conflict level drops to a threshold range, the system outputs a stable fruit ripeness grade result and transmits this result to the harvesting decision module. The harvesting robot then determines whether the fruit meets the harvesting standards, achieving intelligent recognition and autonomous decision-making regarding fruit ripeness.
[0037] Through operation verification under multiple time periods and climatic conditions, the system can maintain high evaluation stability under different light intensities, humidity changes, and shading conditions. Moreover, the ripeness judgment results are highly consistent with human assessments, proving that the present invention can effectively overcome the defects of unstable recognition accuracy, poor environmental adaptability, and low reliability of results in traditional methods. This implementation scheme achieves high-precision assessment and reliable identification of fruit ripeness in the process of intelligent harvesting in orchards, providing more robust technical support for agricultural automation operations, and has significant application value and promotion significance.
[0038] Table 1: Performance Comparison of the Method of the Present Invention and Traditional Fruit Ripeness Assessment Methods
[0039] As can be seen from the data in Table 1, the proposed method of fusing graph reasoning and evidence correction significantly outperforms traditional methods in the fruit ripeness assessment task. Firstly, in terms of average recognition accuracy, the proposed method achieves 96.8%, an improvement of approximately 9.4 percentage points compared to the 87.4% of traditional visual feature recognition methods, and an improvement of approximately 5.3 percentage points compared to the 91.5% of multimodal fusion CNN methods. This indicates that the proposed method has higher discriminative ability in fruit visual feature extraction and environmental semantic fusion. This is because the graph reasoning mechanism employed in this invention can capture high-order semantic relationships between fruit color, texture, and morphological features and environmental temperature, humidity, and light intensity, enabling the model to possess stronger global semantic understanding capabilities when processing multi-source heterogeneous data.
[0040] In terms of stability assessment (confidence fluctuation amplitude), the confidence fluctuation amplitude of the method in this invention is ±0.05, which is significantly lower than that of the traditional method (±0.18) and the multimodal fusion method (±0.12), indicating that the output results are more stable and the prediction fluctuations are smaller. This performance improvement is mainly due to the evidence modeling mechanism based on Dirichlet distribution introduced in this invention, which enables the model to output both confidence and uncertainty simultaneously, thereby maintaining the consistency of the output even when there is interference in the input data. Combined with the adaptive correction mechanism, through the joint normalization of complexity adjustment weights, noise suppression factors, and confidence adjustment coefficients, the model can automatically reduce the impact of high-noise samples on the inference results, enhancing the robustness and consistency of the system.
[0041] In terms of environmental adaptability scoring, the method of this invention achieved a score of 9.4, a significant improvement compared to the traditional method's 6.2 and the multimodal CNN method's 7.8. This indicates that the recognition ability of the model of this invention remains stable even under complex natural scenarios such as changes in light intensity, fluctuations in environmental humidity, and high reflectivity of the fruit surface. The root of this superior performance lies in the fact that the fusion graph inference mechanism forms a directed weighted connection structure between environmental feature nodes and fruit feature nodes. Cross-modal compensation of feature information is achieved through graph convolution propagation, effectively solving the limitation of traditional visual algorithms being susceptible to environmental disturbances.
[0042] In terms of computational efficiency, the method of this invention achieves a processing speed of 12.6 frames per second, slightly lower than the traditional method's 14.3 frames per second, but still meets the response requirements of the harvesting robot's real-time operation. This slight decrease in efficiency is due to the introduction of graph reasoning and evidence correction processes, which increases the information propagation and adaptive computation steps at the feature layer. However, this overhead is compensated by a significant improvement in stability and accuracy, resulting in a marked optimization of the system's overall performance.
[0043] Comprehensive analysis reveals that this invention, while ensuring real-time performance, achieves a comprehensive improvement in recognition accuracy, result stability, and environmental adaptability. The fundamental reason for this performance improvement lies in the fact that this invention integrates visual and environmental multimodal data, utilizes a graph-based reasoning structure for feature relationship modeling, making information dissemination more logically consistent; and introduces an evidence adaptive correction mechanism in the reasoning output stage to dynamically balance complexity, noise, and credibility factors. This mechanism effectively reduces the impact of external interference in fruit ripeness assessment, improves the rationality of feature weight allocation, thereby achieving high-confidence, low-fluctuation assessment results, and significantly enhances the stability and reliability of fruit ripeness recognition under complex natural conditions. In conclusion, this invention outperforms traditional methods in terms of ripeness assessment accuracy, judgment reliability, harvesting efficiency, and harvesting accuracy, fully demonstrating the synergistic advantages of multimodal perception, knowledge graph reasoning, and the KG-EDL model, providing an efficient and reliable technical solution for intelligent agricultural harvesting.
[0044] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for assessing fruit ripeness in a harvesting robot that integrates graph reasoning, characterized in that, Includes the following steps: Image data and environmental data of the target fruit and its surrounding environment are collected, and time synchronization, spatial registration and noise filtering are performed to obtain multimodal sensing data; Preprocessing of image data in multimodal sensing data; The preprocessed image data is input into a deep feature extraction network to extract depth features. At the same time, environmental features are extracted from the multimodal perception data, and all features are integrated to form a multimodal feature vector. Constructing a fusion map of fruit features and environmental features based on multimodal feature vectors; The fused graph is input into the graph reasoning process to perform information propagation and feature updates between nodes. The node embedding vector is calculated through recursive propagation to generate fruit ripeness reasoning embedding features. An improved evidence deep learning algorithm is developed by embedding graph reasoning into feature input. An evidence dynamic adjustment coefficient is introduced to adaptively correct the Dirichlet distribution parameters corresponding to the input features. The confidence and uncertainty of fruit ripeness grade are calculated based on maximizing the evidence likelihood function. Evidence support and evidence conflict are calculated based on confidence and uncertainty. Confidence correction is performed based on evidence support and evidence conflict. When the evidence conflict exceeds a preset threshold, the weight parameters of the feature evidence are iteratively updated to obtain the fruit ripeness assessment result and output the fruit ripeness level.
2. The method for assessing fruit ripeness in a harvesting robot based on fusion graph reasoning as described in claim 1, characterized in that, The generation of the multimodal sensing data includes: Collect image and environmental data of the target fruit and its surrounding environment; The image data and environmental data are uniformly time-stamped to form a multi-source data set corresponding to time. Spatial registration is performed between the synchronized image data and the environmental data; Noise filtering is performed on the image and environmental data that have been synchronized and spatially registered. Environmental data are smoothed using a weighted moving average method. By time synchronization, spatial registration, and noise filtering, multimodal sensing data that is temporally corresponding, spatially aligned, and has low noise interference can be obtained.
3. The method for assessing fruit ripeness in a harvesting robot based on fusion graph reasoning as described in claim 1, characterized in that, The preprocessing includes: Noise removal is performed on image data in multimodal sensing data by using a smoothing algorithm based on pixel neighborhood relationships to reduce high-frequency noise components. Color correction is performed on the image data after noise removal, and the color distribution of the image is normalized and corrected based on the color channel ratio adjustment model. Image enhancement processing is performed on the color-corrected image data, and a local contrast enhancement method is used to strengthen the brightness difference between the fruit area and the background.
4. The method for assessing fruit ripeness in a harvesting robot based on fusion graph reasoning as described in claim 1, characterized in that, The formation of the multimodal feature vector includes: The preprocessed fruit image data is input into a deep feature extraction network; The deep feature extraction network is used to extract fruit color features, fruit surface texture features, and fruit morphology features from fruit image data. Temperature, humidity, and light intensity features are extracted from environmental data. These environmental features are then normalized to unify the numerical range of environmental features of different dimensions between zero and one. By fusing fruit image features with environmental features, a multimodal feature vector is constructed. In the feature fusion process, a feature fusion loss function is defined, and the weight parameters and bias parameters are iteratively optimized based on the loss function to obtain a multimodal feature vector that comprehensively reflects the relationship between fruit appearance features and environmental features.
5. The method for assessing fruit ripeness in a harvesting robot based on fusion graph reasoning as described in claim 1, characterized in that, The generation of the fusion map of fruit characteristics and environmental characteristics includes: The fruit color feature, fruit surface texture feature, fruit shape feature, ambient temperature feature, ambient humidity feature, and ambient light intensity feature in the multimodal feature vector are defined as nodes respectively. Each node stores the statistical attributes of the corresponding feature, forming a node feature set. Establish connections between nodes; Pearson correlation coefficient is used to quantitatively describe the semantic correlation between nodes; Based on the correlation weights between nodes, the connecting edges between nodes are assigned corresponding weights to construct a weighted feature association graph. By combining the weighted feature association graph with the node feature set, a fusion map of fruit features and environmental features is generated.
6. The method for assessing fruit ripeness in a harvesting robot based on fusion graph reasoning as described in claim 1, characterized in that, The generation of the fruit ripeness embedding feature includes: The fusion map of fruit features and environmental features is input into the map reasoning process; The node features are updated layer by layer through a multi-layer recursive propagation mechanism. Each layer of propagation integrates the structural information and feature semantics of the adjacent nodes. After completing multi-layer propagation and feature update, a weighted average aggregation method is used to fuse the embedding vectors of all nodes to generate fruit ripeness inference embedding features.
7. The method for assessing fruit ripeness in a harvesting robot based on fusion graph reasoning as described in claim 1, characterized in that, The generation of the confidence level and uncertainty includes: Using fruit ripeness inference embedding features as input features, an evidence modeling structure based on Dirichlet distribution is established, with different fruit ripeness levels as output nodes; In the process of evidence modeling, a dynamic adjustment coefficient for evidence is introduced. The distribution parameters of each maturity level are adaptively corrected according to the complexity of the input features, the noise level of the samples, and the confidence of the samples. By adjusting the ratio of high-confidence features to low-confidence features, the set of corrected distribution parameters is obtained. An evidence likelihood function is constructed based on the modified set of distribution parameters. The distribution parameters are iteratively optimized by taking the logarithm of the evidence likelihood function and minimizing its negative value, and the optimized distribution parameters are obtained. The confidence level and uncertainty of fruit ripeness grade are calculated based on the optimized distribution parameters.
8. The method for fruit ripeness assessment of a harvesting robot based on fusion graph reasoning according to claim 1, characterized in that, The generation of the fruit ripeness grade includes: Calculate the degree of support and conflict of evidence based on the degree of trust and uncertainty. The stability of the fruit ripeness assessment result is determined by comparing the degree of evidence conflict with the preset threshold. When the degree of evidence conflict is lower than the preset threshold, the current feature evidence weight parameter remains unchanged. When the degree of evidence conflict exceeds the preset threshold, the confidence correction process is initiated. During the confidence correction process, the correction coefficient is calculated based on the degree of conflict. The feature evidence weight parameters are adaptively and iteratively updated based on the correction coefficient. When the evidence conflict degree of a certain fruit feature exceeds the preset threshold, its evidence weight is reduced. When the evidence support degree of a certain fruit feature is higher than the average level, its evidence weight is increased. Through multiple iterations, the overall evidence distribution gradually becomes more consistent. When the evidence conflict level after iterative update is reduced to within a preset threshold range, the confidence correction process is terminated, the final feature evidence set is used as the optimized feature evidence set, the fruit ripeness assessment result is generated based on the optimized feature evidence set, and the fruit ripeness level is output.