Greenhouse plant product harvesting monitoring method and system based on knowledge graph

By constructing a value heatmap and a dynamic fusion mechanism, and combining knowledge graphs to identify key areas and harvesting areas, the accuracy and efficiency issues of multimodal image harvesting systems in greenhouse environments were solved, achieving high-quality image fusion and precise harvesting.

CN121121502BActive Publication Date: 2026-02-10南京市农业装备推广中心 +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511667132.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-10
Estimated Expiration
2045-11-14

AI Technical Summary

Technical Problem

Existing automated harvesting systems cannot fully utilize multi-source information and cannot distinguish the differences in information value between different modal images. This results in low accuracy in the joint analysis of multimodal images, difficulty in adapting to complex greenhouse environments, and low quality of the fused images, which affects plant status identification and harvesting efficiency.

Method used

Based on knowledge graphs, this method analyzes the value of multimodal images through a quality assessment model, constructs a value heatmap, identifies key areas, and configures a dynamic fusion mechanism for image fusion. It also identifies harvesting areas by combining plant harvesting maps.

Benefits of technology

It improves the fusion quality of multimodal images and the accuracy and efficiency of the harvesting process, enhances adaptability to changes in the greenhouse environment, and ensures accurate identification of harvesting areas and efficient harvesting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121121502B_ABST
    Figure CN121121502B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of real-time monitoring, and discloses a greenhouse plant product harvesting monitoring method and system based on a knowledge graph, which comprises the following steps: calculating corresponding image values according to pre-acquired multi-modal images in a greenhouse, and constructing a value heat map; identifying key points of each modal image respectively, and screening out corresponding key regions; configuring a dynamic fusion mechanism, fusing the key regions of each modal image respectively, and obtaining a fused image; analyzing the fused image through a pre-constructed plant harvesting graph, identifying a harvesting region, and harvesting plant products in the harvesting region; through joint analysis of multi-modal images, the adaptability to changes in the greenhouse environment during the harvesting process of the plant products can be improved, low-quality information can be removed while retaining the advantageous features of each modal image, the quality of the fused image can be improved, the harvesting region can be quickly identified in combination with the knowledge graph, and the accuracy and efficiency of the harvesting process of the plant products in the greenhouse can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of real-time monitoring technology, and more specifically to a method and system for monitoring the harvesting of greenhouse plant products based on knowledge graphs. Background Technology

[0002] Currently, modern agriculture is developing towards precision and intelligence, and automated harvesting of greenhouse plant products has become an important research direction in agricultural technology. However, existing automated harvesting systems rely on single-type image data for analysis, failing to fully utilize multi-source information; in the image fusion process, they often use a simple overlay method with fixed weights, lacking evaluation of image quality and targeted processing of key information; and traditional image processing methods struggle to understand the semantics of complex agricultural scenes, resulting in low accuracy and environmental adaptability in harvesting decisions.

[0003] Existing technologies suffer from the following problems: low utilization efficiency of multimodal images, inability to distinguish the differences in information value between different modal images, resulting in low accuracy in joint analysis of multimodal images; low accuracy in identifying key areas when using fixed thresholds or single edge detection, making it difficult to adapt to complex greenhouse environments; direct static fusion of images, inability to dynamically adjust the fusion process based on image content features, resulting in low quality of the fused images, leading to inaccurate plant status identification results and affecting plant harvesting efficiency; to solve at least one of the above problems, this application proposes a greenhouse plant product harvesting monitoring method and system based on knowledge graphs. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this application is to provide a knowledge graph-based method and system for monitoring the harvesting of greenhouse plant products, which can effectively solve the problems in the background technology. The specific technical solution of this application is as follows:

[0005] Knowledge graph-based methods for monitoring greenhouse plant product harvesting include:

[0006] Based on the multimodal images pre-acquired in the greenhouse, the quality of each modal image is analyzed through a pre-set quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed.

[0007] Based on the aforementioned value heatmap, key points of each modal image are identified, and corresponding key regions are selected.

[0008] A dynamic fusion mechanism is configured to fuse key regions of each modality image separately, and the fusion process is dynamically adjusted based on the key points to obtain a fused image;

[0009] By analyzing the fused images using a pre-constructed plant harvesting atlas, harvesting areas are identified, and plant products from these areas are harvested.

[0010] Specifically, based on the pre-acquired multimodal images in the greenhouse, the quality of each modal image is analyzed using a pre-set quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed, including:

[0011] Based on the multimodal images pre-acquired in the greenhouse, the features of each modal image are extracted using a preset feature extraction model, and the corresponding image feature vector is constructed.

[0012] Based on the image feature vectors, the quality of each modal image is analyzed using a preset quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed.

[0013] Specifically, based on the image feature vectors, the quality of each modal image is analyzed using a preset quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed, including:

[0014] Based on the image feature vector of each modal image, the game process between each modal image is simulated through a preset quality assessment model to construct a game matrix;

[0015] Based on the game matrix, the image value of each modal image is calculated to obtain the corresponding global value score;

[0016] For each modal image, the image is segmented according to its structure using a pre-defined image segmentation model. The analytical value of different image structures is analyzed, and a corresponding structural value map is generated.

[0017] For each modal image, the analytical value of the harvesting area in the image is calculated using a preset harvesting analysis model, and a corresponding harvesting value map is generated.

[0018] The first value map is obtained by multiplying each pixel value in the structural value map by the corresponding global value score.

[0019] Based on the harvest value map, a dynamic adjustment factor is calculated for each pixel location. The first value map is then dynamically adjusted according to the dynamic adjustment factor to generate a value heatmap for each modal image.

[0020] Specifically, based on the aforementioned value heatmap, key points are identified for each modal image, and corresponding key regions are selected, including:

[0021] Based on the value heatmap of each modal image, the key points of each modal image are identified, and the first key region set is constructed by combining the corresponding key points.

[0022] For each region in the first set of key regions, calculate the sum of pixel values ​​within the region to obtain the value integral;

[0023] The regions are sorted from highest to lowest according to their value scores to obtain the region order;

[0024] Based on the region order, the spatial overlap between regions is calculated, and regions with a spatial overlap less than a preset overlap threshold are selected to obtain a set of key regions.

[0025] Specifically, based on the value heatmap of each modal image, key points of each modal image are identified, and a first set of key regions is constructed by combining the corresponding key points, including:

[0026] Based on the value heatmap of each modal image, the corresponding key points are identified by a preset key point recognition model to obtain the first set of key points;

[0027] From the first set of key points, select key points whose value is higher than a preset value threshold and whose dynamic adjustment factor at the corresponding pixel position is less than a preset adjustment threshold, and use them as the second set of key points.

[0028] For each key point in the second key point set, adjacent pixels are merged using a preset region generation model until the difference between the value of the adjacent pixel and the value of the key point is greater than a preset value difference threshold, thus obtaining the first key region set.

[0029] Specifically, a dynamic fusion mechanism is configured to fuse key regions of each modality image separately, and the fusion process is dynamically adjusted based on the key points to obtain a fused image, including:

[0030] Based on the key regions of each modal image, overlapping key regions, non-overlapping key regions, and non-key regions between different modal images are selected, and the weight of each modal image in the corresponding region is calculated to generate a weight map.

[0031] According to the weight map, the pixels of each modal image are weighted and fused to obtain the first fused image;

[0032] Based on the key points of each modal image, the pixels of the corresponding regions in the first fused image are dynamically adjusted to obtain the fused image.

[0033] Specifically, based on the key regions of each modal image, overlapping key regions, non-overlapping key regions, and non-key regions between different modal images are selected, and the weight of each modal image in the corresponding region is calculated to generate a weight map, including:

[0034] Based on the key regions of each modal image, overlapping key regions, non-overlapping key regions, and non-key regions between different modal images are selected;

[0035] For overlapping key regions, analyze the value ratio of corresponding key points and calculate the first weight of the corresponding modal image.

[0036] For non-overlapping key regions, the corresponding modal image weight is set to 1, and the weights of the other modal images are set to 0, thus obtaining the second weight.

[0037] For non-critical regions, analyze the image feature vectors of the modal images and calculate the third weight of the corresponding modal images;

[0038] By combining the first weight, the second weight, and the third weight, a weight graph is generated.

[0039] Specifically, the step of dynamically adjusting the pixels in the corresponding region of the first fused image based on the key points of each modal image to obtain the fused image includes:

[0040] Analyze the correspondence between key points in each modal image, match the key points, and obtain the key point matching results;

[0041] Based on the key point matching results, the displacement deviation of the corresponding pixel position is calculated to obtain the displacement deviation result;

[0042] Based on the displacement deviation results, the pixels in the corresponding regions of the first fused image are dynamically adjusted to obtain the fused image.

[0043] Specifically, the step of analyzing the fused image using a pre-constructed plant harvesting atlas to identify harvesting areas and harvesting plant products from those areas includes:

[0044] Harvesting areas are identified using pre-constructed plant harvesting maps based on the fused imagery.

[0045] Based on the harvesting area, the harvesting characteristics are analyzed using a preset harvesting analysis model, and corresponding harvesting results are generated to harvest the plant products in the harvesting area.

[0046] A knowledge graph-based greenhouse plant product harvesting monitoring system is used to implement the aforementioned knowledge graph-based greenhouse plant product harvesting monitoring method, including:

[0047] The value heatmap construction module analyzes the quality of each modal image based on the pre-acquired multimodal images in the greenhouse through a preset quality assessment model, calculates the corresponding image value, and constructs a value heatmap.

[0048] The key region identification module identifies key points for each modal image based on the value heatmap and filters out the corresponding key regions.

[0049] The image fusion module is configured with a dynamic fusion mechanism, which merges the key regions of each modality image separately, and dynamically adjusts the fusion process through the key points to obtain the fused image;

[0050] The plant product harvesting module analyzes the fused image using a pre-constructed plant harvesting atlas to identify harvesting areas and harvest plant products from those areas.

[0051] The beneficial effects of this application are as follows: Based on the game-theoretic process analysis of multimodal images, image quality is analyzed, image value is calculated, and a value heatmap is constructed to quantitatively analyze image quality information; high-value key points are identified based on the value heatmap, and key regions are identified by merging adjacent effective pixels and removing overlapping areas; different weight calculation methods are set for different regions, multimodal images are weighted and fused, and the fused image is calibrated to improve the quality of the fused image; the area requiring harvesting is identified by combining knowledge graphs, and real-time harvesting of plant products is carried out in the harvesting area; the joint analysis of multimodal images can improve the adaptability of plant products to changes in the greenhouse environment during the harvesting process; dynamic weight fusion is set to remove low-quality information while retaining the advantageous features of each modality of image, thereby improving the quality of the fused image; combined with knowledge graphs, the harvesting area can be quickly identified, improving the accuracy and efficiency of the greenhouse plant product harvesting process. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the knowledge graph-based greenhouse plant product harvesting monitoring method in the embodiments of this application.

[0053] Figure 2 This is a flowchart illustrating the weight graph generation process in the embodiments of this application.

[0054] Figure 3 This is a schematic diagram of the plant harvesting atlas in the embodiments of this application;

[0055] Figure 4 This is a schematic diagram of the structure of the knowledge graph-based greenhouse plant product harvesting monitoring system in the embodiments of this application. Detailed Implementation

[0056] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. Identical components are indicated by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom surface," "top surface," "inner," and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.

[0057] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0058] Hereinafter, the terms "first," "second," and other generic terms are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0059] refer to Figure 1 The image shows a specific implementation of the knowledge graph-based greenhouse plant product harvesting monitoring method of this application, including:

[0060] S101. Based on the multimodal images pre-acquired in the greenhouse, analyze the quality of each modal image through a preset quality assessment model, calculate the corresponding image value, and construct a value heat map.

[0061] S102. Based on the value heatmap, identify the key points of each modal image and filter out the corresponding key regions.

[0062] S103. Configure a dynamic fusion mechanism to fuse the key regions of each modal image, and dynamically adjust the fusion process through the key points to obtain a fused image;

[0063] S104. By analyzing the fused image through a pre-constructed plant harvesting atlas, the harvesting area is identified, and the plant products in the harvesting area are harvested.

[0064] Currently, smart agriculture is developing rapidly. In the process of greenhouse cultivation, it is necessary to analyze and identify the status of plant products in real time and harvest mature plants in a timely manner. In the greenhouse environment, plant status is affected by environmental interference such as light fluctuations and leaf shading. The environmental adaptability of plant status recognition based on single-modal image recognition in greenhouses is low and cannot meet the accuracy requirements for harvesting plant products in greenhouses.

[0065] In this embodiment, multimodal images are acquired in a greenhouse, including but not limited to RGB images, multispectral images, and thermal infrared images. Due to differences in lighting and equipment characteristics in the greenhouse environment, the effective information related to plant product harvesting contained in different modal images varies. RGB images reflect color and texture features, while multispectral images reflect the physiological state of crops. However, RGB images may be overexposed and distorted under strong light, while multispectral images are more stable. The quality of the pre-acquired multimodal images is assessed. The quality of each modal image is analyzed using a preset quality assessment model, simulating the game process of information contribution from multimodal images. The corresponding image value is calculated by constructing and solving the game matrix, and a value heatmap is constructed. Each pixel value in the value heatmap represents the information value density at that location.

[0066] It should be noted that by combining multimodal imagery for plant harvesting analysis, the impact of low quality in a single modality on the plant harvesting process can be avoided. By constructing a value heatmap, the pixel regions in each modality of imagery that are valuable for harvesting can be accurately identified, avoiding misjudgments of information caused by quality fluctuations or interference from invalid regions in a single modality. This provides high-quality regions for key area screening and image fusion, thereby improving the effectiveness and efficiency of the plant product harvesting process.

[0067] Specifically, high-value pixel regions in the value heatmap contain key information related to plant product harvesting, but these regions may overlap or be surrounded by low-value regions. Key points for each modality of the image are identified in the value heatmap, and corresponding key regions are selected. Value changes between pixels are calculated in the value heatmap, and peak points of value abrupt changes are located to identify the first set of key points. Key points with values ​​higher than a preset value threshold and small position dynamic adjustment factors are selected from the first set of key points to obtain the second set of key points. For each key point, adjacent pixels with similar values ​​are merged until the difference between the value of an adjacent pixel and the value of the key point exceeds a preset value difference threshold, and overlapping areas are removed to obtain the key region.

[0068] It is important to emphasize that by combining key points with region growth, key regions of varying shapes and sizes can be adaptively generated. By combining this with the actual distribution of plant products in the image, the cutting effect or incomplete information problems caused by using a fixed-size sliding window can be avoided. By screening key regions, the most critical information regions for harvesting can be extracted from multimodal images, redundant and repetitive low-value regions can be removed, the computational load of the image fusion process can be reduced, and the retained regions can be all core information related to the harvesting target, thereby improving the quality of image fusion.

[0069] Key regions of different modal images contain complementary acquired information, but spatial discrepancies exist between them. A dynamic fusion mechanism is configured to fuse the key regions of each modal image separately, dynamically adjusting the fusion process using key points to obtain a fused image. Based on the key regions of each modal image, the image space is divided into overlapping key regions, non-overlapping key regions, and non-key regions. Fusion weights for different modal images within each region are calculated based on region type and key point correspondence. Multimodal images are then weighted and dynamically adjusted according to their corresponding weights to obtain the fused image.

[0070] Unlike traditional simple weighted fusion methods, this embodiment uses dynamic weight fusion and dynamic adjustment to fully utilize the complementary advantages of multimodal information in key areas, while retaining the detailed features of a single modality image in individual key areas and suppressing the contribution of low-quality modalities in non-key areas. Combined with key point calibration, it eliminates image distortion caused by spatial deviation, improves the quality of the fused image, and makes the fused image more accurately reflect the true state of plant products, providing high-quality image evidence for harvesting area identification.

[0071] Specifically, a plant harvesting atlas is constructed based on agricultural knowledge such as plant product varieties, maturity standards, and harvesting rules. Features are extracted from the fused images, and these features are matched against the plant harvesting atlas to identify areas that meet the harvesting criteria, thus determining the harvesting areas. The characteristics of these harvesting areas are analyzed to generate harvesting strategies, controlling the harvesting equipment to harvest the plant products within those areas. By constructing a plant harvesting atlas, the accuracy and reliability of harvesting identification can be improved by combining plant products and their interrelationships within a greenhouse setting. Identifying harvesting areas through the plant harvesting atlas, combined with professional agricultural knowledge, enhances the accuracy of harvesting decisions and the adaptability to different plant products, thereby improving the efficiency and accuracy of the harvesting process.

[0072] This application analyzes image quality based on the game theory process of multimodal images, calculates image value, and constructs a value heatmap to quantitatively analyze image quality information. High-value key points are identified based on the value heatmap, and key regions are identified by merging adjacent effective pixels and removing overlapping areas. Different weight calculation methods are applied to different regions, and multimodal images are weighted and fused, with the fused image calibrated to improve its quality. Knowledge graphs are used to identify harvesting areas, enabling real-time harvesting of plant products. Joint analysis of multimodal images improves the adaptability of plant harvesting to changes in the greenhouse environment. Dynamic weighted fusion retains the advantageous features of each modality while removing low-quality information, improving the quality of the fused image. Combined with knowledge graphs, harvesting areas can be quickly identified, improving the accuracy and efficiency of greenhouse plant harvesting.

[0073] Furthermore, based on the pre-acquired multimodal images in the greenhouse, the quality of each modal image is analyzed using a pre-set quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed, including:

[0074] S201. Based on the pre-acquired multimodal images in the greenhouse, extract the features of each modal image using a preset feature extraction model, and construct the corresponding image feature vector.

[0075] S202. Based on the image feature vector, analyze the quality of each modal image through a preset quality assessment model, calculate the corresponding image value, and construct a value heatmap.

[0076] In this embodiment, based on pre-acquired multimodal images from the greenhouse, features of each modality are extracted using a preset feature extraction model to construct corresponding image feature vectors. The feature extraction model includes, but is not limited to, a deep convolutional neural network model pre-trained using a large amount of image data. Each modality image is input into the model, which extracts edge features, texture features, and other feature information. These features are then arranged and combined sequentially to construct the corresponding image feature vectors. By obtaining the corresponding feature vectors through feature extraction, relevant information for the quality assessment process can be obtained from the images. Calculations and analyses based on these feature vectors can eliminate analytical obstacles caused by differences in image data across different modalities. The feature extraction process can remove redundant information from the images, retain key features relevant to quality assessment, and improve the accuracy of image value calculation.

[0077] Specifically, the quality of different modal images varies due to the influence of greenhouse environment and equipment characteristics, resulting in differences in corresponding image value information. Based on image feature vectors, a pre-set quality assessment model is used to analyze the quality of each modal image, simulating the game relationship of information contribution. By constructing and solving the game matrix, the corresponding image value is calculated, and a value heatmap is constructed. By simulating the game process to calculate the global value score, the overall contribution of each image can be analyzed in a multimodal image environment, avoiding analytical biases caused by analyzing images individually. The constructed value heatmap can reflect the structural integrity and clarity of the image itself, and in connection with the needs of plant product harvesting, it can intuitively display image areas with high image quality and a significant impact on the harvesting process, providing accurate data support for the selection of key areas, thereby improving the efficiency of image analysis and the accuracy of harvesting decisions.

[0078] Furthermore, based on the image feature vectors, the quality of each modal image is analyzed using a preset quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed, including:

[0079] S301. Based on the image feature vector of each modal image, simulate the game process between each modal image through a preset quality assessment model to construct a game matrix;

[0080] S302. Based on the game matrix, calculate the image value of each modal image to obtain the corresponding global value score;

[0081] S303. For each modal image, segment it according to the image structure using a preset image segmentation model, analyze the analytical value of different image structures, and generate the corresponding structural value map.

[0082] S304. For each modal image, the analytical value of the harvesting area in the image is calculated using a preset harvesting analysis model, and a corresponding harvesting value map is generated.

[0083] S305. Multiply each pixel value in the structural value map by the corresponding global value score to obtain the first value map;

[0084] S306. Calculate the dynamic adjustment factor for each pixel position based on the harvest value map, and dynamically adjust the first value map according to the dynamic adjustment factor to generate a value heatmap for each modal image.

[0085] In this embodiment, multimodal images in a greenhouse harvesting scenario exhibit a competitive and complementary relationship in their contribution to identifying harvesting targets. For example, RGB images under strong light suffer from color information distortion and are at a disadvantage in the competition, while multispectral images are stable and have an advantage. Combining RGB and multispectral images can complement each other's missing information. Based on the image feature vector of each modal image, a preset quality assessment model is used to simulate the game process between each modal image, quantify and analyze the information contribution of each modal image, and construct a game matrix.

[0086] The analysis process of the quality assessment model includes: treating each modal image as a game participant, using key features in the image feature vector as corresponding game strategies, analyzing the effectiveness of features in identifying harvesting targets, evaluating game payoffs, simulating the game process between every two modal images. If the feature effectiveness of modal A is higher than that of modal B, the difference in corresponding game payoffs is calculated, resulting in a positive payoff for modal A and a negative payoff for modal B. If the features are complementary, both modal A and modal B receive positive payoffs. Based on the calculated game payoffs, a game matrix is ​​constructed according to the alliance situation between modalities. The game matrix constructed through the simulated game process can dynamically analyze the differences in information contribution of each modal image under the current greenhouse scenario. Analyzing different modal images together avoids the problem of ignoring complementary information in images when analyzing the quality of a single image alone. It can effectively identify images with high quality and those that provide complementary information not found in other modalities, avoiding the one-sidedness caused by quality assessment based solely on isolated features. The constructed game matrix provides data support for calculating the global value score of each modal image and for image fusion.

[0087] Specifically, based on the game theory matrix, the image value of each modality is calculated to obtain the corresponding global value score. The game theory matrix is ​​summed along the row dimension to obtain the total game payoff for each modality, reflecting its overall advantage relative to all other modalities. For each modality image, features with game payoffs greater than a preset threshold are selected. Corresponding high-quality features and their quantities are extracted from the modality feature vector, and the ratio of the number of high-quality features to the total number of features is calculated to obtain the high-quality coefficient. The game payoff threshold can be set according to the accuracy requirements of the image analysis process. The weights of the total game payoff and the high-quality coefficient are determined by analyzing the impact of the overall image advantage and individual image quality on the accuracy of image acquisition and identification in historical data. The total game payoff and the high-quality coefficient are then weighted and summed according to their respective weights to obtain the corresponding global value score. Calculating the global value score for each modality image based on the game theory matrix allows for quantitative analysis of the global value of each modality. By analyzing the overall image advantage and image quality, it is ensured that the calculated global value score reflects both the overall reliability advantage of the modality and does not ignore the key quality information of each modality, thus improving the accuracy of image quality analysis.

[0088] Within the same modal image, different structures exhibit significantly varying informational value for harvesting. These structures include, but are not limited to, mature and immature fruits, leaves, and branches. For each modal image, a pre-defined image segmentation model is used to segment the image according to its structure. This model includes, but is not limited to, a U-Net-based semantic segmentation model pre-trained using a large amount of historical image data. This segmentation yields different image structures for each modal image. The analytical value of each image structure is analyzed based on its distance from the fruit, and the reciprocal of the distance is calculated to obtain the structural value of each structure. A structural value map is then generated based on these structural values. By segmenting the image structurally and calculating its structural value, the value of areas including the fruit to be harvested can be increased, while suppressing background areas with sparse harvesting information. This provides a spatial value distribution reference for constructing a value heatmap, allowing focus on information-rich areas and improving the accuracy of the value assessment process.

[0089] Specifically, for each modal image, the analytical value of harvestable areas in the image is calculated using a pre-defined harvesting analysis model. This model includes, but is not limited to, a random forest regression model pre-trained using a large amount of historical harvesting data. The model learns the correlation analysis process between local features and qualified harvestable areas, analyzing the correlation between each pixel's location features and the harvestable area and calculating the harvesting correlation degree. The harvesting correlation degree ranges from 0 to 1, with higher values ​​indicating a greater probability that the corresponding pixel location belongs to a qualified harvestable area. The harvesting correlation degrees of all pixel locations are arranged according to their spatial location in the image, generating a corresponding harvesting value map. By analyzing the correlation between each pixel location and the harvestable area, the harvesting value of each location can be accurately evaluated and analyzed, precisely locating areas related to the harvesting process and enhancing the accuracy of image value assessment.

[0090] For each modal image, the value of each pixel in the structural value map is multiplied by its corresponding global value score to obtain the first value map. In the first value map, the value of each pixel is simultaneously influenced by the structural value of its local structure and the global value of its modality. By combining the structural value map and the global value score, the global value reflecting the relative importance between modalities can be combined with the value reflecting the internal structure of the image. This allows for the integration of local value assessment and global information assessment, enhancing regions in the first value map that possess both high global value scores and high structural value, thus improving the rationality of the value assessment process.

[0091] Specifically, a dynamic adjustment factor is calculated for each pixel location based on the harvest value map. The first value map is then dynamically adjusted according to this dynamic adjustment factor to generate a value heatmap for each modal image. The pixel correlation in the harvest value map is mapped to the dynamic adjustment factor according to the following rules: if the correlation is greater than or equal to 0.8, the dynamic adjustment factor is set to 1.2; if the correlation is greater than or equal to 0.5 and less than 0.8, the dynamic adjustment factor is set to 1; and if the correlation is less than 0.5, the dynamic adjustment factor is set to 0.5. The pixel value in the first value map is multiplied by the corresponding dynamic adjustment factor to obtain the adjusted pixel value, generating a value heatmap for each modal image. A higher pixel value in the value heatmap indicates a higher overall harvest value for that location. By correcting the first value map using dynamic adjustment factors, the structural value of pixel locations, global quality, and harvesting standards can be combined to obtain accurate value judgments. The constructed value heatmap accurately reflects the true harvesting status of plant products, improving the efficiency and effectiveness of the harvesting process.

[0092] Furthermore, based on the aforementioned value heatmap, key points are identified for each modal image, and corresponding key regions are selected, including:

[0093] S401. Based on the value heatmap of each modal image, identify the key points of each modal image respectively, and construct the first key region set by combining the corresponding key points;

[0094] S402. For each region in the first set of key regions, calculate the sum of pixel values ​​within the region to obtain the value integral.

[0095] S403. Sort each region from high to low according to the value points to obtain the region order;

[0096] S404. Based on the region order, calculate the spatial overlap between regions, filter out regions with a spatial overlap less than a preset overlap threshold, and obtain a set of key regions.

[0097] In this embodiment, based on the value heatmap of each modal image, key points of each modal image are identified, and a first set of key regions is constructed by combining the corresponding key points. By filtering out high-value key points in the value heatmap, the key point information is combined to transform into continuous and complete initial regions, avoiding the omission of key harvesting information. Threshold filtering can eliminate noisy regions with low correlation, reduce the interference of invalid regions on the plant harvesting analysis process, and improve the accuracy and efficiency of the harvesting process analysis results.

[0098] Specifically, for each region in the first set of key regions, the value integral is obtained by traversing all pixels through integral calculation and summing the value values ​​corresponding to the pixels on the value heatmap. By calculating the value integral, the overall value of the region can be quantified, and the regions can be sorted and filtered based on their overall contribution, thereby improving the rationality and accuracy of the key region selection process.

[0099] The regions are sorted in descending order of their calculated value scores to obtain a region order. By determining the region order, higher-value regions can be processed first, improving the effectiveness and efficiency of the processing. During the overlap screening process, high-value regions can be retained while low-value overlapping regions are eliminated, ensuring that high-quality harvesting regions are not lost and improving the quality of the key region set.

[0100] Specifically, based on the regional order, the spatial overlap between regions is calculated, and regions with spatial overlap less than a preset overlap threshold are selected to obtain a set of key regions. Regions are traversed in regional order, and the spatial overlap between the currently traversed region and other regions is calculated. The spatial overlap is obtained by calculating the ratio between the intersection area and the union area of ​​two regions. An overlap threshold is set according to the accuracy requirements of the plant harvesting process. If the spatial overlap is greater than or equal to the overlap threshold, the region is considered redundant and discarded. Regions with spatial overlap less than the overlap threshold are selected and combined to obtain the set of key regions. By filtering based on spatial overlap, redundant information in the set of key regions can be eliminated, reducing the computational load of image fusion, avoiding feature conflicts in overlapping regions, and improving the quality and accuracy of the fused image.

[0101] Furthermore, based on the value heatmap of each modal image, key points of each modal image are identified, and combined with the corresponding key points, a first set of key regions is constructed, including:

[0102] S501. Based on the value heatmap of each modal image, identify the corresponding key points through a preset key point recognition model to obtain the first set of key points;

[0103] S502. Select key points from the first key point set whose value is higher than a preset value threshold and whose dynamic adjustment factor of the corresponding pixel position is less than a preset adjustment threshold, and use them as the second key point set.

[0104] S503. For each key point in the second key point set, adjacent pixels are merged using a preset region generation model until the difference between the value of the adjacent pixel and the value of the key point is greater than a preset value difference threshold, thereby obtaining a first key region set.

[0105] In this embodiment, based on the value heatmap of each modal image, corresponding key points are identified using a preset key point recognition model to obtain a first set of key points. The key point recognition model includes, but is not limited to, an improved Harris corner detection model. This model performs Gaussian smoothing on the value heatmap to reduce noise interference, preserves value gradient features, calculates the difference in value change in the neighborhood of each pixel to obtain a value change gradient matrix, and sets a basic value threshold based on the accuracy requirements of the plant product harvesting process. Key points with significantly higher values ​​than surrounding pixels and value change differences greater than the basic value threshold are selected as the first set of key points. By identifying key points, feature mutation points can be accurately located. Low-value noise points are filtered out using the value threshold, avoiding misclassification of mutation points in irrelevant areas such as leaf edges and background impurities as key points, thus improving the accuracy and quality of the key point recognition process.

[0106] Specifically, from the first set of key points, key points with values ​​higher than a preset value threshold and corresponding pixel positions with dynamic adjustment factors less than a preset adjustment threshold are selected to form the second set of key points. For each key point in the first set, its pixel value in the value heatmap is extracted. Value and adjustment thresholds are set according to the accuracy requirements of the plant product harvesting process. Key points with values ​​higher than the preset value threshold and corresponding pixel positions with dynamic adjustment factors less than the preset adjustment threshold are then selected to obtain the second set of key points. By using the value threshold for selection, meaningless points with local extrema but low absolute value can be removed. By using the dynamic adjustment factor threshold for selection, points with high value but temporary interference or instability in the task context can be effectively removed, improving the value and stability of the retained key points, thereby improving the accuracy and quality of the key areas.

[0107] Specifically, for each keypoint in the second keypoint set, adjacent pixels are merged using a preset region generation model until the difference between the value of an adjacent pixel and the value of the keypoint is greater than a preset value difference threshold, thus obtaining a first key region set. The region generation model uses a region growth algorithm, with each keypoint in the second keypoint set as a seed point. All adjacent pixels of the seed point are checked. In this embodiment, 4-connected neighborhoods are used. If the absolute difference between the value of an adjacent pixel and the value of the seed point is less than or equal to the preset value difference threshold, it indicates that the pixel and the seed point belong to the same value uniform region, and the corresponding pixel is merged into it. The merged pixel is used as a new seed point, and the above growth process is iteratively repeated to continuously expand the region boundary. When the value difference between all adjacent pixels and the current region boundary point is greater than the value difference threshold, the growth process terminates, and the first key region set is obtained. The value difference threshold is set according to the accuracy requirements of the plant product harvesting process.

[0108] It should be noted that by using a region growing algorithm to calculate the value difference threshold to control the expansion process of the region, the value within the generated region is relatively uniform, which can more accurately reflect the actual contours of plant organs in the image. The resulting first key region set not only accurately covers the high-value region, but also has more natural and accurate boundaries, thereby improving the quality of the fused image.

[0109] Furthermore, a dynamic fusion mechanism is configured to fuse key regions of each modality image separately, and the fusion process is dynamically adjusted based on the key points to obtain a fused image, including:

[0110] S601. Based on the key regions of each modal image, filter out the overlapping key regions, non-overlapping key regions, and non-key regions between different modal images, calculate the weight of each modal image in the corresponding region, and generate a weight map.

[0111] S602. According to the weight map, the pixels of each modal image are weighted and fused to obtain the first fused image;

[0112] S603. Based on the key points of each modal image, dynamically adjust the pixels of the corresponding region in the first fused image to obtain the fused image.

[0113] In this embodiment, based on the key regions of each modal image, overlapping key regions, non-overlapping key regions, and non-key regions between different modal images are selected. The weight of each modal image in the corresponding region is calculated to generate a weight map. By dividing the region and formulating different weighting strategies, the advantages of multimodal data can be complemented in the information overlapping region. In the unique information region, it can be ensured that important features are not affected by other modal data. In the ordinary region, higher quality modal data is selected. By partitioning and weighting, dynamic optimization can be performed according to the characteristics of image information during the fusion process, thereby improving the quality of the fused image and improving the efficiency of image information utilization.

[0114] Specifically, according to the weight map, the pixels of each modal image are weighted and fused to obtain the first fused image. For each target pixel position in the image, the pixel values ​​of all modal images at that position are read. Simultaneously, the weight value of each modal image at the corresponding position is read from the generated weight map. The pixel value of each modality is multiplied by the corresponding weight value, and the weighted results of all modalities are summed to obtain the fused pixel value at that position. This process is repeated for each pixel position in the image to obtain the first fused image. By weighting the pixels, it is ensured that the features of the corresponding modality are highlighted in areas with high weights. In areas where weights transition, the information of different modalities can transition naturally, avoiding fusion boundary artifacts caused by abrupt weight changes. This preserves the detailed information of the original data while integrating and enhancing the information through weighted summation.

[0115] Specifically, based on the key points of each modal image, the pixels in the corresponding region of the first fused image are dynamically adjusted to obtain the fused image. Key points are matched, and key point pairs corresponding to the same physical location are selected from the key point sets of different modal images. The coordinate displacement deviation between the key point pairs is calculated, constructing a sparse displacement field. Radial basis function interpolation is used to expand the sparse displacement field into a dense displacement field covering the entire image area. Based on the dense displacement field, the first fused image is resampled using bilinear interpolation, and precise displacement compensation is performed on each pixel position to obtain a fused image with higher geometric alignment accuracy. Through key point matching and displacement field correction, pixel deviations between multimodal images can be effectively compensated, eliminating ghosting and blurring phenomena present in the initial fused image, improving the quality of the fused image, and thus enhancing the accuracy of the plant product harvesting process.

[0116] like Figure 2 As shown, based on the key regions of each modal image, overlapping key regions, non-overlapping key regions, and non-key regions between different modal images are selected. The weight of each modal image in the corresponding region is calculated to generate a weight map, including:

[0117] S701. Based on the key regions of each modal image, filter out the overlapping key regions, non-overlapping key regions, and non-key regions between different modal images;

[0118] S702. For overlapping key regions, analyze the value ratio of the corresponding key points and calculate the first weight of the corresponding modal image.

[0119] S703. For non-overlapping key regions, set the corresponding modal image weight to 1 and the weights of the other modal images to 0 to obtain the second weight.

[0120] S704. For non-critical areas, analyze the image feature vectors of the modal images and calculate the third weight of the corresponding modal images.

[0121] S705. Combine the first weight, the second weight, and the third weight to generate a weight graph.

[0122] In this embodiment, based on the key regions of each modal image, overlapping key regions, non-overlapping key regions, and non-key regions between different modal images are screened. The spatial positions of the identified key regions in each modal image are compared. If the same key region includes two or more modal images, it is classified as an overlapping key region; if the same key region includes only one modal image, it is classified as a non-overlapping key region. Regions outside the overlapping and non-overlapping key regions are classified as non-key regions. By classifying overlapping, non-overlapping, and non-key regions, regions with complementary multimodal information, regions including information unique to a single modality, and regions with low information density can be identified. This provides regional references for the fusion of different modal image regions, improving the accuracy of the fusion process.

[0123] For overlapping key regions, the value ratios of corresponding key points are analyzed to calculate the first weight of the corresponding modal image. For each overlapping key region, the key points of each participating modality within that region are located. For each modality, the values ​​of all key points in the overlapping region are summed to obtain the total key point value of that modality within the overlapping key region. The ratio of this total key point value to the sum of the total key point values ​​of all modalities in the overlapping key region is calculated to obtain the first weight of the corresponding modal image. By calculating the key point value ratio for weight allocation, higher weights can be assigned to modalities with higher information quality and greater value, while allowing other modal images to provide supplementary information. This ensures that the fusion result reflects high-quality information features and improves the quality of the fused images in overlapping regions.

[0124] For non-overlapping key regions, the corresponding modal image weight is set to 1, and the weights of other modal images are set to 0, resulting in a second weight. The second weight retains only the modal image information of the non-overlapping key regions, effectively protecting the unique feature information of the corresponding modality, avoiding interference from other modal image information, and improving the quality of the fused image.

[0125] For non-critical regions, the image feature vectors of the modal images are analyzed, and the third weight of the corresponding modal image is calculated. For each pixel position in the non-critical region, the image feature vectors of each modal image are extracted, and the feature vectors are comprehensively analyzed, including but not limited to calculating the magnitude of the feature amplitude and analyzing the variance of each dimension of the feature. The analysis and calculation results are superimposed and summed to evaluate the overall image quality of each modality in the vicinity of that position. The third weight is assigned to each modality according to the proportion of the superimposed summation result, with higher weights assigned to images with higher superimposed summation results. By assigning weights based on the quality assessment of feature vectors, modal data with lower noise, higher sharpness, or better texture preservation can be selected as the main fusion modal data in background areas where information value is not prominent. This can improve the fusion quality of non-critical regions, avoid the introduction of noise by low-quality modalities, and thus improve the quality of the fused image.

[0126] Specifically, a weight map is generated by combining the first, second, and third weights. Based on the corresponding pixel positions, the calculated first, second, and third weights are arranged into their respective image spatial locations, resulting in a set of weight maps. Each weight map corresponds to a modality, reflecting the specific weight value of that modality when participating in fusion at each pixel position. Generating weight maps ensures optimized fusion of high-value critical areas, information protection of unique areas, and overall quality improvement of ordinary areas. This ensures that each image area can adopt the most suitable fusion method, improving the overall quality of the fused image and enhancing the accuracy and efficiency of the plant product harvesting process.

[0127] Furthermore, based on the key points of each modal image, the pixels of the corresponding regions in the first fused image are dynamically adjusted to obtain the fused image, including:

[0128] S801. Analyze the correspondence between key points in each modal image, match the key points, and obtain the key point matching results;

[0129] S802. Based on the key point matching results, calculate the displacement deviation of the corresponding pixel position to obtain the displacement deviation result;

[0130] S803. According to the displacement deviation result, the pixels in the corresponding area of ​​the first fused image are dynamically adjusted to obtain the fused image.

[0131] In this embodiment, the correspondence between key points in each modal image is analyzed, and key points are matched to obtain key point matching results. Using a pre-trained ORB model with a large amount of regional feature information data, feature descriptors are generated for key points in each modal image. The model extracts feature information such as texture and gradient that can characterize the local area around the key point, and combines multi-dimensional feature data to generate corresponding feature descriptors. The feature Euclidean distance between key point descriptors of different modal images is calculated, and key points are matched according to the feature Euclidean distance; the shorter the distance, the higher the matching degree, thus obtaining key point matching results. Key point matching provides positional references for the fusion process of different modalities, ensuring the accurate spatial correlation of multimodal information.

[0132] Specifically, based on the keypoint matching results, the displacement deviation of the corresponding pixel position is calculated to obtain the displacement deviation result. For each matched keypoint pair, the coordinate difference between the corresponding keypoints in their respective image coordinate systems is calculated to obtain the displacement vector at the matching point. Since the keypoint matching results are sparse, the continuous displacement change of each pixel position is reconstructed through radial basis function interpolation based on the sparse displacement vector sample points, generating a dense displacement field covering the entire image. By establishing a dense displacement field, the complex local geometric differences between multimodal images can be accurately reflected, including various deformation modes such as translation and rotation. The displacement description provides positional reference information for image geometric correction, improving the accuracy and fusion quality of the image fusion process.

[0133] Specifically, based on the displacement deviation results, the pixels in the corresponding regions of the first fused image are dynamically adjusted to obtain the fused image. Using the dense displacement field generated from the displacement deviation results, the source pixel position corresponding to the target pixel position in the first fused image is found at each target pixel position in the fused image. Bilinear interpolation is used to calculate the target pixel value based on the known pixel values ​​around the source pixel position. This process is repeated across all pixel positions in the final fused image to perform geometric correction, generating the fused image. Through precise geometric correction based on the displacement field, positional deviations caused by factors such as shooting angle, lens distortion, or temporal variations can be compensated for. This effectively solves quality problems such as ghosting and edge blurring in the initial fused image, improves the spatial consistency and detail of the fused image, and enhances the quality of the fused image as well as the accuracy and reliability of acquisition decisions.

[0134] Furthermore, by analyzing the fused image using a pre-constructed plant harvesting atlas, harvesting areas are identified, and plant products from these areas are harvested, including:

[0135] S901. Based on the fused imagery, the harvesting area is identified using a pre-constructed plant harvesting atlas;

[0136] S902. Based on the harvesting area, analyze the harvesting characteristics through a preset harvesting analysis model, generate corresponding harvesting results, and harvest the plant products in the harvesting area.

[0137] In this embodiment, the harvesting area is identified based on the fused imagery using a pre-constructed plant harvesting atlas; such as Figure 3As shown, based on knowledge of plant harvesting, the constraints of the harvesting process are analyzed to construct target entities. These entities include plant product varieties, maturity levels, and harvesting rules. The correspondence between maturity levels and harvesting is analyzed to construct entity relationships, resulting in a plant harvesting atlas. A CNN model is trained using extensive historical plant harvesting data to obtain a pre-trained CNN model. The model extracts features from fused images to generate fused image feature vectors, which are then matched against the plant harvesting atlas to analyze whether the corresponding plant products and maturity levels meet the harvesting criteria. The model outputs the harvesting areas that meet the criteria. By constructing a plant harvesting atlas and analyzing the semantic attributes and interrelationships of entity objects, the accuracy and reliability of harvesting area identification are significantly improved.

[0138] Specifically, based on the harvesting area, a pre-set harvesting analysis model analyzes harvesting characteristics and generates corresponding harvesting results for harvesting plant products in the area. The harvesting analysis model includes, but is not limited to, a convolutional neural network (CNN) model. This model is trained using a large amount of harvesting feature data to obtain a pre-trained CNN model. The model analyzes the harvesting characteristics of the harvesting area and outputs corresponding harvesting results, including but not limited to harvesting order and harvesting location. Analyzing the harvesting area using this model can improve the success rate of the harvesting process and product quality, effectively increasing the overall harvesting efficiency of products in the greenhouse.

[0139] like Figure 4 As shown, a knowledge graph-based greenhouse plant product harvesting monitoring system is used to implement a knowledge graph-based method for monitoring greenhouse plant product harvesting, including:

[0140] The value heatmap construction module analyzes the quality of each modal image based on the pre-acquired multimodal images in the greenhouse through a preset quality assessment model, calculates the corresponding image value, and constructs a value heatmap.

[0141] The key region identification module identifies key points for each modal image based on the value heatmap and filters out the corresponding key regions.

[0142] The image fusion module is configured with a dynamic fusion mechanism, which merges the key regions of each modality image separately, and dynamically adjusts the fusion process through the key points to obtain the fused image;

[0143] The plant product harvesting module analyzes the fused image using a pre-constructed plant harvesting atlas to identify harvesting areas and harvest plant products from those areas.

[0144] In this embodiment, the value heatmap construction module obtains image feature vectors through a preset feature extraction model, calculates the global value score by simulating the game process between images of different modalities using a preset quality assessment model, and generates a value heatmap by combining the structural value map and the harvesting value map. Through image game analysis, high-quality, high-value image regions are selected, providing corresponding regional references for image analysis and improving image processing efficiency. The key region identification module extracts value peak points and key regions through a preset key point identification model. A set of key regions is obtained by sorting by value integral and filtering by spatial overlap. Extracting key regions reduces data processing volume, provides high-value regions for image fusion, and improves the fusion effect and the accuracy of harvesting region identification.

[0145] Specifically, the image fusion module divides the image into overlapping key regions, non-overlapping key regions, and non-key regions, calculates fusion weights for each, generates a weighted map for preliminary weighted fusion, and corrects the preliminary fused image based on key point matching results. Through partitioned weight allocation and correction, the complementary nature of multimodal information is utilized while effectively eliminating spatial misalignment between images, generating a fused image that is complete in information, clear in detail, and spatially consistent, thus improving the quality of the fused image. The plant product harvesting module analyzes and identifies the fused image using a pre-constructed plant harvesting atlas to obtain areas that meet harvesting standards, and uses a harvesting analysis model to refine the analysis of harvesting characteristics for plant harvesting. The identification and analysis of harvesting areas can improve the accuracy and efficiency of harvesting decisions, and increase the success rate and efficiency of plant harvesting operations.

[0146] The above description is merely a preferred embodiment of this application. The scope of protection of this application is not limited to the above embodiments. All technical solutions falling within the scope of this application's concept are within the scope of protection of this application. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of this application should also be considered within the scope of protection of this application.

Claims

1. A knowledge graph-based method for monitoring the harvesting of greenhouse plant products, characterized in that, include: Based on the multimodal images pre-acquired in the greenhouse, the quality of each modal image is analyzed through a pre-set quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed. Based on the aforementioned value heatmap, key points of each modal image are identified, and corresponding key regions are selected. Based on the key regions of each modal image, overlapping key regions, non-overlapping key regions, and non-key regions between different modal images are selected; For overlapping key regions, analyze the value ratio of corresponding key points and calculate the first weight of the corresponding modal image. For non-overlapping key regions, the corresponding modal image weight is set to 1, and the weights of the other modal images are set to 0, thus obtaining the second weight. For non-critical regions, analyze the image feature vectors of the modal images and calculate the third weight of the corresponding modal images; By combining the first weight, the second weight, and the third weight, a weight graph is generated; According to the weight map, the pixels of each modal image are weighted and fused to obtain the first fused image; Analyze the correspondence between key points in each modal image, match the key points, and obtain the key point matching results; Based on the key point matching results, the displacement deviation of the corresponding pixel position is calculated to obtain the displacement deviation result; Based on the displacement deviation results, the pixels in the corresponding region of the first fused image are dynamically adjusted to obtain the fused image. By analyzing the fused images using a pre-constructed plant harvesting atlas, harvesting areas are identified, and plant products from these areas are harvested.

2. The method for monitoring greenhouse plant product harvesting based on knowledge graphs according to claim 1, characterized in that, Based on the pre-acquired multimodal images in the greenhouse, the quality of each modal image is analyzed using a pre-set quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed, including: Based on the multimodal images pre-acquired in the greenhouse, the features of each modal image are extracted using a preset feature extraction model, and the corresponding image feature vector is constructed. Based on the image feature vectors, the quality of each modal image is analyzed using a preset quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed.

3. The method for monitoring greenhouse plant product harvesting based on knowledge graphs according to claim 2, characterized in that, Based on the image feature vectors, the quality of each modal image is analyzed using a preset quality assessment model, the corresponding image value is calculated, and a value heatmap is constructed, including: Based on the image feature vector of each modal image, the game process between each modal image is simulated through a preset quality assessment model to construct a game matrix; Based on the game matrix, the image value of each modal image is calculated to obtain the corresponding global value score; For each modal image, the image is segmented according to its structure using a pre-defined image segmentation model. The analytical value of different image structures is analyzed, and a corresponding structural value map is generated. For each modal image, the analytical value of the harvesting area in the image is calculated using a preset harvesting analysis model, and a corresponding harvesting value map is generated. The first value map is obtained by multiplying each pixel value in the structural value map by the corresponding global value score; Based on the harvest value map, a dynamic adjustment factor is calculated for each pixel location. The first value map is then dynamically adjusted according to the dynamic adjustment factor to generate a value heatmap for each modal image.

4. The method for monitoring greenhouse plant product harvesting based on knowledge graphs according to claim 1, characterized in that, Based on the aforementioned value heatmap, key points for each modal image are identified, and corresponding key regions are selected, including: Based on the value heatmap of each modal image, the key points of each modal image are identified, and the first key region set is constructed by combining the corresponding key points. For each region in the first set of key regions, calculate the sum of pixel values ​​within the region to obtain the value integral; The regions are sorted from highest to lowest according to their value scores to obtain the region order; Based on the region order, the spatial overlap between regions is calculated, and regions with a spatial overlap less than a preset overlap threshold are selected to obtain a set of key regions.

5. The method for monitoring greenhouse plant product harvesting based on knowledge graphs according to claim 4, characterized in that, The value heatmap based on each modal image identifies key points for each modal image, and constructs a first set of key regions by combining the corresponding key points, including: Based on the value heatmap of each modal image, the corresponding key points are identified by a preset key point recognition model to obtain the first set of key points; From the first set of key points, select key points whose value is higher than a preset value threshold and whose dynamic adjustment factor at the corresponding pixel position is less than a preset adjustment threshold, and use them as the second set of key points. For each key point in the second key point set, adjacent pixels are merged using a preset region generation model until the difference between the value of the adjacent pixel and the value of the key point is greater than a preset value difference threshold, thus obtaining the first key region set.

6. The method for monitoring greenhouse plant product harvesting based on knowledge graphs according to claim 1, characterized in that, The process involves analyzing the fused image using a pre-constructed plant harvesting atlas to identify harvesting areas and harvesting plant products from those areas, including: Based on the fused imagery, the harvesting area is identified using a pre-constructed plant harvesting atlas; Based on the harvesting area, the harvesting characteristics are analyzed using a preset harvesting analysis model, and corresponding harvesting results are generated to harvest the plant products in the harvesting area.

7. A knowledge graph-based greenhouse plant product harvesting monitoring system, characterized in that, The method for monitoring the harvesting of greenhouse plant products based on knowledge graphs as described in any one of claims 1 to 6 includes: The value heatmap construction module analyzes the quality of each modal image based on the pre-acquired multimodal images in the greenhouse through a preset quality assessment model, calculates the corresponding image value, and constructs a value heatmap. The key region identification module identifies key points for each modal image based on the value heatmap and filters out the corresponding key regions. The image fusion module is configured with a dynamic fusion mechanism, which merges the key regions of each modal image and dynamically adjusts the fusion process through the key points to obtain a fused image. The plant product harvesting module analyzes the fused image using a pre-constructed plant harvesting atlas to identify harvesting areas and harvest plant products from those areas.

Citation Information

Patent Citations

  • Tea tender shoot picking point positioning method based on deep learning algorithm

    CN114842188A

  • Greenhouse pest and disease damage intelligent monitoring and early warning system and method based on Internet of Things

    CN120452173A