Feature fingerprint-based agricultural product origin identification traceability system and method
By combining explicit and implicit feature fingerprinting with a successive verification mode, the problem of inaccurate traceability of agricultural product origins has been solved, achieving more efficient and accurate identification and traceability of agricultural product origins.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2026-03-31
AI Technical Summary
Existing agricultural product origin traceability technologies cannot guarantee the accuracy of traceability, leading to inaccurate classification of agricultural product origins and potentially causing misjudgments.
A method for identifying and tracing the origin of agricultural products based on feature fingerprints is adopted. The agricultural product sample set is divided into primary and secondary categories by explicit and implicit feature fingerprint markings, and the final origin traceability results are obtained by combining successive verification or individual verification modes.
It improves the accuracy and efficiency of agricultural product origin identification, reduces misjudgments caused by single characteristics or single analysis methods, enhances the reliability and objectivity of classification results, and ensures the quality control and safety of agricultural products.
Smart Images

Figure CN119807816B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural product origin identification and traceability technology, and in particular to an agricultural product origin identification and traceability system and method based on feature fingerprints. Background Technology
[0002] Currently, when tracing the origin of agricultural products based on their characteristics, the inaccurate classification of their origins due to one-sided descriptions of the product's features affects the accuracy of the traceability process. Furthermore, a traceability process that does not involve classification and accuracy assessment may lead to erroneous results. Agricultural products may have similar characteristics, and it may be difficult to distinguish between products from different origins. Direct traceability without classification may result in misclassifying agricultural products from different origins as having the same source.
[0003] Chinese Patent Application Publication No. CN113406245A discloses a method for soybean origin traceability identification based on a combination of MALDI□TOF / TOF and IRMS technologies. This method utilizes MALDI□TOF / TOF, a novel soft ionization biomass spectrometry technique, combined with stable isotope ratio analysis. It establishes a soybean origin traceability identification method based on the fusion of MALDI□TOF / TOF and stable isotope content analysis. By referencing high-resolution mass spectrometry data of triglyceride compounds in soybean oil samples and stable isotope ratios of water-soluble proteins in soybean flour samples, the two data sets are fused and imported into the soybean origin traceability identification model to predict the origin of the soybeans being tested.
[0004] This shows that current agricultural product origin traceability technology cannot guarantee the accuracy of traceability. Summary of the Invention
[0005] Therefore, the purpose of this invention is to provide an agricultural product origin identification and traceability system and method based on feature fingerprints, which can overcome the problem that current agricultural product origin traceability technologies cannot guarantee the accuracy of traceability.
[0006] To achieve the above objectives, the present invention provides a method for identifying and tracing the origin of agricultural products based on feature fingerprints, comprising:
[0007] Collect target agricultural product samples with a preset initial sample size to obtain the target sample set;
[0008] The target sample set is fingerprinted based on the dominant features of the target agricultural products to obtain dominant labeling results. The target sample set is then divided into several sample clusters based on the dominant labeling results.
[0009] The target sample set is fingerprinted with latent feature based on the latent features of the target agricultural products to obtain latent labeling results. The target sample set is then divided into two categories based on the latent labeling results to obtain several binary sample clusters.
[0010] Compare the results of the first-class and second-class classifications to determine whether to enable the successive verification mode or the individual verification mode.
[0011] Based on the successive verification mode, the verification results of each step are obtained, and it is determined whether to output the actual place of origin traceability results based on the verification results of each step.
[0012] Based on the separate verification mode, the first-class classification result and the second-class classification result are verified separately to obtain the first-class verification result and the second-class verification result. Based on the first-class verification result and the second-class verification result, it is determined whether to output the actual place of origin traceability result.
[0013] Furthermore, the process of performing dominant feature fingerprinting on the target sample set based on the dominant features of the target agricultural products to obtain dominant labeling results, and then classifying the target sample set into several sample clusters based on the dominant labeling results, includes:
[0014] Based on the dominant characteristics of the target agricultural products, determine each dominant feature vector, and mark each target sample point with a dominant feature fingerprint based on each dominant feature vector to determine the dominant identifier of each target sample point.
[0015] Based on the explicit identifier, each target sample point is classified into one class according to the initial explicit classification criteria to obtain a sample cluster corresponding to different origin types.
[0016] The target sample set includes several target sample points; the sample cluster includes several target sample points with explicit identifiers.
[0017] Furthermore, the process of performing latent feature fingerprinting on the target sample set based on the latent features of the target agricultural products to obtain latent labeling results, and then performing binary classification on the target sample set based on the latent labeling results to obtain several binary sample clusters includes:
[0018] The latent feature vector is determined based on the latent characteristics of the target agricultural product, and the latent feature fingerprint is marked on each target sample point based on the latent feature vector to determine the latent identifier of each target sample point.
[0019] Based on the aforementioned implicit identifier, each target sample point is divided into two categories according to the initial implicit classification criteria to obtain each category of sample clusters corresponding to different origin types;
[0020] The second type of sample cluster includes several target sample points with implicit identifiers.
[0021] Furthermore, the process of comparing the results of the first-class and second-class classifications to determine whether to initiate the successive verification mode or the individual verification mode includes:
[0022] Identify a class partition to determine the number of partitions and the size of each class's sample cluster;
[0023] Identify binary partitions to determine the number of binary partitions and the size of each binary sample cluster;
[0024] A comparison is made based on the number of classifications in the first category and the number of classifications in the second category to determine whether to activate the supplementary comparison mode.
[0025] Based on enabling the supplementary comparison mode, a second comparison is performed according to the capacity of each type I sample cluster and the capacity of each type II sample cluster, and the successive verification mode or the individual verification mode is determined to be enabled based on the results of the second comparison.
[0026] The first-class partitioning result includes: the number of first-class partitions and the capacity of each first-class sample cluster; the second-class partitioning result includes: the number of second-class partitions and the capacity of each second-class sample cluster.
[0027] Furthermore, the process of performing a secondary comparison based on the size of each type I sample cluster and the size of each type II sample cluster includes:
[0028] Obtain the largest and smallest class sample sizes among the sizes of each class of sample clusters;
[0029] Obtain the maximum and minimum binary sample sizes among the sizes of each binary sample cluster;
[0030] The largest sample size of the first class is compared with the largest sample size of the second class to obtain the maximum comparison result;
[0031] The minimum sample size of the first class is compared with the minimum sample size of the second class to obtain the minimum comparison result;
[0032] Determine the number of items that meet a single judgment condition in the maximum comparison result and the minimum comparison result.
[0033] Furthermore, the process of determining whether to output the actual place of origin traceability result based on the results of each verification includes:
[0034] Based on the successive verification mode, the initial sample size is adjusted according to the first verification standard to obtain the verification sample size;
[0035] The verification sample size is used to perform a primary verification based on the dominant and latent characteristics of the target agricultural product, resulting in several primary verification first-class sample clusters and several primary verification second-class sample clusters.
[0036] The number of explicit partitions and the number of implicit partitions in a single validation are determined based on the first-class sample clusters and the second-class sample clusters in each validation.
[0037] The verification result is determined based on the number of explicit partitions and the number of implicit partitions in a single verification.
[0038] Furthermore, the process of determining whether to output the actual place of origin traceability result based on the results of each verification also includes:
[0039] Based on the aforementioned successive verification mode, the target classification standard is adjusted according to the secondary verification standard to obtain the actual classification standard;
[0040] The initial sample size is divided into several secondary verification sample clusters and several secondary verification sample clusters of type I.
[0041] The number of explicit partitions and the number of implicit partitions in the secondary validation are determined based on each type I sample cluster and each type II sample cluster in the secondary validation.
[0042] The secondary verification result is determined based on the number of explicit partitions and the number of implicit partitions in the secondary verification.
[0043] Determine whether the first verification result and the second verification result are consistent. If the verification results are consistent, output the actual place of origin traceability result.
[0044] Furthermore, the process of determining the verification result based on the number of explicit partitions and the number of implicit partitions in a single verification includes:
[0045] The first verification result is determined to be either consistent with the initial result or inconsistent with the initial result based on the absolute value of the second difference and the preset difference evaluation value.
[0046] If the absolute value of the second difference is greater than or equal to the difference evaluation value, then the first verification result is determined to be consistent with the initial result.
[0047] If the absolute value of the second difference is less than the difference evaluation value, then the first verification result is determined to be inconsistent with the initial result.
[0048] Wherein, the absolute value of the second difference is the absolute value of the difference between the number of explicit partitions in the first verification and the number of implicit partitions in the first verification.
[0049] Furthermore, the process of determining the secondary verification result based on the number of explicit partitions and the number of implicit partitions in the secondary verification includes:
[0050] Based on the absolute value of the third difference and the difference evaluation value, the result of the secondary verification is determined to be either consistent with the initial result or inconsistent with the initial result.
[0051] If the absolute value of the third difference is greater than or equal to the difference evaluation value, then the result of the second verification is determined to be consistent with the initial result.
[0052] If the absolute value of the third difference is less than the difference evaluation value, then the result of the second verification is determined to be inconsistent with the initial result.
[0053] The absolute value of the third difference is the absolute value of the difference between the number of explicit partitions in the secondary verification and the number of implicit partitions in the secondary verification.
[0054] Another aspect of the present invention provides an agricultural product origin identification and traceability system based on feature fingerprints, comprising:
[0055] The sample collection module is used to collect target agricultural product samples according to a preset initial sample size to obtain a target sample set;
[0056] The dominant feature analysis module is used to perform dominant feature analysis on the target sample set to obtain dominant feature fingerprints;
[0057] The latent feature analysis module is used to perform latent feature analysis on the target sample set to obtain latent feature fingerprints;
[0058] The partitioning and clustering module is used to partition samples into primary and secondary classes based on each feature fingerprint, so as to obtain several primary sample clusters and several secondary sample clusters.
[0059] The verification mode selection module is used to compare the results of the first-class partitioning and the second-class partitioning to determine whether to use the successive verification mode or the individual verification mode.
[0060] The successive verification module is used to perform multiple divisions through multiple verifications to obtain the verification results of each verification, and decide whether to output the actual place of origin traceability results based on the verification results of each verification.
[0061] The separate verification module is used to independently verify the first-class and second-class classification results, and combine the two verification results to determine whether to output the actual place of origin traceability results.
[0062] Compared with existing technologies, the beneficial effects of this invention are that by combining dominant physical characteristics and latent chemical characteristics, it can more comprehensively describe the characteristics of soybeans, thereby improving the accuracy of origin identification. It uses dominant and latent characteristics to classify soybeans into primary and secondary categories respectively, and verifies the results through successive verification or individual verification modes, enhancing the reliability of the classification results. Through a multi-level verification process, it reduces misjudgments caused by single characteristics or single analysis methods, improving the overall accuracy of traceability. By quantifying the physical and chemical characteristics of soybeans, origin identification becomes more objective, reducing the uncertainty caused by subjective judgment. Accurate origin traceability results can help control agricultural product quality, ensure product quality and safety, and improve the accuracy and efficiency of soybean origin identification.
[0063] By comprehensively describing soybean samples, the accuracy of classification is improved. Dominant identifiers, acting as characteristic fingerprints, can uniquely identify the dominant features of each target sample point, facilitating subsequent classification and traceability. Based on the initial dominant classification criteria using dominant identifiers, soybean samples with similar characteristics can be quickly grouped into one category, improving classification efficiency. Dominant features are closely related to the soybean's growth environment, and classification can reflect the soybean's adaptability to the environment, helping to understand the characteristics of different production areas. Through dominant identifiers, the origin of soybean samples can be traced, enhancing product traceability. Effectively utilizing the dominant features of soybeans for classification and traceability improves the accuracy and efficiency of classification and traceability.
[0064] By analyzing the latent characteristics of soybean mineral content, we can gain a deeper understanding of the intrinsic quality of soybeans and the impact of their growing environment, thereby improving the depth and accuracy of classification. Combining dominant and latent characteristics allows for the classification of soybean samples from multiple dimensions, enhancing the comprehensiveness and reliability of the classification. Latent identifiers, as a condensed representation of mineral content, can uniquely identify the mineral composition of each target sample point, facilitating accurate classification and traceability. Mineral content is closely related to the soybean growing environment. Classification using latent characteristics can analyze the adaptability of soybeans to specific soil and climatic conditions, providing guidance for agricultural production. Latent characteristic fingerprinting helps to more accurately track the origin of soybeans, especially when dominant characteristics are similar, latent characteristics can provide additional distinguishing information.
[0065] By comparing the results of the first and second category classifications, it is possible to effectively determine whether to adopt a successive verification mode or a single verification mode, thereby improving the efficiency of the entire verification process. By comparing the number and capacity of sample clusters, verification resources can be better allocated. By supplementing the comparison mode, inconsistencies between the classification results of dominant and latent features can be identified, thereby taking measures to improve the consistency and reliability of classification. Through detailed comparative analysis, misjudgments caused by single feature classification can be reduced, ensuring that the final traceability results are more accurate. Based on the setting of the difference evaluation value and the standard capacity difference, the system can flexibly adapt to different classification needs and sample characteristics, improving overall adaptability. Through precise comparison and verification mode selection, the accuracy of soybean origin traceability can be further improved.
[0066] By verifying and comparing the results of different verification steps, the accuracy of soybean origin traceability can be significantly improved, and the errors that may be introduced by a single classification can be reduced. Multiple verifications ensure the consistency and stability of the classification criteria, thereby enhancing the reliability of the overall classification system. By adjusting the classification criteria, the classification can be made more refined and the resolution of the classification can be improved. By comparing the results of the first and second verifications, misjudgments in the classification process can be effectively identified and reduced. By comparing the results of different verification steps, the consistency of the data can be ensured, which is crucial for establishing a reliable traceability system. By comprehensively considering the two characteristics, the samples can be evaluated more comprehensively, and the comprehensiveness of the classification can be improved.
[0067] By reanalyzing and validating the dominant and latent features of the samples, the accuracy of the classification results can be ensured and the classification error reduced. Validating the dominant and latent features of each sample cluster individually can verify the stability and reliability of the classification results, improve the accuracy of traceability, and assess the stability of the classification results by checking the stability of features within the sample clusters. This provides a basis for the final traceability results and ensures the accuracy and reliability of the traceability results. Attached Figure Description
[0068] Figure 1 This is a flowchart of the method for identifying and tracing the origin of agricultural products based on feature fingerprints, as described in an embodiment of the present invention.
[0069] Figure 2 This is a comparative flowchart of the method for identifying and tracing the origin of agricultural products based on feature fingerprints in this embodiment of the invention;
[0070] Figure 3 This is a flowchart of the successive verification process in the agricultural product origin identification and traceability method based on feature fingerprints according to an embodiment of the present invention;
[0071] Figure 4 This is a system block diagram of an agricultural product origin identification and traceability system based on feature fingerprints, according to an embodiment of the present invention. Detailed Implementation
[0072] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0073] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0074] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0075] See Figure 1 As shown, this invention provides a method for identifying and tracing the origin of agricultural products based on feature fingerprints;
[0076] In this embodiment, the target agricultural product is soybeans, but other agricultural products can also be selected. When identifying and tracing the origin of agricultural products based on their characteristic fingerprints, this embodiment does not specifically limit the type of agricultural product, and will not elaborate further here.
[0077] The method includes:
[0078] Step S100: Collect soybean samples with a preset initial sample size to obtain the target sample set;
[0079] Step S200: The target sample set is fingerprinted according to the dominant features of soybeans to obtain dominant marking results, and the target sample set is divided into several sample clusters according to the dominant marking results.
[0080] Step S300: The target sample set is fingerprinted with latent features based on the latent features of soybeans to obtain latent labeling results, and the target sample set is divided into two categories based on the latent labeling results to obtain several binary sample clusters.
[0081] Step S400: Compare the results of the first-class classification with the results of the second-class classification to determine whether to enable the successive verification mode or the individual verification mode.
[0082] Step S500: Based on the successive verification mode, obtain the verification results of each verification, and determine whether to output the actual place of origin traceability results based on the verification results of each verification.
[0083] Step S600: Based on the separate verification mode, the first-class classification result and the second-class classification result are verified separately to obtain the first-class verification result and the second-class verification result, and it is determined whether to output the actual place of origin traceability result based on the first-class verification result and the second-class verification result.
[0084] In this embodiment, the dominant features are the physical characteristics of soybeans, extracted using image processing techniques, such as shape, size, color, and texture. The latent features are the chemical characteristics of soybeans, extracted using chemical analysis techniques, such as protein content, fat content, fatty acid composition, amino acid composition, or mineral content. In specific implementation, image processing techniques can acquire soybean image data using an industrial camera and process the image data accordingly. The specific processing includes: segmenting the soybean image into different regions, such as the bean region and the background region; and extracting the physical characteristics of soybeans from the segmented image, for example, using geometric features (such as area, perimeter, and roundness) or shape descriptors (such as Hue, Zn, etc.). Image processing techniques, such as color moments, are used to describe the shape of soybeans. The size of soybeans is described by the number of pixels or the area of soybeans in the image. The color of soybeans is described by methods such as color histograms, color moments, or color space transformation. The texture of soybeans is described by methods such as gray-level co-occurrence matrix, local binary mode (LBP), or wavelet transform. Chemical analysis techniques can be used to analyze the chemical composition of soybeans using techniques such as spectral analysis, chromatographic analysis, or mass spectrometry. In this embodiment, image processing techniques and chemical analysis techniques are not specifically limited and will not be elaborated here.
[0085] Specifically, this invention combines dominant physical characteristics and latent chemical characteristics to more comprehensively describe the properties of soybeans, thereby improving the accuracy of origin identification. It uses dominant and latent characteristics to classify soybeans into primary and secondary categories, respectively, and verifies the results through successive or individual verification methods, enhancing the reliability of the classification results. This multi-level verification process reduces misjudgments caused by single characteristics or single analysis methods, improving the overall accuracy of traceability. Quantifying the physical and chemical characteristics of soybeans makes origin identification more objective, reducing uncertainty caused by subjective judgment. Accurate origin traceability results can help control agricultural product quality, ensuring product quality and safety, and improving the accuracy and efficiency of soybean origin identification.
[0086] Specifically, in this embodiment, the process of performing dominant feature fingerprinting on the target sample set based on the dominant features of soybeans to obtain dominant labeling results, and then dividing the target sample set into several sample clusters based on the dominant labeling results, includes:
[0087] Step S210: Determine each dominant feature vector based on the dominant features of soybean, and perform dominant feature fingerprinting on each target sample point based on each dominant feature vector to determine the dominant identifier of each target sample point.
[0088] Step S220: Based on the explicit identifier, classify each target sample point into one class according to the initial explicit classification criteria to obtain each sample cluster corresponding to different origin types;
[0089] The target sample set includes several target sample points; the sample cluster includes several target sample points with explicit identifiers.
[0090] In this embodiment, the dominant feature vectors include: shape vector, size vector, color vector, and texture vector. For any target sample point, the dominant features of the target sample point are first identified. The specific identification process includes: 1) For the shape vector: obtaining the appearance shape of the soybean, such as round, oval, or kidney-shaped. In the specific implementation, this is determined by calculating the geometric parameters of the soybean outline; 2) For the size vector: obtaining the size of the soybean, such as length, width, and thickness. In the specific implementation, this is quantified by measuring these physical dimensions to obtain soybeans of different sizes; 3) For the color vector: referring to the color attribute of the soybean. In the specific implementation, this can be represented by coordinates in the color space (such as RGB and HSV values); 4) For the texture vector: reflecting the texture and pattern of the soybean surface. In the specific implementation, texture analysis techniques (such as gray-level co-occurrence matrix) can be used to extract the texture features of the soybean to determine soybeans of different textures.
[0091] The feature data of the target sample point is collected, and the feature data is encoded to generate a comprehensive feature fingerprint, i.e., a dominant identifier. In this embodiment, the dominant identifier is a condensed representation of the features of the target sample point, which can uniquely identify the dominant features of the target sample point. The dominant identifier for any target sample point includes the shape, size, color, and texture information of the target sample point. For example, the dominant identifier corresponding to any target sample point is: [circle, size: 8.2mm × 6.5mm × 4.3mm, ...]. HSV: (60°, 80%, 90%), smooth texture, where size: 8.2mm × 6.5mm × 4.3mm refers to the length of the target sample point being 8.2mm, the width being 6.5mm, and the thickness being 4.3mm; HSV: (60°, 80%, 90%) refers to H (hue): representing the type of soybean color, S (saturation): representing the purity of the soybean color, with high saturation meaning a more vibrant color, and V (brightness): representing the brightness of the soybean color, i.e., the corresponding soybean is a mature yellow;
[0092] In this embodiment, the initial dominant classification criteria are set based on the dominant characteristics of soybeans. These criteria are used to group target sample points with similar characteristics into the same category. The specific setting process includes: a specific type of shape parameter, a specific range of size parameter, a similarity threshold for color features, and a classification standard for texture features. Sample points whose dominant identifiers meet the same classification criteria are grouped into one category, forming a sample cluster. For example, dominant identifiers with consistent shape vector types, size parameters corresponding to size vectors within a set fluctuation range, HSV value differences corresponding to color vectors within a set similarity threshold, and consistent texture features are grouped into one category to form a sample cluster. Since the dominant characteristics of soybeans are closely related to their growing environment (i.e., place of origin), environmental conditions in different places (such as soil type and climate conditions) will affect the shape, size, color, and texture of soybeans. Therefore, soybeans with similar dominant characteristics come from the same place of origin. Classifying them by dominant identifiers actually groups soybeans with similar environmental adaptability together, thus forming a sample cluster corresponding to different place types.
[0093] Specifically, this invention improves classification accuracy by comprehensively describing soybean samples. Dominant identifiers, acting as characteristic fingerprints, uniquely identify the dominant features of each target sample point, facilitating subsequent classification and traceability. Based on the initial dominant classification criteria using these identifiers, soybean samples with similar characteristics can be quickly grouped into one category, improving classification efficiency. Dominant features are closely related to the soybean's growth environment; classification reflects the soybean's adaptability to the environment and helps understand the characteristics of different production areas. Dominant identifiers allow for tracing the origin of soybean samples, enhancing product traceability. Effectively utilizing the dominant features of soybeans for classification and traceability improves the accuracy and efficiency of classification and traceability.
[0094] Specifically, in this embodiment, the process of performing latent feature fingerprinting on the target sample set based on the latent features of soybeans to obtain latent labeling results, and then performing binary classification on the target sample set based on the latent labeling results to obtain several binary sample clusters includes:
[0095] Step S310: Determine the latent feature vector based on the latent features of soybeans, and perform latent feature fingerprinting on each target sample point based on the latent feature vector to determine the latent identifier of each target sample point.
[0096] Step S220: Based on the implicit identifier, divide each target sample point into two categories according to the initial implicit classification criteria to obtain each binary sample cluster corresponding to different origin types;
[0097] The second type of sample cluster includes several target sample points with implicit identifiers.
[0098] In this embodiment, the latent feature vector of soybeans is selected as mineral content. The latent feature (mineral content) of soybeans is identified, and appropriate chemical analysis methods (such as atomic absorption spectrometry or inductively coupled plasma mass spectrometry) are used to determine the mineral content in soybean samples, including potassium, calcium, magnesium, iron, and zinc. The mineral content data for each target sample point is recorded, and a latent identifier is generated. For example, a latent identifier might be represented as [potassium: 100 mg / kg, calcium: 200 mg / kg, magnesium: 150 mg / kg, iron: 30 mg / kg, zinc: 20 mg / kg]. The latent identifier is a condensed representation of the mineral content of the target sample point, uniquely identifying its mineral composition. Based on the distribution and correlation of mineral content, an initial latent classification standard is set. The specific setting process includes: a similarity threshold for mineral content, and comparing the latent feature fingerprint with the initial latent classification standard. By comparing the criteria, sample points that meet the same classification standard are grouped into the same category, forming binary sample clusters. For example, sample points with potassium content in the range of 90-110 mg / kg, calcium content in the range of 190-210 mg / kg, magnesium content in the range of 140-160 mg / kg, iron content in the range of 20-40 mg / kg, and zinc content in the range of 10-30 mg / kg are grouped into the same binary sample cluster. This clusters soybean samples with similar mineral content characteristics together, forming several binary sample clusters. Since mineral content is closely related to the soybean's growing environment (origin), the soil mineral content and fertility of different origins will affect the soybean's mineral absorption. Therefore, classification by latent feature fingerprints actually groups soybeans with similar mineral compositions together. These soybeans may come from similar growing environments, thus forming binary sample clusters corresponding to different origin types.
[0099] Specifically, this invention analyzes the latent characteristics of soybean mineral content to gain a deeper understanding of the intrinsic quality of soybeans and the influence of the growing environment, thereby improving the depth and accuracy of classification. By combining dominant and latent characteristics, soybean samples can be classified from multiple dimensions, enhancing the comprehensiveness and reliability of classification. Latent identifiers, as a concentrated representation of mineral content, can uniquely identify the mineral composition of each target sample point, facilitating accurate classification and traceability. Mineral content is closely related to the soybean growing environment. Classification using latent characteristics can analyze the adaptability of soybeans to specific soil and climate conditions, providing guidance for agricultural production. Latent characteristic fingerprinting helps to more accurately track the origin of soybeans, especially when dominant characteristics are similar, latent characteristics can provide additional distinguishing information.
[0100] See Figure 2 As shown, this is a comparison flowchart, in which the process of comparing the results of the first-class classification and the second-class classification to determine whether to initiate the successive verification mode or the individual verification mode includes:
[0101] Step S410: Identify a class of partitions to determine the number of partitions and the size of each class of sample clusters;
[0102] Step S420: Identify binary partitions to determine the number of binary partitions and the capacity of each binary sample cluster;
[0103] Step S430: Compare the number of first-class divisions and the number of second-class divisions to determine whether to enable the supplementary comparison mode;
[0104] Step S440: Based on enabling the supplementary comparison mode, perform a second comparison according to the capacity of each type I sample cluster and the capacity of each type II sample cluster, and determine whether to enable the successive verification mode or the individual verification mode based on the results of the second comparison.
[0105] The first-class partitioning result includes: the number of first-class partitions and the capacity of each first-class sample cluster; the second-class partitioning result includes: the number of second-class partitions and the capacity of each second-class sample cluster.
[0106] Specifically, in this embodiment, the process of comparing the number of divisions in the first category and the number of divisions in the second category to determine whether to activate the supplementary comparison mode includes:
[0107] Step S421: Determine whether to enable the supplementary comparison mode based on the absolute value of the first difference and the preset difference evaluation value;
[0108] Step S422: Based on the fact that the absolute value of the first difference is greater than or equal to the difference evaluation value, it is determined that the successive verification mode is enabled.
[0109] Step S423: Based on the fact that the absolute value of the first difference is less than the difference evaluation value, it is determined that the supplementary comparison mode is activated;
[0110] Wherein, the absolute value of the first difference is the absolute value of the difference between the number of first-class divisions and the number of second-class divisions.
[0111] In the specific implementation process, the difference evaluation value is used to measure the degree of difference between the two classification results. Specifically, it measures the inconsistency or difference between the first-class classification (based on dominant features) and the second-class classification (based on latent features). If the difference in the number of the two classes is large, it indicates that there is a significant difference between the origin classification results of soybean samples based on different features, and further verification analysis is required. The setting of the difference evaluation value is affected by the actual classification accuracy and the initial sample size. In this embodiment, the difference evaluation value is set to 10. For example, if the first-class classification has 50 sample clusters and the second-class classification has 45 sample clusters, the first absolute difference value is 5. If the preset difference evaluation value is 10, then since the absolute difference value is less than the difference evaluation value, it will be determined that the supplementary comparison mode will be activated.
[0112] Specifically, in this embodiment, the process of performing a secondary comparison based on the size of each type I sample cluster and the size of each type II sample cluster includes:
[0113] Step S4411: Obtain the largest and smallest sample size among the sample cluster sizes of each class.
[0114] Step S4412: Obtain the maximum and minimum binary sample sizes among the sizes of each binary sample cluster.
[0115] Step S4413: Compare the maximum sample size of the first class with the maximum sample size of the second class to obtain the maximum comparison result;
[0116] Step S4414: Compare the minimum first-class sample size with the minimum second-class sample size to obtain the minimum comparison result;
[0117] Step S4415: Determine the number of items that meet a single judgment condition in the maximum comparison result and the minimum comparison result.
[0118] Specifically, in this embodiment, the process of determining the number of items that meet a single judgment condition in the maximum comparison result and the minimum comparison result includes:
[0119] Calculate the maximum capacity difference between the maximum first-class sample size and the maximum second-class sample size, and determine whether the single judgment condition is met based on the maximum capacity difference;
[0120] Calculate the minimum size difference between the minimum size of the first class of samples and the minimum size of the second class of samples, and determine whether the single judgment condition is met based on the minimum size difference.
[0121] Specifically, if the maximum capacity difference is less than or equal to a preset standard capacity difference, the maximum comparison result is determined to meet the single determination condition; if the maximum capacity difference is greater than the standard capacity difference, the maximum comparison result is determined to not meet the single determination condition. Similarly, if the minimum capacity difference is less than or equal to the standard capacity difference, the minimum comparison result is determined to meet the single determination condition; if the minimum capacity difference is greater than the standard capacity difference, the minimum comparison result is determined to not meet the single determination condition.
[0122] In the specific implementation process, the standard capacity difference is used to measure whether the difference between the maximum or minimum sample capacity in the first-class sample cluster and the second-class sample cluster is within an acceptable range. The specific setting is affected by the actual classification accuracy and the initial sample capacity. In this embodiment, the standard capacity difference is set to 5.
[0123] Specifically, in this embodiment, the process of determining whether to activate the successive verification mode or the individual verification mode based on the results of the second comparison includes:
[0124] Step S4421: When the number of items is 2, it is determined that the individual verification mode is enabled.
[0125] Step S4422: If the number of items is less than 2, then it is determined that the successive verification mode is enabled.
[0126] Specifically, by comparing the classification results of Class I and Class II, the embodiments of the present invention can effectively determine whether to adopt a successive verification mode or a single verification mode, thereby improving the efficiency of the entire verification process. By comparing the number and capacity of sample clusters, verification resources can be better allocated. By supplementing the comparison mode, inconsistencies between the classification results of dominant and latent features can be identified, thereby taking measures to improve the consistency and reliability of classification. Through detailed comparative analysis, misjudgments caused by single feature classification can be reduced, ensuring that the final traceability results are more accurate. Based on the setting of the difference evaluation value and the standard capacity difference value, the system can flexibly adapt to different classification needs and sample characteristics, improving the overall adaptability. Through precise comparison and verification mode selection, the accuracy of soybean origin traceability can be further improved.
[0127] See Figure 3 As shown, this is a flowchart of the successive verification process. The process of determining whether to output the actual origin traceability result based on the results of each verification includes:
[0128] Step S510: Based on the successive verification mode, adjust the initial sample size according to the single verification standard to obtain the verification sample size;
[0129] Step S520: Based on the dominant and recessive characteristics of soybeans, the verification sample capacity is used to perform a primary verification division to obtain several primary verification first-class sample clusters and several primary verification second-class sample clusters.
[0130] Step S530: Determine the number of explicit partitions and the number of implicit partitions in each validation based on the first-class sample clusters and the second-class sample clusters in each validation.
[0131] Step S540: Determine the verification result based on the number of explicit partitions and the number of implicit partitions in the verification.
[0132] Step S550: Based on the successive verification mode, adjust the target classification standard according to the secondary verification standard to obtain the actual classification standard;
[0133] Step S560: Perform secondary verification on the initial sample capacity to obtain several secondary verification first-class sample clusters and several secondary verification second-class sample clusters.
[0134] Step S570: Determine the number of explicit partitions and the number of implicit partitions in the secondary validation based on each secondary validation class 1 sample cluster and each secondary validation class 2 sample cluster.
[0135] Step S580: Determine the secondary verification result based on the number of explicit partitions and the number of implicit partitions in the secondary verification.
[0136] Step S590: Determine whether the first verification result and the second verification result are consistent. If the verification results are consistent, output the actual place of origin traceability result.
[0137] In this implementation, the first verification standard is to increase the initial sample size to 1.5-2 times;
[0138] The process of determining the verification result based on the number of explicit partitions and the number of implicit partitions in a single verification includes:
[0139] Step S541: Determine whether the first verification result is consistent with the initial result or inconsistent with the initial result based on the absolute value of the second difference and the difference evaluation value.
[0140] Step S542: Based on the fact that the absolute value of the second difference is greater than or equal to the difference evaluation value, the first verification result is determined to be consistent with the initial result;
[0141] Step S423: Based on the fact that the absolute value of the second difference is less than the difference evaluation value, it is determined that the first verification result is inconsistent with the initial result;
[0142] Wherein, the absolute value of the second difference is the absolute value of the difference between the number of explicit partitions in the first verification and the number of implicit partitions in the first verification.
[0143] Specifically, in this embodiment, the process of adjusting the target classification standard according to the secondary verification standard based on the successive verification mode to obtain the actual classification standard includes:
[0144] When the absolute value of the first difference is greater than or equal to the difference evaluation value, the classification standard corresponding to the larger value between the number of first-class divisions and the number of second-class divisions is selected as the target classification standard. For example, if the number of first-class divisions is greater than the number of second-class divisions, an initial implicit classification standard is selected as the target classification standard; if the number of first-class divisions is less than the number of second-class divisions, an initial explicit classification standard is selected as the target classification standard.
[0145] In this embodiment, when the initial implicit classification standard is selected as the target classification standard, the secondary verification standard is to reduce the similarity threshold of mineral content, so that the implicit classification standard is more refined; when the initial explicit classification standard is selected as the target classification standard, the secondary verification standard is to reduce the set fluctuation range and the set similarity threshold, so that the explicit classification standard is more refined.
[0146] The process of determining the secondary verification result based on the number of explicit partitions and the number of implicit partitions in the secondary verification includes:
[0147] Step S581: Based on the absolute value of the third difference and the difference evaluation value, determine whether the secondary verification result is consistent with the initial result or inconsistent with the initial result.
[0148] Step S582: Based on the fact that the absolute value of the third difference is greater than or equal to the difference evaluation value, the result of the second verification is determined to be consistent with the initial result.
[0149] Step S583: Based on the fact that the absolute value of the third difference is less than the difference evaluation value, it is determined that the secondary verification result is inconsistent with the initial result;
[0150] The absolute value of the third difference is the absolute value of the difference between the number of explicit partitions in the secondary verification and the number of implicit partitions in the secondary verification.
[0151] Specifically, the process of determining whether the first verification result and the second verification result are consistent, and outputting the actual origin traceability result when the verification results are consistent, includes: if both the first verification result and the second verification result are consistent with or inconsistent with the initial result, then the division is determined to be accurate, and the actual origin traceability result is output; if one of the first verification result and the second verification result is consistent with the initial result and the other is inconsistent with the initial result, then the division is determined to be inaccurate, and the actual origin traceability result is not output.
[0152] In this embodiment, successive verification reduces the error that may be caused by a single division by dividing and comparing multiple times, and improves the accuracy of classification. By comparing the results of different verification steps, the consistency and stability of the classification criteria can be checked. If the results of multiple verifications are consistent, it indicates that the division result is accurate. If the results of two verifications are consistent, the actual origin traceability result is output. The specific output process is as follows: the initial division result is compared with the database corresponding to the standard soybean origin and characteristics. The feature identifier of each sample cluster is matched with the record in the database to output the actual origin traceability result.
[0153] Specifically, the embodiments of the present invention can significantly improve the accuracy of soybean origin traceability by verifying and comparing the results of different verification steps, reducing the errors that may be introduced by a single classification. Multiple verifications ensure the consistency and stability of the classification standards, thereby enhancing the reliability of the overall classification system. By adjusting the classification standards, the classification can be made more detailed, improving the resolution of the classification. By comparing the results of the first and second verifications, misjudgments in the classification process can be effectively identified and reduced. By comparing the results of different verification steps, the consistency of the data can be ensured, which is crucial for establishing a reliable traceability system. By comprehensively considering both characteristics, the samples can be evaluated more comprehensively, improving the comprehensiveness of the classification.
[0154] Specifically, in this embodiment, the process of separately verifying the first-class classification result and the second-class classification result based on the separate verification mode to obtain the first-class verification result and the second-class verification result, and determining whether to output the actual place of origin traceability result based on the first-class verification result and the second-class verification result, includes:
[0155] Samples from one type of cluster are re-analyzed to confirm whether their dominant characteristics are consistent with the initial label. The shape, size, color, and texture of soybeans are verified through visual inspection, instrument measurement, or other verification methods. The stability of dominant characteristics within each cluster is checked to ensure that no samples are misclassified. Samples from the second type of cluster are chemically analyzed to re-determine their latent mineral content. The re-measured data are compared with the initial label results to check for significant differences and to assess the stability of the clusters based on latent characteristics.
[0156] The results of verification for each class of sample clusters based on dominant features show no significant differences, indicating that the class of sample clusters accurately reflects the dominant features of soybeans. The results of verification for each class of sample clusters based on latent features show no significant differences, indicating that the class of sample clusters accurately reflects the latent features of soybeans. If both the first-class and second-class verification results show no significant differences, the sample cluster division is considered accurate and stable, and the division result is considered reliable. The verified sample features are then matched with a standard soybean origin and feature database. Based on the matching results, the actual origin traceability information for each sample is output, i.e., the actual origin traceability result is output. If either the first-class or second-class verification result shows a significant difference, the division result is considered unreliable.
[0157] Specifically, by re-analyzing and verifying the dominant and latent features of the samples, the embodiments of the present invention can ensure the accuracy of the classification results and reduce classification errors. By individually verifying the dominant and latent features of each sample cluster, the stability and reliability of the classification results can be verified, improving the accuracy of tracing. By checking the stability of features within the sample cluster, the stability of the classification results can be evaluated, providing a basis for the output of the final tracing results and ensuring the accuracy and reliability of the tracing results.
[0158] See Figure 4 The present invention also provides an agricultural product origin identification and traceability system based on feature fingerprints, comprising:
[0159] The sample collection module is used to collect soybean samples according to a preset initial sample size in order to obtain the target sample set;
[0160] The dominant feature analysis module is used to perform dominant feature analysis on the target sample set to obtain dominant feature fingerprints;
[0161] The latent feature analysis module is used to perform latent feature analysis on the target sample set to obtain latent feature fingerprints;
[0162] The partitioning and clustering module is used to partition samples into primary and secondary classes based on each feature fingerprint, so as to obtain several primary sample clusters and several secondary sample clusters.
[0163] The verification mode selection module is used to compare the results of the first-class partitioning and the second-class partitioning to determine whether to use the successive verification mode or the individual verification mode.
[0164] The successive verification module is used to perform multiple divisions through multiple verifications to obtain the verification results of each verification, and decide whether to output the actual place of origin traceability results based on the verification results of each verification.
[0165] The separate verification module is used to independently verify the first-class and second-class classification results, and combine the two verification results to determine whether to output the actual place of origin traceability results.
[0166] The calculation compensation parameters and calculation adjustment parameters described in this invention serve two purposes: first, to balance the left and right dimensions of the formula; and second, to adjust the numerical results. In this embodiment, no specific values are assigned. Furthermore, in this embodiment, each calculation formula is used to intuitively reflect the adjustment relationship between the values, such as positive correlation or negative correlation. Unless otherwise specified, the values of parameters that are not specifically limited are all taken as positive.
[0167] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0168] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying the origin of agricultural products based on characteristic fingerprints, comprising: collecting target agricultural product samples with a preset initial sample capacity to obtain a target sample set; labeling the target sample set with dominant characteristic fingerprints according to the dominant characteristics of the target agricultural products to obtain a dominant labeling result, and performing one-class partitioning on the target sample set according to the dominant labeling result to obtain a plurality of one-class sample clusters; labeling the target sample set with recessive characteristic fingerprints according to the recessive characteristics of the target agricultural products to obtain a recessive labeling result, and performing two-class partitioning on the target sample set according to the recessive labeling result to obtain a plurality of two-class sample clusters; comparing the one-class partitioning result and the two-class partitioning result to determine whether to start a successive verification mode or a separate verification mode; based on the successive verification mode, obtaining each verification result, and determining whether to output an actual origin traceability result according to each verification result; based on the separate verification mode, separately verifying the one-class partitioning result and the two-class partitioning result to obtain a one-class verification result and a two-class verification result, and determining whether to output the actual origin traceability result according to the one-class verification result and the two-class verification result; the process of comparing the one-class partitioning result and the two-class partitioning result to determine whether to start the successive verification mode or the separate verification mode comprises: identifying one-class partitioning to determine the number of one-class partitioning and the capacity of each one-class sample cluster; identifying two-class partitioning to determine the number of two-class partitioning and the capacity of each two-class sample cluster; based on the first difference absolute value being greater than or equal to the difference evaluation value, it is determined that the successive verification mode is started; based on the first difference absolute value being less than the difference evaluation value, it is determined that a supplementary comparison mode is started, and based on the supplementary comparison mode being started, a second comparison is performed according to the capacity of each one-class sample cluster and the capacity of each two-class sample cluster, comprising: obtaining the maximum one-class sample capacity and the minimum one-class sample capacity in the capacity of each one-class sample cluster; obtaining the maximum two-class sample capacity and the minimum two-class sample capacity in the capacity of each two-class sample cluster; comparing the maximum one-class sample capacity with the maximum two-class sample capacity to obtain a maximum comparison result; comparing the minimum one-class sample capacity with the minimum two-class sample capacity to obtain a minimum comparison result; determining the number of items that meet a single determination condition in the maximum comparison result and the minimum comparison result; based on the number of items being 2, it is determined that the separate verification mode is started; based on the number of items being less than 2, it is determined that the successive verification mode is started; wherein the first difference absolute value is the absolute value of the difference between the number of one-class partitioning and the number of two-class partitioning, the dominant characteristics are the physical characteristics of soybeans, and the physical characteristics of soybeans are extracted using image processing technology; the recessive characteristics are the chemical characteristics of soybeans, and the chemical characteristics of soybeans are extracted using chemical analysis technology.
2. The feature fingerprint based agricultural produce geographical origin identification traceability method according to claim 1, characterized in that, the process of labeling the target sample set with dominant characteristic fingerprints according to the dominant characteristics of the target agricultural products to obtain a dominant labeling result, and performing one-class partitioning on the target sample set according to the dominant labeling result to obtain a plurality of one-class sample clusters comprises: According to the explicit feature of the target agricultural product, an explicit feature vector is determined, and each target sample point is marked with an explicit feature fingerprint according to the explicit feature vector, so as to determine an explicit identifier of each target sample point; Based on the explicit identifier, each target sample point is classified into one category according to an initial explicit classification standard, so as to obtain each one-category sample cluster corresponding to different types of production places; The target sample set includes a plurality of target sample points, and the one-category sample cluster includes a plurality of target sample points with explicit identifiers.
3. The feature fingerprint based agricultural produce geographical origin identification traceability method as claimed in claim 2, wherein, The process of marking the target sample set with an implicit feature fingerprint according to the implicit feature of the target agricultural product to obtain an implicit marking result, and classifying the target sample set into two categories according to the implicit marking result to obtain a plurality of two-category sample clusters includes: According to the implicit feature of the target agricultural product, an implicit feature vector is determined, and each target sample point is marked with an implicit feature fingerprint according to the implicit feature vector, so as to determine an implicit identifier of each target sample point; Based on the implicit identifier, each target sample point is classified into two categories according to an initial implicit classification standard, so as to obtain each two-category sample cluster corresponding to different types of production places; The two-category sample cluster includes a plurality of target sample points with implicit identifiers.
4. The feature fingerprint based agricultural produce geographical origin identification traceability method as claimed in claim 3, wherein, The process of determining whether to output the actual origin traceability result according to each verification result includes: Based on the successive verification mode, the initial sample capacity is adjusted according to a one-time verification standard to obtain a verification sample capacity; The verification sample capacity is used to perform one-time verification division according to the explicit feature and the implicit feature of the target agricultural product, to obtain a plurality of one-time verification one-category sample clusters and a plurality of one-time verification two-category sample clusters; The number of one-time verification explicit divisions and the number of one-time verification implicit divisions are determined according to each one-time verification one-category sample cluster and each one-time verification two-category sample cluster; The one-time verification result is determined according to the number of one-time verification explicit divisions and the number of one-time verification implicit divisions.
5. The feature fingerprint based agricultural produce geographical origin identification traceability method as claimed in claim 4, wherein, The process of determining whether to output the actual origin traceability result according to each verification result also includes: Based on the successive verification mode, the target classification standard is adjusted according to a two-time verification standard to obtain an actual classification standard; The initial sample capacity is subjected to two-time verification division to obtain a plurality of two-time verification one-category sample clusters and a plurality of two-time verification two-category sample clusters; The number of two-time verification explicit divisions and the number of two-time verification implicit divisions are determined according to each two-time verification one-category sample cluster and each two-time verification two-category sample cluster; The two-time verification result is determined according to the number of two-time verification explicit divisions and the number of two-time verification implicit divisions; It is determined whether the one-time verification result and the two-time verification result are consistent, and when the verification results are consistent, the actual origin traceability result is output.
6. The feature fingerprint based agricultural produce geographical origin identification traceability method as claimed in claim 5, wherein, The process of determining the one-time verification result according to the number of one-time verification explicit divisions and the number of one-time verification implicit divisions includes: According to the second difference absolute value combined with a preset difference evaluation value, the one-time verification result is determined to be consistent with the initial or inconsistent with the initial; Based on the second difference absolute value being greater than or equal to the difference evaluation value, it is determined that the one-time verification result is consistent with the initial. If the second difference absolute value is less than the difference evaluation value, it is determined that the first verification result is inconsistent with the initial result. The second difference absolute value is an absolute value of a difference between the first verification explicit division number and the first verification implicit division number.
7. The feature fingerprint based agricultural produce geographical origin identification traceability method as claimed in claim 6, wherein, The process of determining the second verification result according to the second verification explicit division number and the second verification implicit division number comprises: According to the third difference absolute value and the difference evaluation value, it is determined that the second verification result is consistent with the initial result or inconsistent with the initial result. If the third difference absolute value is greater than or equal to the difference evaluation value, it is determined that the second verification result is consistent with the initial result. If the third difference absolute value is less than the difference evaluation value, it is determined that the second verification result is inconsistent with the initial result. The third difference absolute value is an absolute value of a difference between the second verification explicit division number and the second verification implicit division number.
8. A characteristic fingerprint-based agricultural product origin identification traceability system, using the characteristic fingerprint-based agricultural product origin identification traceability method according to any one of claims 1 to 7, characterized in that, The method comprises: a sample collection module configured to collect target agricultural product samples according to a preset initial sample capacity to obtain a target sample set; an explicit feature analysis module configured to perform explicit feature analysis on the target sample set to obtain an explicit feature fingerprint; an implicit feature analysis module configured to perform implicit feature analysis on the target sample set to obtain an implicit feature fingerprint; a division and clustering module configured to perform one-class division and two-class division on the samples according to the feature fingerprints to obtain a plurality of one-class sample clusters and a plurality of two-class sample clusters; a verification mode selection module configured to compare the one-class division result and the two-class division result to determine whether to adopt a successive verification mode or a separate verification mode; a successive verification module configured to perform multiple divisions through multiple verifications to obtain a plurality of verification results and determine whether to output an actual origin traceability result according to the verification results; a separate verification module configured to independently verify the one-class division result and the two-class division result and determine whether to output an actual origin traceability result according to the two verification results.
Citation Information
Patent Citations
Soybean origin traceability identification method based on combination of MALDI-TOF / TOF and IRMS technologies
CN113406245A
Agricultural product image processing and blockchain interaction identification method and system
CN111159458A
Dairy product tracing method based on block chain
CN118822560A