Methods for constructing, compressing, and matching a garlic origin traceability feature database

CN122492246BActive Publication Date: 2026-09-01SHANDONG XINNUO FOOD TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610922369.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-01
Estimated Expiration
2046-06-25

AI Technical Summary

Technical Problem

当特征库直接保存所有字段及其完整精度时,虽然有利于提高匹配依据的完备性,但会造成字段数量多、取值范围细、产地记录冗余、匹配计算量大等问题;当为降低存储量而简单删除字段、合并字段或者降低字段精度时,又可能改变历史样品的候选产地集合、被排除产地或者最终产地,导致压缩后的特征库出现误召回、漏排除或最终产地偏移,影响溯源结论的稳定性

Benefits of technology

[0019] This invention first acquires traceability data of garlic samples from known origins and breaks it down into multiple feature fields according to feature categories. Field identifiers, value ranges, and origin identifiers are written into the corresponding origin records to form an uncompressed feature library. This uncompressed feature library is then used to match historical samples and record the candidate origin set, excluded origins, and final origin. Thus, before compression, a replayable and comparable baseline matching result can be generated, ensuring that subsequent feature library compression is constrained by the principle that the origin identification result does not substantially change.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492246B_ABST
    Figure CN122492246B_ABST
Patent Text Reader

Abstract

This invention relates to a method for constructing, compressing, and matching a garlic origin traceability feature library, comprising: acquiring traceability data of garlic samples from known origins, splitting it into multiple feature fields according to feature categories, writing field identifiers, field value ranges, and origin identifiers into corresponding origin records to form an uncompressed feature library, and using the uncompressed feature library to match historical samples, recording candidate origin sets, excluded origins, and final origins; performing field deletion, field merging, or field precision reduction processing on each feature field, statistically analyzing the differences in matching results before and after processing and classifying compression loss levels, generating a compressed feature library, and activating it after verifying consistency through historical sample playback matching; acquiring traceability data of garlic samples to be traced, generating matching data and matching it with the activated compressed feature library, and outputting the origin traceability results, thereby reducing the amount of feature library data and the matching computation burden while improving the stability of origin identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural product origin traceability and data processing technology, specifically to a method for constructing, compressing, and matching a garlic origin traceability feature database. Background Technology

[0002] In existing technologies, the traceability of origin for specialty agricultural products such as garlic typically relies on origin information, testing information, coding identifiers, database records, or multi-dimensional sample characteristics for management. Such solutions can achieve source query and anti-counterfeiting identification to a certain extent. For example, the method, device, electronic equipment, and storage medium for agricultural product anti-counterfeiting traceability disclosed in Publication No. CN120338803A generate a first agricultural product code based on each origin information, encrypts the code, generates a color QR code, and then matches the second agricultural product code obtained by scanning it through a user terminal with an agricultural product code database to return the corresponding origin information. Another example is the agricultural product origin traceability method and system based on big data disclosed in Publication No. CN116883026A, which generates agricultural product origin traceability results by constructing a basic information database, extracting specific characteristics of agricultural products, setting limit intervals for sample characteristics in the same region, and establishing an origin traceability network. The aforementioned public documents indicate that existing traceability technologies have gradually shifted from simple circulation record queries to a comprehensive judgment based on origin codes, sample characteristics, and database matching.

[0003] However, tracing the origin of garlic differs from general agricultural product coding for anti-counterfeiting or distribution traceability. Identifying the origin of garlic samples often requires simultaneous reference to multiple categories of data, including physicochemical indicators, mineral elements, stable isotopes, climate and soil correlation parameters, appearance quality parameters, and batch testing parameters. Furthermore, there may be some overlap in characteristic ranges between different origins. While directly storing all fields and their complete precision in the feature library improves the completeness of matching criteria, it also leads to problems such as a large number of fields, narrow value ranges, redundant origin records, and high computational load. Conversely, simply deleting, merging, or reducing the precision of fields to reduce storage may alter the candidate origin set, excluded origins, or final origin of historical samples, resulting in false recalls, missed exclusions, or final origin shifts in the compressed feature library, affecting the stability of the traceability conclusions.

[0004] Meanwhile, existing publicly available technologies focus more on origin coding generation, barcode scanning, basic database establishment, or multi-dimensional feature recognition, lacking a constraint mechanism for the long-term compression of feature libraries to ensure consistency of matching results. Especially with the increasing number of garlic producing areas, the continuous accumulation of historical samples, and the expanding range of detection fields, if it is not possible to identify which fields anchor the final origin and which fields are crucial for excluding similar origins, it becomes difficult to maintain the original discrimination logic while reducing storage and matching burdens. Therefore, it is necessary to provide a technical solution that can evaluate field compression loss based on historical sample playback results, and verify the consistency of candidate origins, excluded origins, and the final origin before the compressed feature library is activated. Summary of the Invention

[0005] The purpose of this invention is to provide a method for constructing, compressing, and matching a garlic origin traceability feature database, thereby addressing some of the drawbacks and shortcomings pointed out in the background art.

[0006] The present invention adopts the following technical solution to solve the above-mentioned technical problems:

[0007] The traceability data of garlic samples from known production areas is obtained and split into multiple feature fields according to feature categories. The field identifier, field value range and production area identifier of each feature field are written into the corresponding production area record to form an uncompressed feature library. The uncompressed feature library is then used to match historical samples and record the candidate production area set, excluded production area and final production area corresponding to each historical sample.

[0008] Each feature field is processed by deleting, merging, or reducing precision. The differences between the candidate origin set, excluded origin, and final origin before and after processing are statistically analyzed. Based on the differences, the feature fields are divided into compression loss levels. The compression loss level is used to determine the field retention, merging, or precision reduction method. A compressed feature library is generated and verified by replaying historical samples. If the candidate origin set, excluded origin, and final origin obtained from the replay match the results of the uncompressed feature library, the compressed feature library is enabled. If they do not match, the compression loss level of the corresponding feature field is adjusted and the replay is performed again.

[0009] Obtain the traceability data of the garlic sample to be traced, generate matching data according to the feature category, match the matching data with the enabled compressed feature library, and output the origin traceability result.

[0010] Furthermore, after recording the candidate origin set, excluded origin, and final origin corresponding to each historical sample, an exclusion constraint is established. The exclusion constraint is associated with the excluded origin, the feature field that triggers the exclusion, and the value range of the field. When the playback matching result of the compressed feature library contains an origin that has been excluded in the matching result of the uncompressed feature library, the corresponding exclusion constraint is written into the anti-recall field of the origin record. The anti-recall field is used to exclude origin records that are mistakenly recalled after compression. When matching garlic samples to be traced, the anti-recall field is called to remove the corresponding origin record before determining the final origin.

[0011] Furthermore, when classifying feature fields into compression loss levels, feature fields that enable historical samples to enter the final origin are set as result anchor fields, and feature fields that enable candidate origins to become excluded origins are set as exclusion anchor fields. The result anchor fields and the exclusion anchor fields are configured to retain compression loss levels with higher priority than other feature fields, and the corresponding field identifiers, field value ranges, and corresponding origin records are retained in the compression feature library.

[0012] Furthermore, after obtaining the traceability data of the garlic sample to be traced, the feature fields involved in the matching in the garlic sample to be traced are compared with the feature fields involved in the matching in the compressed feature library. The completeness comparison includes the identification of missing values, null values, or invalid values. When the garlic sample to be traced lacks a feature field configured with a compression loss level, the missing field is not numerically filled in. Instead, a group of alternative feature fields that has been pre-recorded in the historical sample playback matching and maintains the same final origin when the same feature field is missing is called to participate in the matching.

[0013] Furthermore, when the same excluded origin corresponds to multiple exclusion constraints, the multiple exclusion constraints are merged according to the feature field that triggers the exclusion and the value range of the field; when the value ranges of multiple fields are continuous or overlap and correspond to the same excluded origin, the value ranges of multiple fields are merged into a single anti-recall field to reduce duplicate exclusion data in the compressed feature library.

[0014] Furthermore, before the exclusion constraint is written into the anti-recall field, the number of times the exclusion constraint is triggered in the historical sample playback matching is counted; when the number reaches a preset threshold, the exclusion constraint is written into the anti-recall field; when the number does not reach the preset threshold, the exclusion constraint is retained as an exclusion constraint to be verified. The exclusion constraint to be verified is stored in the compressed feature library but does not participate in the origin exclusion in the matching of garlic samples to be traced.

[0015] Furthermore, after the anti-recall field is invoked, the removed origin records and the feature fields that triggered the removal are recorded; when the same origin record is removed by the same feature field in the matching of garlic samples to be traced within a preset sample quantity threshold, the field value range corresponding to the feature field is updated to the priority exclusion interval of the origin record, and the priority exclusion interval is invoked before the ordinary exclusion constraint item in the matching stage.

[0016] Furthermore, before calling the alternative feature field group to participate in matching, the number of times the alternative feature field group was substituted and the number of times the place of origin was consistent in the historical sample playback matching are read, and a substitution credibility identifier is generated based on the ratio of the number of times the place of origin is consistent to the number of times the substitution was substituted; when the ratio corresponding to the substitution credibility identifier reaches a preset credibility threshold, the corresponding alternative feature field group is called to participate in matching.

[0017] Furthermore, when multiple alternative feature field groups exist, the feature category to which the missing field belongs is read, and alternative feature field groups that belong to different feature categories from the missing field and maintain the same final origin in historical sample playback matching are preferentially called. The alternative feature field groups that have not been called are written into the supplementary matching record. The supplementary matching record is used to verify the matching results when supplementing sample data later.

[0018] Furthermore, when the same group of alternative feature fields corresponds to more than two final origins in the historical sample playback matching, the number of substitutions and the number of times the origin matches are counted separately according to the final origin, and the alternative trust identifier is split according to the final origin to generate a sub-origin alternative trust identifier associated with the origin record; when matching garlic samples to be traced, the corresponding sub-origin alternative trust identifier is first screened according to the candidate origin, and then the alternative feature field group associated with the sub-origin alternative trust identifier that meets the preset trust threshold is called.

[0019] This invention first acquires traceability data of garlic samples from known origins and breaks it down into multiple feature fields according to feature categories. Field identifiers, value ranges, and origin identifiers are written into the corresponding origin records to form an uncompressed feature library. This uncompressed feature library is then used to match historical samples and record the candidate origin set, excluded origins, and final origin. Thus, before compression, a replayable and comparable baseline matching result can be generated, ensuring that subsequent feature library compression is constrained by the principle that the origin identification result does not substantially change.

[0020] This invention performs field deletion, field merging, or field precision reduction on each feature field, and statistically analyzes the differences between the candidate origin set, excluded origin, and final origin before and after processing. Based on this, it classifies the compression loss level to determine the method of field retention, field merging, or field precision reduction. This setting allows for the retention of key fields that significantly influence the final origin determination, while compressing less influential or redundant fields. This ensures the reliability of garlic origin traceability results while reducing the amount of feature database data and the burden of matching calculations.

[0021] This invention employs historical sample replay matching verification after generating the compressed feature library. This verification is only activated if the candidate origin set, excluded origins, and final origin obtained from the compressed feature library match the matching results of the uncompressed feature library. If they do not match, the compression loss level of the corresponding feature fields is adjusted, and the replay matching is performed again. This reduces the risks of false recall, false exclusion, and final origin shift, improving the stability of the compressed feature library when used for matching garlic samples for traceability. Attached Figure Description

[0022] Figure 1 The flowchart for constructing, compressing, and matching the garlic origin traceability feature library of this invention is shown below.

[0023] Figure 2 This is a diagram showing the classification of feature field compression loss levels in Embodiment 1 of the present invention.

[0024] Figure 3 This is a verification diagram of the feature library compression rate and playback consistency in Embodiment 1 of the present invention.

[0025] Figure 4 This is a call diagram of the anti-recall field and the priority exclusion interval in Embodiment 1 of the present invention.

[0026] Figure 5 This is a comparison chart of the integrity of the fields of the sample to be traced in Embodiment 2 of the present invention.

[0027] Figure 6 This is a screening diagram of the confidence value of the alternative feature field group in Embodiment 2 of the present invention.

[0028] Figure 7 This is a diagram showing the trusted identifier and verification path for place-of-origin substitution in Embodiment 2 of the present invention. Detailed Implementation

[0029] As attached Figure 1As shown, this embodiment provides a method for constructing, compressing, and matching a garlic origin traceability feature database. This method can be executed by a server, terminal device, or computer system with data processing capabilities. First, the computer system acquires traceability data for garlic samples from known origins. The traceability data may include one or more of the following: origin environment data, physicochemical testing data, elemental testing data, spectral data, variety data, harvest time data, and storage and transportation data. Origin environment data may include soil type, soil pH, water source type, climate zone, and altitude range data for the corresponding planting area of ​​the sample; physicochemical testing data may include data on moisture, reducing sugar, soluble solids, volatile flavor substances, or sulfur-containing compounds; elemental testing data may include data on potassium, calcium, magnesium, iron, zinc, selenium, and stable isotopes; spectral data may include characteristic peaks, spectral intensity, or spectral morphology data formed after preprocessing from near-infrared spectroscopy, Raman spectroscopy, or hyperspectral detection data. After acquiring the aforementioned traceability data, the computer system performs sample number association, unit unification, batch identification writing, and anomaly data verification on data from different sources. The verified data is used as the data source for subsequent feature field generation. The computer system splits the traceability data according to pre-defined feature categories, resulting in multiple feature fields. Each feature field corresponds to a field identifier, a field value range, and an origin identifier. The field value range can be formed by the verified value boundaries, category value sets, or spectral feature intervals of known samples from the same origin, and is stored together with the corresponding detection method identifier and data source identifier so that comparisons can be made based on the same field meaning during subsequent matching. Subsequently, the computer system writes the field identifier, field value range, and origin identifier into the corresponding origin record, forming an uncompressed feature library from multiple origin records. The uncompressed feature library retains the original matching features of garlic samples from known origins, serving as benchmark data for subsequent compressed verification.

[0030] After forming an uncompressed feature library, the computer system uses this library to match historical samples. During the matching process, the computer system compares the value ranges of each feature field of the historical sample with the field values ​​in each origin record to obtain a set of candidate origins for the historical sample. When a field value of a historical sample falls within the value range, category set, or spectral feature interval of the corresponding origin record, the computer system marks that field as meeting the matching condition of the corresponding origin record; when a field value exceeds the corresponding range, is inconsistent with the category set, or is inconsistent with the spectral feature interval, the computer system marks that field as not meeting the matching condition. Simultaneously, the computer system records origins excluded due to field values ​​not meeting the matching conditions and determines the final origin based on the matching results in the candidate origin set. When determining the final origin, the computer system can prioritize candidate origins that meet a large number of criteria, satisfy the result anchoring fields, and have not triggered exclusion anchoring fields. If multiple candidate origins meet the same criteria, the system further reads and compares the harvest time, stored transportation data, and historical stable matching records sequentially. If a unique determination is still not possible, multiple candidate origins and their corresponding verification identifiers are output. Thus, for each historical sample, the system generates three types of matching records: a candidate origin set, excluded origins, and the final origin. These matching records are used to evaluate the contribution of each feature field to the origin identification result.

[0031] After completing the historical sample matching records, the computer system performs field deletion, field merging, or field precision reduction on each feature field. Field deletion removes the corresponding feature field from the matching process. Field merging combines multiple similar or related feature fields into a single composite field. During field merging, the computer system retains the correspondence between the original field identifier and the composite field identifier, and considers fields of the same feature category, from the same detection source, or that have changed together in historical sample matching as mergeable fields. Field precision reduction reduces the precision or resolution of the field's value range. When field precision is reduced, the computer system adjusts the field's value range according to the number of significant digits, interval granularity, or category level allowed by the detection method, and retains the source record of the field's value range before adjustment. The computer system then statistically analyzes the differences in the candidate origin set, the differences in excluded origins, and the differences in the final origin before and after the above processing, and classifies the feature fields into different compression loss levels based on these differences. If the candidate origin set, excluded origin, and final origin remain the same before and after field processing, the feature field is classified as having a low compression loss level. If the candidate origin set or excluded origin changes after field processing but the final origin remains the same, the feature field is classified as having a medium compression loss level. If the final origin changes after field processing, or causes previously excluded origins to re-enter the final origin determination process, the feature field is classified as having a high compression loss level. The compression loss level is used to represent the degree of impact of the feature field on the matching results after compression processing, and is used to determine whether to use field retention, field merging, or field precision reduction compression methods for the feature field in the future.

[0032] After determining the compression loss level of each feature field, the computer system generates a compressed feature library based on the compression loss level. For feature fields that maintain stable matching results after compression, the system can store them using field merging or reduced field precision. For feature fields that would cause changes in the candidate origin set, excluded origin, or final origin after compression, the system increases their retention priority to maintain the accuracy of origin tracing results. When generating the compressed feature library, the computer system saves the compressed field identifier, the compressed field value range, the original field mapping relationship, the compression method identifier, and the compression loss level in the origin record, so that the original field meaning can be restored based on the compressed field when matching samples to be traced. After the compressed feature library is generated, the computer system performs playback matching again using historical samples and compares the candidate origin set, excluded origin, and final origin obtained from the playback with the matching results of the uncompressed feature library. If the two are consistent, it indicates that the compressed feature library can maintain stable tracing results while reducing data storage volume, and the computer system enables the compressed feature library. If the two are inconsistent, the computer system adjusts the compression loss level of the feature fields causing the discrepancy, regenerates the compressed feature library, and performs playback matching again until the activation conditions are met. During adjustment, the computer system can restore fields causing the final origin difference to the field retention method, restore fields causing the excluded origin difference to a higher precision field value range, or cancel the field merging relationship that caused false recall, so that the compressed feature library meets the playback consistency requirements before activation.

[0033] When matching garlic samples for traceability, the computer system acquires the traceability data of the garlic samples and generates matching data according to the same feature categories as garlic samples from known origins. During the generation of matching data, the computer system standardizes the detection methods, field units, field formats, and sample numbers of the garlic samples for traceability, and only writes fields corresponding to field identifiers in the compressed feature library or the original field mapping relationship into the matching data. Subsequently, the computer system matches the matching data with the already activated compressed feature library. During matching, the system compares the feature field values ​​in the matching data with the field value ranges of each origin record in the compressed feature library, filters candidate origin records, and outputs the origin traceability result based on the degree of matching of the candidate origin records. The origin traceability result may include the final origin, the candidate origin set, the feature fields that trigger matching, the feature fields that trigger exclusion, and the verification identifier. Through the above method, this embodiment can reduce the data size of the feature library while maintaining the consistency of historical sample matching results, and improve the matching efficiency and origin identification stability of garlic samples for traceability.

[0034] In one specific implementation, after recording the candidate origin set, excluded origin, and final origin for each historical sample, the computer system establishes exclusion constraints based on the historical sample matching process. These exclusion constraints describe the judgment conditions corresponding to the exclusion of a particular origin record. Specifically, the computer system associates and stores the excluded origin, the feature field that triggered the exclusion, and the corresponding field value range, ensuring that each exclusion constraint reflects the reason for the exclusion of the origin record. The exclusion constraints can also store the sample number, batch identifier, field source identifier, and the exclusion occurrence stage, which includes one of the following: candidate origin screening stage, candidate origin set verification stage, and final origin determination stage. By storing this information, the computer system can locate the field and origin record causing the discrepancy when differences occur during compressed feature library playback matching. When a compressed feature library is subsequently generated and historical sample playback matching is performed, the computer system compares the playback matching results of the compressed feature library with the matching results of the uncompressed feature library. If the playback matching results of the compressed feature library contain an origin that has been excluded in the uncompressed feature library matching results, it indicates that the origin record was falsely recalled due to feature compression. At this point, the computer system writes the exclusion constraint corresponding to the origin record into the anti-recall field of that origin record. The anti-recall field is used to store exclusion conditions for falsely recalled origin records when the compressed feature library participates in matching. Thus, during the matching process of garlic samples to be traced, the computer system calls the anti-recall field before determining the final origin, and removes the corresponding origin records in the candidate origin set according to the feature fields that trigger exclusion and the field value range stored in the anti-recall field, thereby avoiding the compression process causing excluded origins to re-participate in the final origin determination.

[0035] In another specific implementation, when the same excluded origin corresponds to multiple exclusion constraints, the computer system merges these multiple exclusion constraints. During merging, the computer system reads the feature field that triggers the exclusion and the field value range in each exclusion constraint, and determines whether the multiple field value ranges are continuous or overlapping. If the multiple field value ranges are continuous or overlapping, and all multiple field value ranges correspond to the same excluded origin, the computer system merges the multiple field value ranges into one field value range and writes the merged field value range into the recall field of the origin record. If multiple exclusion constraints correspond to different feature fields, the computer system merges them according to the feature fields respectively; if multiple exclusion constraints correspond to the same feature field but have different detection sources, the computer system retains the detection source identifier and writes it into the recall field separately to avoid data from different sources being incorrectly merged. Through this processing method, the compressed feature library does not need to repeatedly store multiple exclusion conditions with the same target and adjacent value ranges, thereby reducing the storage volume of duplicate exclusion data and reducing the number of times the recall field is read during the matching process.

[0036] In another specific implementation, before writing the exclusion constraint into the recall field, the computer system counts the number of times the exclusion constraint is triggered in historical sample playback matching. If the number reaches a preset threshold, it indicates that the exclusion constraint has stable triggering characteristics in historical samples, and the computer system writes the exclusion constraint into the recall field, allowing it to participate in the origin exclusion in the matching of garlic samples to be traced. If the number does not reach the preset threshold, it indicates that the triggering stability of the exclusion constraint is insufficient, and the computer system retains the exclusion constraint as an exclusion constraint to be verified. The exclusion constraint to be verified is stored in a compressed feature library, but is not used for origin exclusion when matching garlic samples to be traced. The exclusion constraint to be verified can be stored together with its corresponding origin record, trigger field, and historical sample number. When new samples with known origins are added or sample data is manually verified, the computer system recounts its triggering status and transfers it to the recall field after meeting the preset threshold. Thus, the computer system can prevent low-frequency, occasional exclusion conditions from directly affecting the final matching result of garlic samples to be traced.

[0037] In another specific implementation, after invoking the anti-recall field, the computer system records the removed origin records and the feature fields that triggered the removal. For the same origin record, the computer system counts the number of samples removed by the same feature field in the matching of garlic samples to be traced. If the number of samples reaches a preset sample number threshold, the computer system updates the value range of the field corresponding to the feature field to the priority exclusion interval of the origin record. The priority exclusion interval is used to be invoked before ordinary exclusion constraints during the matching stage. Specifically, when generating a candidate origin set or filtering the candidate origin set, the computer system first determines whether the corresponding feature field of the garlic sample to be traced falls into the priority exclusion interval. If it falls into the priority exclusion interval, the corresponding origin record is removed from the candidate origin set first, and then the ordinary exclusion constraint judgment is executed. If the corresponding feature field of the garlic sample to be traced is null, missing, or invalid, the computer system does not invoke the priority exclusion interval for removal, but instead switches to the field integrity processing process to avoid incorrect exclusion due to insufficient data. Through the above methods, the system can dynamically strengthen the exclusion conditions triggered at high frequencies according to the actual matching process, thereby improving the exclusion efficiency of the compressed feature library in subsequent matching and the stability of the source tracing results.

[0038] In one specific implementation, when classifying each feature field into compression loss levels, the computer system determines the role of each feature field in the origin identification result by combining the matching process of historical samples. Specifically, the computer system reads the matching records of historical samples in the uncompressed feature library and analyzes the role of each feature field in the candidate origin set screening, exclusion origin determination, and final origin determination processes. When the value of a feature field causes a historical sample to be further identified as the final origin from the candidate origin set, the computer system sets that feature field as the result anchor field. The result anchor field is used to maintain the stability of the final origin identification result. When the value of a feature field causes a candidate origin to become an excluded origin, the computer system sets that feature field as the exclusion anchor field. The exclusion anchor field is used to maintain the accuracy of the candidate origin exclusion process. For feature fields that participate in both final origin determination and candidate origin exclusion, the computer system records both their result anchor attribute and exclusion anchor attribute, and participates in subsequent compression processing according to the higher retention priority.

[0039] After determining the result anchoring fields and exclusion anchoring fields, the computer system assigns a compression loss level to these fields with a higher priority than that of other feature fields. Specifically, for result anchoring fields and exclusion anchoring fields, the computer system prioritizes retaining their field identifiers, value ranges, and corresponding origin records when generating the compressed feature library. This is to prevent compression processing from causing changes in the final origin or causing origins that should have been excluded to re-enter the candidate origin set. For other feature fields, the computer system can choose to compress them by merging fields, reducing field precision, or retaining fields, based on their impact on the candidate origin set, excluded origins, and final origin. For fields that have already been set as result anchoring fields or exclusion anchoring fields, if the computer system reduces field precision, it verifies whether the corresponding anchoring effect is still maintained during historical sample playback matching. If the anchoring effect disappears, the precision reduction processing for that field is canceled, and the original field value range is restored. Thus, the compressed feature library can retain feature fields that have a key impact on the origin identification results while reducing the data size, thereby improving the matching stability after compression.

[0040] In one specific implementation, after acquiring the traceability data of the garlic sample to be traced, the computer system performs an integrity comparison of the feature fields in the garlic sample to be traced, according to the feature fields involved in the matching in the compressed feature library. Specifically, the computer system reads each feature field of the garlic sample to be traced and determines whether each feature field has missing values, null values, or invalid values. Missing values ​​refer to data for which the garlic sample to be traced does not provide corresponding feature fields. Null values ​​refer to feature fields that have field identifiers but whose values ​​are not recorded. Invalid values ​​refer to feature field values ​​that do not conform to preset data formats, detection ranges, or matching rules. Through the above integrity comparison, the computer system can determine whether the garlic sample to be traced is missing feature fields configured with compression loss levels. For cases where the detection units are inconsistent, the field names are different but the field meanings are the same, or the spectral data preprocessing methods are different, the computer system first performs conversion or standardization according to the field mapping relationship and the detection method identifier; if the conversion or standardization still cannot meet the field format requirements, it is recorded as an invalid value.

[0041] When a garlic sample to be traced lacks a feature field configured with a compression loss level, the computer system does not perform numerical imputation on the missing field. Instead, it calls upon a set of alternative feature fields for matching. This set of alternative feature fields is a combination of fields pre-recorded during historical sample playback matching. When forming the alternative feature field set, the computer system simulates the matching state of historical samples lacking the same feature field and determines whether the final origin can still be consistent even without that feature field. If a certain field combination can maintain consistency in the final origin even without the same feature field, then that field combination is recorded as the alternative feature field set for the corresponding missing field. The alternative feature field set can include fields from different detection categories, as well as fields that have a stable co-occurrence relationship with the missing field in historical samples. It also saves the corresponding missing field identifier, applicable origin record, and historical sample source. Therefore, when a garlic sample to be traced lacks a field, the computer system can use field combinations validated by historical samples for matching, avoiding the introduction of uncertain data through numerical imputation.

[0042] Before invoking the alternative feature field group for matching, the computer system reads the number of times the alternative feature field group has been substituted and the number of times its place of origin matches in historical sample playback matching. The number of substitutions refers to the number of times the alternative feature field group has been used to substitute for missing fields in historical sample playback matching. The number of times its place of origin matches refers to the number of times the final place of origin obtained after playback matches the final place of origin in the non-missing state. The computer system generates a substitution confidence identifier based on the ratio of the place of origin matching count to the substitution count. The substitution confidence identifier is used to represent the substitution stability of the alternative feature field group in historical samples. When the ratio corresponding to the substitution confidence identifier reaches a preset confidence threshold, the computer system invokes the corresponding alternative feature field group to participate in the matching of the garlic sample to be traced. If the ratio corresponding to the substitution confidence identifier does not reach the preset confidence threshold, the computer system does not invoke the corresponding alternative feature field group to avoid the influence of field combinations with insufficient substitution stability on the place of origin traceability results. The computer system can also write uninvoked alternative feature field groups and corresponding missing fields together into the pending review record, so that the substitution matching results can be reviewed when supplementing test data later.

[0043] When multiple alternative feature field groups exist, the computer system reads the feature category to which the missing field belongs and compares the feature categories to which fields in each alternative feature field group belong. If an alternative feature field group belongs to a different feature category than the missing field, and this alternative feature field group can maintain consistency in the final origin during historical sample playback matching, the computer system prioritizes calling this alternative feature field group for matching. By prioritizing the use of alternative feature field groups of different feature categories, information bias caused by the absence of similar fields can be reduced. For alternative feature field groups that are not called, the computer system writes them into the supplementary matching record. The supplementary matching record is used to verify the matching results when supplementing sample data later. Specifically, when the missing fields of the garlic sample to be traced are subsequently supplemented or new test data is added, the computer system can read the supplementary matching record and use the uncalled alternative feature field groups to verify the original matching results. If the verification result is inconsistent with the original matching result, the computer system marks the garlic sample to be traced as a sample requiring manual verification and retains the original matching result, the verified matching result, and the feature field that caused the difference.

[0044] When the same set of alternative feature fields corresponds to more than two final origins in historical sample playback matching, the computer system counts the number of substitutions and the number of times the origin matches, respectively, and splits the substitution credibility identifier by final origin to generate a sub-origin substitution credibility identifier associated with the origin record. The sub-origin substitution credibility identifier is used to represent the substitution stability of the same set of alternative feature fields in different candidate origins. When matching garlic samples to be traced, the computer system first filters the sub-origin substitution credibility identifiers corresponding to the current candidate origin, and then determines whether the filtered sub-origin substitution credibility identifiers meet the preset credibility threshold. If the preset credibility threshold is met, the alternative feature field group associated with the sub-origin substitution credibility identifier is called to participate in the matching. If the preset credibility threshold is not met, the alternative feature field group is not called. If there is no corresponding sub-origin substitution credibility identifier for the current candidate origin, the computer system does not use the alternative feature field group as the basis for determining the final origin, and writes the candidate origin into the candidate origin record to be reviewed. In this way, the computer system can select a group of alternative feature fields with higher stability based on historical playback matching results when fields are missing, thereby improving the reliability of the origin matching results of garlic samples to be traced.

[0045] Example 1:

[0046] This embodiment provides a method for constructing, compressing, and re-recalling a garlic origin traceability feature database, which is executed by a computer system. The computer system acquires traceability data of garlic samples from known origins. The traceability data includes origin environment data, physicochemical test data, elemental test data, spectral data, harvest time data, and storage and transportation data. The data is then split into multiple feature fields according to preset feature categories. The field identifier, field value range, and origin identifier of each feature field are written into the corresponding origin record to form an uncompressed feature database.

[0047] In this embodiment, the known production areas include production areas A, B, and C, with a total of 90 historical samples, 30 of which correspond to each production area. The computer system extracts characteristic fields from the historical samples, including soil pH, selenium content, potassium content, soluble solids, near-infrared spectral absorption intensity at 1450 nm, near-infrared spectral absorption intensity at 1930 nm, peak area of ​​sulfur-containing compounds, and harvest month. The soil samples from production area A had a pH of 7.1 to 7.6, a selenium content of 0.042 mg / kg to 0.068 mg / kg, and an absorption intensity of 0.63 to 0.71 in the near-infrared 1450 nm band. The soil samples from production area B had a pH of 6.6 to 7.0, a selenium content of 0.025 mg / kg to 0.044 mg / kg, and an absorption intensity of 0.54 to 0.62 in the near-infrared 1450 nm band. The soil samples from production area C had a pH of 7.3 to 7.8, a selenium content of 0.018 mg / kg to 0.036 mg / kg, and an absorption intensity of 0.59 to 0.66 in the near-infrared 1450 nm band.

[0048] The computer system uses an uncompressed feature library to match 90 historical samples, recording the candidate origin set, excluded origin, and final origin for each historical sample. During matching, when a field value of a historical sample falls within the value range of a field in a certain origin record, the computer system retains that origin record in the candidate origin set; when a field value does not fall within the corresponding value range, the computer system treats that origin record as an excluded origin and records the feature field that triggered the exclusion.

[0049] After completing the historical sample matching, the computer system performs field deletion, field merging, or field precision reduction on each feature field, and statistically analyzes the differences between the candidate origin set, excluded origins, and final origin before and after processing. Field deletion refers to removing the corresponding feature field from the matching process; field merging refers to combining multiple fields with the same detection source or similar matching function into a composite field; field precision reduction refers to reducing the resolution or significant digits of a field's value range, but still retaining the field for matching. For example... Figure 2 As shown, the computer system maps soil pH, selenium content, potassium content, soluble solids, near-infrared 1450nm absorption intensity, near-infrared 1930nm absorption intensity, sulfur-containing compound peak area, and harvest month into compression loss values. It distinguishes between the processing directions of field merging, field precision reduction, and field retention by using a low loss threshold of 0.25 and a high loss threshold of 0.60, thus providing a comparable quantitative basis for the compression processing of different feature fields.

[0050] The computer system calculates the compression loss value of the feature field based on the matching difference before and after compression. The compression loss value satisfies the following relationship:

[0051]

[0052] in, This represents the compression loss value. This indicates the number of differences in the candidate origin set before and after compression. This indicates the difference in the number of production areas excluded before and after compression. Indicates whether the final origin changes before and after compression; if the final origin does not change. When the final place of origin changes to 0. Take 1, , , These represent the weights of the differences in the candidate origin set, the differences in the excluded origin, and the differences in the final origin, respectively.

[0053] For example, for the near-infrared 1450nm absorption intensity field, after the computer system reduces the precision of its value range, the number of differences in the candidate origin set during historical sample playback matching increases. The number of excluded origin differences is 0. The value is 1, and the final place of origin remains unchanged. The value is 0. Let it be 0. It is 0.2. It is 0.3. If the value is 0.5, then substituting it in, we get:

[0054]

[0055] The computer system classifies the field with a compression loss value of 0.3 as a medium compression loss level and processes it by reducing the field precision, but does not delete the field. Figure 2 The absorption intensity field in the mid-to-near infrared 1450nm band is located between the low loss threshold of 0.25 and the high loss threshold of 0.60, which indicates that although this field can reduce accuracy, it still needs to be included in the matching. The compression loss values ​​of selenium content and sulfur-containing compound peak area are 0.82 and 0.76, respectively, both higher than the high loss threshold of 0.60, indicating that they have a strong constraint effect on the identification of origin and the exclusion of candidate origins, and should be retained first.

[0056] When classifying compression loss levels, the computer system also reads the matching records of historical samples in the uncompressed feature library. Feature fields that lead historical samples to the final origin are set as result anchor fields, while feature fields that turn candidate origins into excluded origins are set as exclusion anchor fields. For result anchor fields and exclusion anchor fields, the computer system assigns them a higher retention priority than other feature fields and prioritizes retaining the corresponding field identifiers, field value ranges, and corresponding origin records in the compressed feature library.

[0057] When generating the compressed feature library, the computer system stores fields with low compression loss levels by merging or reducing their precision. For fields with medium compression loss levels, the system retains their field identifiers and limits the reduction in precision. For fields with high compression loss levels, the system retains their original value range. After generating the compressed feature library, the system performs playback matching again using 90 historical samples and compares the candidate origin set, excluded origins, and final origin obtained from the playback with the matching results of the uncompressed feature library. If the three types of results match, the compressed feature library is activated. If they do not match, the compression loss level of the feature fields causing the discrepancies is increased, and the compressed feature library is regenerated before playback matching is performed again.

[0058] The computer system also calculates the data saving ratio of the compressed feature library relative to the uncompressed feature library, with the compression ratio satisfying the following relationship:

[0059]

[0060] in, Indicates compression ratio. This indicates the amount of data in the uncompressed feature library. This indicates the amount of data in the compressed feature library. For example, the amount of data in the uncompressed feature library. The compressed feature library data size is 480KB. If the value is 315KB, then substituting it in, we get:

[0061]

[0062] That is, the compression ratio is approximately 34.4%. For example... Figure 3 As shown, the uncompressed feature library has a data size of 480KB, while the compressed feature library has a data size of 315KB, indicating a significant reduction in data size after compression; simultaneously, Figure 3 The consistency rate of the three types of replays—the candidate origin set, the excluded origin, and the final origin—all remained at 100%, indicating that this embodiment does not simply compress field data, but rather reduces the amount of feature library data under the constraint of replay consistency.

[0063] During the playback matching process of the compressed feature library, if a certain origin record is excluded in the uncompressed feature library matching results but enters the candidate origin set in the compressed feature library playback matching results, the computer system determines that the origin record has been falsely recalled. At this time, the computer system reads the trigger exclusion field and field value range corresponding to when the origin record was excluded from historical matching records, establishes an exclusion constraint, and writes this exclusion constraint into the anti-recall field of the origin record. The anti-recall field is used to exclude origin records falsely recalled due to compression processing when matching garlic samples to be traced.

[0064] When multiple exclusion constraints correspond to the same excluded origin, the computer system merges them according to the feature field that triggered the exclusion and the range of field values. If multiple exclusion constraints correspond to the same feature field, and the ranges of multiple field values ​​are consecutive or overlapping, and all correspond to the same excluded origin, the computer system merges the ranges of multiple field values ​​into a single recall field. For exclusion constraints with a low trigger count, the computer system does not directly write them into the recall field, but first counts the number of times they are triggered in historical sample playback matching; when the number of triggers reaches a preset threshold of 5, they are written into the recall field; when the number of triggers does not reach 5, they are retained as exclusion constraints to be verified. These exclusion constraints to be verified are stored in a compressed feature library, but do not participate in the origin exclusion in the matching of garlic samples to be traced.

[0065] For example, after reducing the precision of the near-infrared 1450nm absorption intensity field, region C was mistakenly recalled into the candidate origin set during some historical sample playback matching. The computer system reads the selenium content field and the near-infrared 1450nm absorption intensity field corresponding to the exclusion of region C from the uncompressed feature library, and associates them as exclusion constraints, writing them into the anti-recall field of the region C origin record. When the selenium content or near-infrared 1450nm absorption intensity of subsequent traceable samples meets the exclusion conditions in this anti-recall field, the computer system removes region C from the candidate origin set before determining the final origin. Figure 4 The execution order of compressed playback false recall, exclusion constraint writing, anti-recall field invocation, candidate origin removal and final origin output is further illustrated. This order shows that the anti-recall field is not an independent verification step, but is embedded in the candidate origin correction process before the final origin determination.

[0066] When matching garlic samples for traceability, the computer system acquires the traceability data of the samples and generates matching data according to the compressed feature library. One sample has a soil pH of 7.4, a selenium content of 0.051 mg / kg, and a near-infrared 1450 nm absorption intensity of 0.68. Because the near-infrared 1450 nm absorption intensity field for region C in the compressed feature library has undergone precision reduction processing, both region C and region A are included in the preliminary candidate origin set. Subsequently, the computer system calls the exclusion anchoring field and the recall field for verification. Since the selenium content field for region C ranges from 0.018 mg / kg to 0.036 mg / kg, it does not meet the matching condition of the selenium content of 0.051 mg / kg in the sample. Therefore, the computer system excludes region C and outputs region A as the final origin.

[0067] After the recall field is invoked, the computer system also records the removed origin records and the feature fields that triggered the removal. If the same origin record is removed from the near-infrared 1450nm spectral absorption intensity field in the subsequent 20 samples to be traced, and the number of samples reaches the preset sample quantity threshold of 10, the computer system updates the value range of the field corresponding to that field to the priority exclusion interval of that origin record. Figure 4 The cumulative removal amount corresponding to the sample sequence to be traced represents the dynamic update process. When the cumulative removal amount reaches the threshold of 10, the computer system forms a priority exclusion interval. In subsequent matching, the priority exclusion interval is called first to remove the corresponding place of origin record, and then the ordinary exclusion constraint is called, thereby reducing the number of invalid candidate place of origin records participating in the final place of origin determination.

[0068] Through this embodiment, the computer system can maintain consistency between the historical sample candidate origin set, excluded origin, and final origin playback, while merging fields, reducing field precision, and retaining necessary fields in the feature library, thereby reducing the amount of feature library data. Furthermore, by reducing the number of falsely recalled origin records after compression through anti-recall fields, merging exclusion constraints, trigger number thresholds, and priority exclusion intervals, the system can improve the matching efficiency and origin identification stability of garlic samples to be traced.

[0069] Example 2:

[0070] This embodiment provides a method for handling missing fields and calling alternative feature field groups in garlic origin traceability matching. This method is executed by a computer system. The computer system acquires the traceability data of the garlic sample to be traced and compares the completeness of the fields of the garlic sample to be traced according to the feature fields participating in the matching in the already enabled compressed feature library.

[0071] In this embodiment, the garlic sample to be traced is numbered G-2023-071. The fields obtained by the computer system include soil pH value (7.35), potassium content (1.82 g / kg), soluble solids (31.6%), near-infrared 1450 nm absorption intensity (0.67), and harvest month (May). The selenium content field was not obtained. The fields used for matching in the compressed feature library include soil pH value, selenium content, potassium content, near-infrared 1450 nm absorption intensity, sulfur compound peak area, and harvest month. Candidate production area records include production areas A, B, and C.

[0072] The computer system reads the characteristic fields of the garlic sample to be traced and determines whether there are missing, null, or invalid values ​​in each field. A missing value means that the garlic sample to be traced does not provide data for the corresponding characteristic field; a null value means that the corresponding characteristic field has a field identifier but its value is not recorded; an invalid value means that the value of the corresponding characteristic field does not conform to the preset detection range, data format, or matching rules. For sample G-2023-071, the computer system identified the selenium content field as a missing value. Figure 5 Using the fields from the compressed feature library for matching as the horizontal objects, soil pH, potassium content, near-infrared 1450nm spectral absorption intensity, and harvest month were marked as valid fields, selenium content as a missing field, and peak area of ​​sulfur-containing compounds as a field to be supplemented. This demonstrates that the field integrity comparison does not only determine whether a field name exists, but also distinguishes between valid fields, missing fields, and fields to be verified later.

[0073] When a garlic sample to be traced lacks a feature field configured with a compression loss level, the computer system does not fill in the missing field with the mean, nor does it use empirical values ​​to complete it. Instead, it calls upon a pre-recorded set of alternative feature fields from historical sample playback matching to participate in the matching. The set of alternative feature fields is a combination of fields formed under the condition that historical samples lack the same field, and this combination of fields can maintain consistency in the final origin during historical sample playback matching. Figure 5 The missing state corresponding to the selenium content field is below the valid field judgment line. Based on this, the computer system switches to the alternative feature field group calling process instead of filling the selenium content field with an estimated value.

[0074] Before invoking the alternative feature field group, the computer system reads the number of substitutions and the number of times the origin matches in the historical sample playback matching, and generates a substitution confidence value. The substitution confidence value satisfies the following relationship:

[0075]

[0076] Where T represents the substitution confidence value, M represents the number of times the substitution feature field group remained consistent with the final origin after participating in historical sample playback matching, and N represents the total number of times the substitution feature field group was used to substitute for missing fields. The higher the value of T, the higher the substitution stability of the substitution feature field group in historical samples.

[0077] In this embodiment, the alternative feature field group A consists of potassium content, near-infrared 1450nm absorption intensity, and harvest month. This alternative feature field group was used to replace the selenium content field a total of 40 times in historical samples. The final origin after 36 playbacks was consistent with the final origin in the non-missing state. Substituting the data, we obtain:

[0078]

[0079] The computer system sets the preset confidence threshold to 0.85. Since the surrogate confidence value of 0.90 for surrogate feature field group A reaches the preset confidence threshold, the computer system identifies surrogate feature field group A as a callable field group, allowing it to participate in the origin matching of garlic sample G-2023-071 to be traced. Figure 6 The substitution confidence value and the number of replay statistics are used to represent the selection process of substitution field groups. Substitution feature field group A corresponds to 36 times of consistent origin and 40 substitution calls, with a confidence value higher than the confidence threshold of 0.85. Substitution feature field group B corresponds to 25 times of consistent origin and 34 substitution calls, with a confidence value of approximately 0.74, which is lower than the confidence threshold of 0.85. Therefore, it is only retained for subsequent review and is not used as the priority field group for this call.

[0080] When multiple alternative feature field groups exist, the computer system reads the feature category to which the missing field belongs and compares the feature categories of fields in each alternative feature field group. The selenium content field belongs to the element detection category. Alternative feature field group A includes potassium content, near-infrared 1450nm absorption intensity, and harvest month; alternative feature field group B includes soil pH, soluble solids, and peak area of ​​sulfur compounds. The computer system prioritizes using alternative feature field group A, which has a significantly different category from the missing field and has stable historical playback results, and adds the unused alternative feature field group B to the pending matching record. Figure 6 The difference in confidence values ​​between alternative feature field group A and alternative feature field group B enables the computer system to select which field to call based on historical playback stability rather than the number of fields when multiple alternative field groups exist.

[0081] After the computer system retrieved the substitution feature field group A, it matched the soil pH value (7.35), potassium content (1.82 g / kg), near-infrared 1450 nm absorption intensity (0.67), and harvest month (May) of sample G-2023-071 with the origin records in the compressed feature library, obtaining production areas A and C as candidate production areas. Since the same substitution feature field group corresponded to more than two candidate production areas in this matching, the computer system further retrieved the sub-origin substitution confidence identifiers. Figure 7 The sequence of field missing identification, replacement field group A call, origin-based reliable filtering, output of origin A and supplementary verification indicates that after the replacement field group call, it still needs to go through origin-based reliable identification filtering, and cannot directly use all the origins in the candidate origin set as the final output.

[0082] The confidence values ​​for place-of-origin substitution satisfy the following relationship:

[0083]

[0084] Where Tp represents the sub-origin substitution confidence value corresponding to a candidate origin, Mp represents the number of times the substitution feature field group maintains the final origin consistency in the candidate origin, and Np represents the number of times the substitution feature field group is used to substitute for missing fields in the candidate origin. The sub-origin substitution confidence value is used to represent the substitution stability of the same substitution feature field group in different candidate origins.

[0085] In this embodiment, for production area A, the number of times the substitute feature field group A maintains the final production location consistency in this candidate production area is Mp, which is 18, and the number of times it is used to substitute missing fields in this candidate production area is Np, which is 20. Substituting the data, we get:

[0086]

[0087] For region C, the number of times the substitution feature field group A maintains the final origin consistency in this candidate region is Mp, and the number of times it is used to substitute missing fields in this candidate region is Np, which is 12. Substituting the data, we get:

[0088]

[0089] The computer system sets the preset confidence threshold to 0.85. The confidence value for the sub-origin substitution of region A (0.90) reaches the preset confidence threshold, while the confidence value for the sub-origin substitution of region C (0.583) does not. Therefore, the computer system uses the substitution feature field group A associated with the sub-origin substitution confidence identifier of region A to participate in the final origin determination, but does not use the substitution feature field group A associated with the sub-origin substitution confidence identifier of region C as the basis for final origin determination, and adds region C to the pending matching record. Figure 7 The confidence value of place of origin substitution in production area A is above the confidence threshold of 0.85, while the confidence value of place of origin substitution in production area C is below the confidence threshold of 0.85. This difference indicates that the confidence level of the same substitution feature field group can be different in different candidate production areas.

[0090] After the selenium content of garlic sample G-2023-071 was subsequently supplemented to 0.052 mg / kg, the computer system read the supplemented matching record and re-matched it using the supplemented complete fields. If the final origin obtained by re-matching is still production area A, the computer system marks the original matching result as passed verification; if the final origin obtained by re-matching changes, the computer system generates a manual verification mark and saves the original matching result, the verified matching result, the missing fields, and the characteristic fields that caused the difference. Figure 7 The supplementary verification path in the document indicates that although production area C was not output as the final production location, it is still retained with the matching record to be supplemented, so that the original matching result can be traced and verified after the selenium content field is supplemented.

[0091] In this way, when garlic samples to be traced contain missing, null, or invalid values, the computer system can avoid introducing uncertain data through numerical completion. Instead, it uses historically validated alternative feature field groups, alternative trusted identifiers, and origin-specific alternative trusted identifiers for matching. Figure 5 The field integrity comparison shown Figure 6 The alternative field group confidence value filtering shown and Figure 7 The provided reliable identification of production areas and supplementary verification paths enhance the reliability and verifiability of garlic origin tracing results in scenarios with missing data.

Claims

1. A method for constructing, compressing, and matching a garlic origin traceability feature database, applied to a computer system, characterized by: The traceability data of garlic samples from known production areas is obtained and split into multiple feature fields according to feature categories. The field identifier, field value range and production area identifier of each feature field are written into the corresponding production area record to form an uncompressed feature library. The uncompressed feature library is then used to match historical samples and record the candidate production area set, excluded production area and final production area corresponding to each historical sample. Each feature field is processed by deleting, merging, or reducing precision. The differences between the candidate origin set, excluded origin, and final origin before and after processing are statistically analyzed. Based on the differences, the feature fields are divided into compression loss levels. The compression loss level is used to determine the field retention, merging, or precision reduction method. A compressed feature library is generated and verified by replaying historical samples. If the candidate origin set, excluded origin, and final origin obtained from the replay match the results of the uncompressed feature library, the compressed feature library is enabled. If they do not match, the compression loss level of the corresponding feature field is adjusted and the replay is performed again. Obtain the traceability data of the garlic sample to be traced, generate matching data according to the feature category, match the matching data with the enabled compressed feature library, and output the origin traceability result.

2. The method according to claim 1, characterized in that: After recording the candidate origin set, excluded origin, and final origin for each historical sample, an exclusion constraint is established. The exclusion constraint is associated with the excluded origin, the feature field that triggered the exclusion, and the value range of the field. When the playback matching result of the compressed feature library contains an origin that has been excluded in the matching result of the uncompressed feature library, the corresponding exclusion constraint is written into the anti-recall field of the origin record. The anti-recall field is used to exclude origin records that are mistakenly recalled after compression. When matching garlic samples to be traced, the anti-recall field is called to remove the corresponding origin record before determining the final origin.

3. The method according to claim 1, characterized in that: When classifying feature fields into compression loss levels, feature fields that enable historical samples to enter the final origin are set as result anchor fields, and feature fields that enable candidate origins to be excluded origins are set as exclusion anchor fields. The result anchor fields and the exclusion anchor fields are configured to retain compression loss levels with higher priority than other feature fields, and the corresponding field identifiers, field value ranges, and corresponding origin records are retained in the compression feature library.

4. The method according to claim 1, characterized in that: The feature fields in the garlic sample to be traced are compared with the feature fields in the compressed feature library for completeness. The completeness comparison includes the identification of missing values, null values ​​or invalid values. When the garlic sample to be traced is missing a feature field configured with a compression loss level, the missing field is not numerically filled in. Instead, a group of alternative feature fields that has been pre-recorded in the historical sample playback matching and maintains the same final origin when the same feature field is missing is called to participate in the matching.

5. The method according to claim 2, characterized in that: When the same excluded origin corresponds to multiple exclusion constraints, the multiple exclusion constraints are merged according to the feature field that triggers the exclusion and the value range of the field; when the value ranges of multiple fields are consecutive or overlap and correspond to the same excluded origin, the value ranges of multiple fields are merged into a single anti-recall field.

6. The method according to claim 2, characterized in that: The number of times the exclusion constraint is triggered in historical sample playback matching is counted; when the number reaches a preset threshold, the exclusion constraint is written into the anti-recall field; when the number does not reach the preset threshold, the exclusion constraint is retained as an exclusion constraint to be verified. The exclusion constraint to be verified is stored in the compressed feature library but does not participate in the origin exclusion in the matching of garlic samples to be traced.

7. The method according to claim 2, characterized in that: Record the removed origin records and the feature fields that triggered the removal; when the same origin record is removed by the same feature field in the matching of garlic samples to be traced within a preset sample quantity threshold, update the field value range corresponding to the feature field to the priority exclusion interval of the origin record. The priority exclusion interval is called before the ordinary exclusion constraint item in the matching stage.

8. The method according to claim 4, characterized in that: The number of substitutions and the number of times the place of origin matches in the historical sample playback matching of the substitution feature field group are read, and a substitution credibility identifier is generated based on the ratio of the number of times the place of origin matches to the number of substitutions; when the ratio corresponding to the substitution credibility identifier reaches a preset credibility threshold, the corresponding substitution feature field group is called to participate in the matching.

9. The method according to claim 4, characterized in that: When multiple alternative feature field groups exist, the feature category to which the missing field belongs is read. Alternative feature field groups that belong to different feature categories than the missing field and have maintained the same final origin in historical sample playback matching are called first. Alternative feature field groups that have not been called are written into the supplementary matching record. The supplementary matching record is used to verify the matching results when supplementing sample data later.

10. The method according to claim 8, characterized in that: When the same group of substitution feature fields corresponds to more than two final origins in historical sample playback matching, the substitution count and the number of times the origin matches are counted separately according to the final origin, and the substitution credibility identifier is split according to the final origin to generate a sub-origin substitution credibility identifier associated with the origin record; When matching garlic samples to be traced, first select the corresponding alternative trusted identifiers based on the candidate origin, and then call the alternative feature field group associated with the alternative trusted identifiers that meet the preset trusted threshold.

Citation Information

Patent Citations

  • Agricultural product producing area tracing method and system based on big data

    CN116883026A

  • Agricultural product anti-counterfeiting traceability method and device, electronic equipment and storage medium

    CN120338803A

  • Safety production informatization data management method and system

    CN120429283A

  • Propolis component intelligent identification and traceability system and method

    CN120948407A