Aquatic product producing area tracing method and system
The classification model is constructed through mineral element detection and random forest algorithms, which solves the problem of differentiation between counterfeit and multi-production areas in the traceability of hairy crabs in Yangcheng Lake, realizes high-precision traceability, and provides reliable traceability technical support.
Patent Information
- Application Number
- CN202510546168.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
AI Technical Summary
The existing technology is difficult to effectively distinguish between Yangcheng Lake hairy crabs and peripheral breeding farms. Traditional anti-counterfeiting technology is prone to counterfeiting, and the existing methods are insufficient to handle nonlinear interaction effects in the multi-production area, resulting in insufficient spatial resolution.
By using mineral element detection combined with random forest algorithm, a multi-level random forest classification model is constructed, mineral elements with high importance scores are selected as core features, and a random forest classification model is established to achieve accurate traceability of aquatic products.
It has achieved accurate distinction between the core production areas of Yangcheng Lake and the peripheral breeding farms, provided a verifiable and difficult to imitate traceability solution, and the spatial resolution reaches ≤10km, improving the protection ability of geographical indication products.
Smart Images

Figure CN120450718A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of product origin tracing, and in particular relates to a method and system for tracing the origin of aquatic products. Background Art
[0002] Yangcheng Lake hairy crabs, a key geographical indication product, saw an industry valued at 6.5 billion yuan in 2022. However, counterfeit and substandard products account for over 80% of the products circulating in the market. Traditional crab clasp anti-counterfeiting systems, due to inherent flaws such as ease of duplication and lack of essential relevance to product quality, are no longer sufficient to protect geographical indication products.
[0003] Existing patent 1 (publication number CN114241248B) discloses a method and system for tracing the origin of river crabs. The method includes: obtaining an image of the carapace of the purchased target river crab; uploading the image of the carapace of the target river crab to the blockchain, and determining the target origin traceability information of the target river crab by calling the traceability smart contract deployed on the blockchain.
[0004] Existing patent 2 (publication number CN119313357A) is a method, system and crab clasp for anti-counterfeiting and traceability of hairy crabs. The initial information of hairy crabs is entered into the hairy crab anti-counterfeiting and traceability system; a unique anti-counterfeiting and traceability code is generated according to the initial information of each hairy crab, and the anti-counterfeiting and traceability code is stored in the chip of the crab clasp. The crab clasp is bundled with the hairy crab corresponding to the anti-counterfeiting and traceability code; effective anti-counterfeiting and traceability of each hairy crab is achieved, thereby effectively preventing counterfeit products from entering the market.
[0005] Patents 1-CN108985410A and 2-CN119313357A both disclose methods for anti-counterfeiting and traceability of hairy crabs. Existing hairy crab traceability technologies fall into three main categories: image recognition based on morphological characterization, anti-counterfeiting traceability code analysis, and elemental traceability based on environmental response characteristics. However, existing research is largely limited to traditional statistical methods such as linear discriminant analysis (LDA), which struggle to effectively analyze the nonlinear interactions between high-yield intervals within the same watershed, resulting in insufficient spatial resolution. This present invention is based on these findings. Summary of the Invention
[0006] In order to solve the above-mentioned technical problems existing in the prior art, the present invention provides a method and system for tracing the origin of aquatic products, so as to solve the problem that the traditional anti-counterfeiting technology of Yangcheng Lake hairy crabs in the prior art is easy to be counterfeited, and at the same time solve the problem of accurately distinguishing multiple production areas in the basin, and provide a verifiable and difficult-to-imitate technical solution.
[0007] To achieve the above object, the technical solution of the present invention is as follows:
[0008] In a first aspect, a method for tracing the origin of aquatic products comprises:
[0009] S1. Processing the collected aquatic products to obtain aquatic product samples;
[0010] S2. Performing mineral element testing on the aquatic product sample to obtain the mineral element content of the aquatic product sample;
[0011] S3. Calculate the characteristic importance score of the mineral element content using a random forest algorithm, screen out mineral elements whose importance score is greater than a preset score, and use the mineral elements as core features;
[0012] S4. Establishing a random forest classification model based on the core features;
[0013] S5. Obtain aquatic product samples that need to be traced to their origin, process the aquatic product samples that need to be traced to their origin based on the random forest classification model, obtain a source origin classification result of the aquatic product samples that need to be traced, and realize origin traceability for the aquatic product samples.
[0014] Furthermore, before establishing the random forest classification model, the optimization of the random forest classification model is also included, including:
[0015] According to the random forest algorithm, different numbers of decision trees are selected to import the random forest classification model;
[0016] Train the model and predict the model accuracy score. The test model accuracy scores vary when different numbers of decision trees are selected.
[0017] Combined with the model accuracy results output by Python using test data with different features, and taking into account the maintenance of the prediction performance of the random forest classification model, a suitable decision tree is selected to import the random forest classification model.
[0018] Going further, the specific steps for selecting the number of decision trees include:
[0019] Select different numbers of decision trees and calculate the model accuracy score based on different numbers of decision trees;
[0020] When the accuracy score of the model shows a gradually increasing trend, continue to increase the number of decision trees;
[0021] After comprehensive consideration, select an appropriate number of decision trees to import into the random forest classification model.
[0022] Furthermore, the specific steps of obtaining the core features include:
[0023] Rank all mineral element characteristics in order of importance;
[0024] Remove features in the model whose importance scores are less than a first preset score, rebuild a classification model with a preset number of decision trees, and calculate the model accuracy score;
[0025] Remove features in the model whose importance scores are less than a second preset score, re-establish a classification model with a preset number of decision trees, and calculate the model accuracy score;
[0026] The matching model is selected by evaluating the accuracy score of the model, and the mineral element characteristics in the selected model are used as the core characteristics.
[0027] Furthermore, the specific steps of the mineral element detection include:
[0028] S21, weighing the aquatic product sample and placing it in a digestion tank;
[0029] S22, wetting the sample with ultrapure water, adding nitric acid, covering the sample, and letting it stand for a preset time;
[0030] S23, heating and digesting the mixture after standing for a preset time;
[0031] S24. After digestion, open the lid to remove the gas in the tank and transfer the digestion tank to the acid remover for heating;
[0032] S25, allowing the heated digestion tank to stand and then adding ultrapure water to a fixed volume to obtain a solution to be tested;
[0033] S26. Using a mineral element detection device to perform mineral element detection on the solution to be detected.
[0034] Furthermore, the mineral elements include any one or more of calcium, copper, magnesium, sodium, lanthanum, zinc, potassium, gadolinium, barium, cerium, mercury, manganese, vanadium, thulium, praseodymium, dysprosium, samarium, arsenic, selenium, europium, cobalt, iron, cadmium, ytterbium, erbium, neodymium, strontium, chromium and titanium.
[0035] Furthermore, the mineral element detection equipment adopts: an inductively coupled plasma emission spectrometer and an inductively coupled plasma mass spectrometer.
[0036] Furthermore, the detection conditions of the inductively coupled plasma emission spectrometer are: RF power 1100-1200 W, atomizing gas flow rate 0.65-0.75 L / min, cooling gas flow rate 12.0-13.0 L / min, auxiliary gas flow rate 0.45-0.55 L / min, and peristaltic pump speed 40-50 r / min.
[0037] Furthermore, the detection conditions of the inductively coupled plasma mass spectrometer are as follows: RF power 1500-1600 W, cooling gas flow rate 12-16 L / min, auxiliary gas flow rate 0.7-0.9 L / min, nebulizing gas flow rate 0.8-1.2 L / min, helium flow rate 4.2-4.8 mL / min, residence time 0.01-0.03 s, scan number 95-105 times, injection time 42-47 s, and peristaltic pump speed 40-50 r / min.
[0038] On the other hand, a traceability system for aquatic products includes:
[0039] Detection module: used to obtain aquatic product samples after processing the collected aquatic product samples, and perform mineral element detection on the aquatic product samples to obtain the mineral element content of the aquatic product samples;
[0040] Data processing module: used to calculate the characteristic importance score of the sample data of the detection module through the random forest algorithm, screen out the mineral elements whose importance score is greater than the preset score, and use the mineral elements as the core features;
[0041] Classification model module: used to establish a random forest classification model based on the core features of the data processing module;
[0042] Classification output module: used to import aquatic product samples that need to be traced into the classification model module, and output the source origin classification results of the aquatic product samples that need to be traced.
[0043] Compared with the existing technology, the origin traceability method and system of the aquatic products provided by the present invention innovatively integrates mineral element characteristics and random forest algorithms. By constructing a multi-level random forest classification model, it can achieve accurate distinction between the core production area of Yangcheng Lake and the peripheral farms with a spatial resolution of ≤10km. This is the first time that accurate distinction between the core production area of Yangcheng Lake and the peripheral farms is achieved, providing a verifiable and difficult-to-imitate technical solution for geographical indication protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A schematic diagram of the process for tracing the origin of aquatic products provided by the present invention;
[0045] Figure 2 A framework diagram of the origin tracing system for aquatic products provided by the present invention;
[0046] Figure 3 This is a schematic diagram of the output model accuracy score provided in Example 1 of the present invention;
[0047] Figure 4 A schematic diagram of the importance scores of features provided in the first embodiment of the present invention;
[0048] Figure 5 A schematic diagram of the output model accuracy score provided in the second embodiment of the present invention;
[0049] Figure 6 A schematic diagram of feature importance scores provided in the second embodiment of the present invention. DETAILED DESCRIPTION
[0050] The technical solution of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are not all embodiments of the present invention. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0051] It should be noted that, unless otherwise specifically stated, the relative arrangements of components and steps, and numerical expressions set forth in these embodiments should not be construed as limiting the scope of the present invention.
[0052] The following description of exemplary embodiments is merely illustrative and is not intended to limit the present invention, its application, or use in any sense. Technologies, methods, and apparatus known to those skilled in the art may not be discussed in detail herein, but to the extent applicable, such technologies, methods, and apparatuses should be considered part of this specification.
[0053] See Figure 1 , Figure 1 This is a method for tracing the origin of aquatic products proposed in an embodiment of the present invention, which detects the origin classification of the aquatic products. The specific processing steps include:
[0054] S11. Data collection and processing: Processing the collected aquatic products to obtain aquatic product samples;
[0055] The crabs were collected from Yangcheng Lake, surrounding ponds, Taihu Lake, Gaoyou Lake, Gucheng Lake, and Xinghua. A total of 904 healthy, mature male crabs were collected. The crabs had intact carapaces and weighed between 180 and 200 grams. They were dried at 75-85°C, ground through an 80-mesh sieve, and stored in polyethylene ziplock bags until ready for use.
[0056] S12. Data mineral element detection: Perform mineral element detection on the aquatic product sample to obtain the mineral element content of the aquatic product sample; the specific steps of the mineral element detection include:
[0057] S121. Weigh 0.1 g of aquatic product sample and place it in a microwave digestion tank. The weight of the aquatic product sample should be accurate to 0.001 g.
[0058] S122. Wet the sample with a small amount of ultrapure water, add 6.5 ml of nitric acid, cover and let stand for a preset time. Usually the preset time is overnight, but the time can be adjusted according to specific needs, but it should not be less than 12 hours.
[0059] S123, after standing, heating the microwave digestion tank for digestion; the heating digestion is performed according to a certain heating program. The program provided in this embodiment is as follows: 115-125°C for 5 minutes, 155-165°C for 10 minutes, and 175-185°C for 18-22 minutes;
[0060] S124. After the microwave digestion tank has cooled, take it out and open the lid of the microwave digestion tank to remove the gas inside the tank. Transfer the digestion tank to the acid remover for heating at 95-105°C for 30 minutes.
[0061] S125, allowing the heated microwave digestion tank to stand, then adding ultrapure water to dilute to 50 ml, and mixing evenly to obtain a solution to be tested;
[0062] S126, using a mineral element detection device to perform mineral element detection on the solution to be detected;
[0063] Among them, potassium (K), sodium (Na), calcium (Ca), magnesium (Mg), iron (Fe), manganese (Mn), copper (Cu), zinc (Zn), titanium (Ti), strontium (Sr), and barium (Ba) in the solution to be tested are detected using an inductively coupled plasma optical emission spectrometer; the inductively coupled plasma optical emission spectrometer can be the iCAP PRO X device from Thermo Fisher Scientific.
[0064] Inductively coupled plasma optical emission spectrometer (ICP-EOS) detection conditions were: RF power 1100-1200 W, nebulizing gas flow rate 0.65-0.75 L / min, cooling gas flow rate 12.0-13.0 L / min, auxiliary gas flow rate 0.45-0.55 L / min, and peristaltic pump speed 40-50 rpm. Potassium, sodium, calcium, and magnesium were observed in vertical mode, while iron, manganese, copper, zinc, strontium, titanium, and barium were observed in horizontal mode.
[0065] In addition, an inductively coupled plasma mass spectrometer is used to detect vanadium (V), chromium (Cr), cobalt (Co), arsenic (As), selenium (Se), cadmium (Cd), mercury (Hg), lanthanum (La), cerium (Ce), praseodymium (Pr), neodymium (Nd), samarium (Sm), europium (Eu), gadolinium (Gd), dysprosium (Dy), erbium (Er), thulium (Tm), and ytterbium (Yb) in the solution to be tested; the inductively coupled plasma mass spectrometer can be the iCAPQ equipment from Thermo Fisher Scientific.
[0066] The detection conditions of the inductively coupled plasma mass spectrometer are as follows: RF power 1500-1600 W, cooling gas flow rate 12-16 L / min, auxiliary gas flow rate 0.7-0.9 L / min, nebulizing gas flow rate 0.8-1.2 L / min, helium flow rate 4.2-4.8 mL / min, dwell time 0.01-0.03 s, scan number 95-105 times, injection time 42-47 s, and peristaltic pump speed 40-50 r / min.
[0067] Scallop biological component analysis standard substances were used as quality control samples, and the test results were all within the allowable range of the standard values.
[0068] S13, mineral element screening: using a random forest algorithm to calculate the characteristic importance score of the mineral element content, screening out mineral elements with the importance score greater than a preset score, and taking the mineral elements with the importance score greater than the preset score as core features;
[0069] All data analyses were performed using the Random Forest algorithm in the Jupyter Notebook application. To ensure randomness and consistency in data partitioning, aquatic product samples were divided into training and test sets in a ratio of 7:3, ensuring a consistent proportion of crab samples from different production areas.
[0070] S131: Build a random forest model based on all features:
[0071] According to the random forest algorithm, we first select different numbers of decision trees to import into the random forest classifier, train the model on the training set, and predict the model accuracy score. The results are as follows Figure 1 As shown in the figure, the test model accuracy scores are slightly different when different numbers of decision trees are selected. The highest output model accuracy score is 0.9737 when there are 50 decision trees. The model accuracy scores are all 0.9699 when there are 10, 100, 150 and 200 decision trees. When the number of decision trees is increased to 250, the model accuracy score is 0.9662.
[0072] In combination with Python's output of model accuracy results using test data with different features, in order to maintain the predictive performance of the classification model, 100 decision trees were selected to be imported into the random forest classifier and the model was trained on the training set. The number of decision trees can also be adjusted according to the actual number of collected samples and is not limited to the above number.
[0073] S132. Filter core features based on importance scores
[0074] Random Forest can simplify the model and improve predictive performance by calculating feature importance and selecting the most predictive features. Based on the importance of the metal elements output by the Random Forest classification, all feature importance scores entering the Random Forest classification model are ranked. Features with importance scores less than 0.02-0.03 in the model are extracted to form the core features with the highest predictive power. The screening score can be adjusted based on the sample data type and number, and is not limited to 0.02-0.03.
[0075] S14. Building a classification model: Building a random forest classification model based on the core features;
[0076] By selecting the most predictive features, the model can be simplified and the prediction performance can be improved.
[0077] S15. Output model results: Process the aquatic product samples that require origin traceability based on the random forest classification model, obtain the origin classification results of the aquatic product samples that require origin traceability through a classification report, and implement origin traceability for the aquatic product samples;
[0078] The classification report is an important tool for evaluating the performance of random forest classification models. Its core indicators include accuracy, precision, recall, and F1 score. Accuracy refers to the proportion of samples correctly predicted by the model to the total number of samples, reflecting the overall prediction ability.
[0079] Precision measures the proportion of samples predicted as positive that are actually positive, reflecting the model's accuracy in identifying positive examples. Recall indicates the proportion of actual positive examples that are correctly predicted, reflecting the model's ability to capture positive examples. The F1 score is the harmonic mean of precision and recall, comprehensively balancing the performance of both. A higher value indicates better model performance. By analyzing these metrics, we can comprehensively evaluate the classification performance of a model across different categories.
[0080] The random forest algorithm performed exceptionally well in this paper, automatically selecting important features and handling complex nonlinear relationships. Feature importance analysis screened for discriminative mineral elements, avoiding multicollinearity. The model performed reliably across diverse origin classification tasks and demonstrated high generalization capabilities. The ensemble learning mechanism of the random forest algorithm reduced overfitting, and the model performed well in external validation.
[0081] This method innovatively adopts a dynamic sample weight mechanism and a visual decision support system, and has been successfully applied in geographical indication protection projects. It provides reliable classification solutions for areas with different sample sizes. Its technical indicators are better than industry benchmarks and it has significant technical innovation and application value.
[0082] Second, see Figure 2The present invention provides a system for tracing the origin of aquatic products, which is used to implement the above-mentioned method for tracing the origin of aquatic products. The system includes:
[0083] Detection module: used to obtain aquatic product samples after processing the collected aquatic product samples, and perform mineral element detection on the aquatic product samples to obtain the mineral element content of the aquatic product samples;
[0084] Data processing module: used to calculate the characteristic importance score of the sample data of the detection module through the random forest algorithm, screen out the mineral elements whose importance score is greater than the preset score, and use the mineral elements as the core features;
[0085] Classification model module: used to establish a random forest classification model based on the core features of the data processing module;
[0086] Classification output module: used to import the aquatic product samples that need to be traced to the origin into the classification model module, and output the source origin classification results of the aquatic product samples that need to be traced.
[0087] Example 1
[0088] An embodiment of the present invention provides a method for tracing the origin of an aquatic product. In this embodiment, the aquatic product is Chinese mitten crab, and the detection result of the origin classification of the aquatic product is Yangcheng Lake area and non-Yangcheng Lake area. The specific processing steps include:
[0089] S21. 904 healthy adult male crabs were collected from the Yangcheng Lake area, surrounding ponds, Taihu Lake, Gaoyou Lake, Gucheng Lake, and Xinghua area. The crabs were dried at 80°C, ground through an 80-mesh sieve, and stored in polyethylene ziplock bags for later use.
[0090] S22. Using an inductively coupled plasma emission spectrometer and an inductively coupled plasma mass spectrometer, perform mineral element detection on the aquatic product sample to obtain the mineral element content of the aquatic product sample;
[0091] S23, using a random forest algorithm to calculate the characteristic importance score of the mineral element content, screening out mineral elements whose importance score is greater than a preset score, and using the mineral elements as core features;
[0092] S24. Establishing a random forest classification model based on the core features;
[0093] Based on the random forest algorithm, we first select different numbers of decision trees and import them into the random forest classifier. The model is then trained on the training set and the model accuracy score is predicted. Random forest uses decision trees as its base model and improves the model's robustness and generalization capabilities by randomly selecting samples and features. The multiple decision trees in a random forest can collaborate with each other, enabling the model to better handle complex datasets and tasks. Compared to a single decision tree, random forests generally achieve better performance on classification and regression problems.
[0094] See for example Figure 3 As shown in the figure, the output model accuracy scores vary slightly when different numbers of decision trees are selected. The highest output model accuracy score is 0.9737 when there are 50 decision trees. The model accuracy scores are all 0.9699 when there are 10, 100, 150, and 200 decision trees. When the number of decision trees is increased to 250, the model accuracy score is 0.9662. The accuracy of the model containing the same conditional features is tested in sequence with 50 decision trees, 100 decision trees, and 150 decision trees. When all features are included, the accuracy scores of the models with 50, 100, and 150 decision trees are 0.9737, 0.9737, and 0.9699, respectively; when features with importance scores <0.02 are removed, the accuracy scores of the models with 50, 100, and 150 decision trees are 0.9586, 0.9662, and 0.9586, respectively; when features with importance scores <0.03 are removed, the accuracy scores of the models with 50, 100, and 150 decision trees are 0.9624, 0.9737, and 0.9586, respectively. It can be seen that under the same conditions, the test data model with 100 decision trees has the highest accuracy.
[0095] In summary, combined with Python's output of model accuracy results using test data with different features, in order to maintain the predictive performance of the classification model, 100 decision trees were selected to import the random forest classifier and the model was trained on the training set.
[0096] See Figure 4 As shown, all the features in order of importance are Ca, Cu, Mg, Na, La, Zn, K, Gd, Ba, Ce, Hg, Mn, V, Tm, Pr, Dy, Sm, As, Se, Eu, Co, Fe, Cd, Yb, Er, Nd, Sr, Cr, Ti.
[0097] After removing the eight features Fe, Cd, Yb, Er, Nd, Sr, Cr, and Ti with importance scores less than 0.02 from the model, a new classification model with 100 decision trees was established, and the model accuracy score was 0.9662. At this time, 21 features entered the model, namely Ca, Cu, Mg, Na, La, Zn, K, Gd, Ba, Ce, Hg, Mn, V, Tm, Pr, Dy, Sm, As, Se, Eu, and Co.
[0098] After removing the 18 features Mn, V, Tm, Pr, Dy, Sm, As, Se, Eu, Co, Fe, Cd, Yb, Er, Nd, Sr, Cr, and Ti with importance scores less than 0.03 from the model, a new classification model with 100 decision trees was established, with a model accuracy score of 0.9737. At this time, 11 features entered the model, namely Ca, Cu, Mg, Na, La, Zn, K, Gd, Ba, Ce, and Hg. It can be seen that after removing the 8 features with importance scores less than 0.02 from the model, and then removing the 10 features with importance scores between 0.02 and 0.03, the accuracy of the re-modeled output model increased.
[0099] Considering the accuracy of the random forest classification model and the economic cost, time efficiency and workload of future actual verification work, a random forest classification model with 11 features (Ca, Cu, Mg, Na, La, Zn, K, Gd, Ba, Ce, Hg) was selected.
[0100] S25. Process the aquatic product samples based on the random forest classification model to obtain the source origin classification results of the aquatic product samples that need to be traced, thereby realizing the origin traceability of the aquatic product samples.
[0101] Table 1 shows the classification results for the Yangcheng Lake region and non-Yangcheng Lake region classification models. The model for the non-Yangcheng Lake region achieved an accuracy of 0.98, while the Yangcheng Lake region achieved an accuracy of 0.94. Leveraging its data volume advantage and using feature importance analysis to select key metrics, the non-Yangcheng Lake region model achieved 98% accuracy, 99% recall, and 98% F1 score. Both models, through cross-validation to ensure generalization, achieved an overall accuracy of 97% across 266 samples. A dual-dimensional evaluation system was established to achieve a precise balance between positive and negative classification.
[0102] Table 1 Classification report of the classification model for Yangcheng Lake area and non-Yangcheng Lake area
[0103]
[0104] In order to verify the accuracy of this method, 18 pre-reserved blind samples of river crabs were imported into a random forest classification model with 100 decision trees, 11 features, and an accuracy score of 0.9737 to classify and identify their origins. The blind sample verification results of the random forest classification model are shown in Table 2.
[0105] Table 2 Blind sample verification results
[0106]
[0107]
[0108] As shown in Table 2, the random forest classification model achieved an average prediction accuracy score of 0.8794 for 18 river crab samples collected from six locations in Jiangsu Province. Comparing the samples to their actual origins revealed that five predictions were accurate. However, the random forest model identified river crab samples collected from the Yangcheng Lake area as not from the Yangcheng Lake region, which was inconsistent with the actual predictions. The accuracy of the actual validation results was 83.3%. The performance was satisfactory in both the Yangcheng Lake and non-Yangcheng Lake regions, demonstrating the advantages of the random forest algorithm in handling complex nonlinear relationships.
[0109] This study, based on mineral element fingerprint analysis technology and a random forest algorithm, established a classification model for the origin of river crabs at different scales, providing a scientific basis for tracing the origin of Yangcheng Lake hairy crabs. The results showed that the random forest model, based on mineral element fingerprints, can effectively distinguish river crabs from those in the Yangcheng Lake area and those outside the Yangcheng Lake area, with an accuracy rate of up to 97.37%. Based on the 11 characteristic elements screened, a portable XRF detector can be developed for on-site screening and identification by market regulators.
[0110] Example 2
[0111] An embodiment of the present invention provides a method for tracing the origin of an aquatic product. In this embodiment, the aquatic product is Chinese mitten crab, and the origin classification result of the aquatic product is detected as Yangcheng Lake hairy crab or non-Yangcheng Lake hairy crab. The specific processing steps include:
[0112] S31. 904 healthy adult male crabs were collected from the Yangcheng Lake area, surrounding ponds, Taihu Lake, Gaoyou Lake, Gucheng Lake, and Xinghua area. The crabs were dried at 80°C, ground through an 80-mesh sieve, and stored in polyethylene ziplock bags for later use.
[0113] S32. Using an inductively coupled plasma emission spectrometer and an inductively coupled plasma mass spectrometer, perform mineral element detection on the aquatic product sample to obtain the mineral element content of the aquatic product sample;
[0114] S33, using a random forest algorithm to calculate the characteristic importance score of the mineral element content, screening out mineral elements whose importance score is greater than a preset score, and using the mineral elements as core features;
[0115] S34. Establishing a random forest classification model based on the core features;
[0116] According to the random forest algorithm, we first select different numbers of decision trees to import into the random forest classifier, train the model on the training set, and predict the model accuracy score.
[0117] See for example Figure 5 As shown, with 10 decision trees, the model accuracy score is 0.9662. With 50, 100, and 150 decision trees, the model accuracy scores are 0.9474, 0.9511, and 0.9549, respectively, showing a gradually increasing trend. Further increasing the number of decision trees, the model accuracy score decreases when there are 200 decision trees, but remains constant at 0.9436 when the number of decision trees reaches 250. Considering Python's ability to output model accuracy results using test data with different features, to maintain the predictive performance of the classification model, 150 decision trees were selected for import into the random forest classifier and the model was trained on the training set.
[0118] See Figure 6 As shown, all the features in order of importance are Ba, Ca, As, Mn, Zn, Hg, Mg, Na, Sr, Eu, Co, V, Fe, Cu, K, Se, La, Yb, Cr, Pr, Dy, Er, Tm, Ti, Cd, Sm, Nd, Ce, Gd.
[0119] After removing the 15 features K, Se, La, Yb, Cr, Pr, Dy, Er, Tm, Ti, Cd, Sm, Nd, Ce, and Gd with importance scores less than 0.02 from the model, a classification model with 150 decision trees was re-established, and the model accuracy score was 0.9549. At this time, 14 features entered the model, namely Ba, Ca, As, Mn, Zn, Hg, Mg, Na, Sr, Eu, Co, V, Fe, and Cu.
[0120] After removing the 19 features (Co, V, Fe, Cu, K, Se, La, Yb, Cr, Pr, Dy, Er, Tm, Ti, Cd, Sm, Nd, Ce, and Gd) with importance scores less than 0.03 from the model, a new classification model with 150 decision trees was built, achieving an accuracy score of 0.9519. Ten features (Ba, Ca, As, Mn, Zn, Hg, Mg, Na, Sr, and Eu) were then added to the model. It can be seen that after removing the 13 features with importance scores less than 0.02 from the model, and then removing the four features with importance scores between 0.02 and 0.03, the accuracy of the rebuilt output model remained unchanged. Considering the testing cost and workload for future validation work, a random forest classification model with 10 important features (Ba, Ca, As, Mn, Zn, Hg, Mg, Na, Sr, and Eu) was selected.
[0121] S35. Process the aquatic product samples based on the random forest classification model to obtain the source origin classification results of the aquatic product samples that need to be traced, thereby realizing the origin traceability of the aquatic product samples.
[0122] The classification report of the classification model for Yangcheng Lake hairy crabs and non-Yangcheng Lake hairy crabs is shown in Table 3. From the table, we can see that in the model established by the classification method of Yangcheng Lake hairy crabs and non-Yangcheng Lake hairy crabs, the accuracy and F1 score of the Yangcheng Lake hairy crab samples are 0.97 and 0.96 respectively, which are slightly better than those of the non-Yangcheng Lake hairy crab samples. In terms of recall rate, it is slightly lower than that of the non-Yangcheng Lake hairy crabs. In general, the model has a good classification effect.
[0123] Table 3 Classification report of Yangcheng Lake hairy crab and non-Yangcheng Lake hairy crab classification models
[0124]
[0125] In order to verify the accuracy of this method, 18 pre-reserved blind samples of river crabs were imported into a random forest classification model with 150 decision trees, 10 features, and an accuracy score of 0.9519 to classify and identify their origins. The blind sample verification results of the random forest classification model are shown in Table 4.
[0126] Table 4 Blind sample verification results
[0127]
[0128] As shown in Table 4, the 18 crab samples originated from six production areas in Jiangsu Province. The average prediction accuracy score of the random forest classification model was 0.75. Comparison with the actual origins of the samples revealed that five predictions were accurate. Among them, a crab sample collected from the Yangcheng Lake area was identified as a non-Yangcheng Lake hairy crab, inconsistent with the actual predictions. The actual validation accuracy was 83.3%. The classification performance for Yangcheng Lake hairy crabs was slightly better than that for non-Yangcheng Lake hairy crabs, indicating that their mineral element fingerprints have strong regional characteristics, supporting the feasibility of this model for verifying origin.
[0129] This study, based on mineral element fingerprint analysis technology and a random forest algorithm, established a classification model for the origin of river crabs at different scales, providing a scientific basis for tracing the origin of Yangcheng Lake hairy crabs. The results showed that the random forest model, based on mineral element fingerprints, can effectively distinguish river crabs from those in the Yangcheng Lake area and those outside the Yangcheng Lake area, with an accuracy rate of up to 95.19%. Based on the 10 characteristic elements screened, a portable XRF detector can be developed for on-site screening and identification by market regulators.
[0130] The above specific embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to examples, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the scope of the technical solutions of the present invention, and all of these should be included in the scope of the claims of the present invention.
Claims
1. A method for tracing the origin of aquatic products, characterized in that: include: S1. Processing the collected aquatic products to obtain aquatic product samples; S2. Performing mineral element testing on the aquatic product sample to obtain the mineral element content of the aquatic product sample; S3. Calculate the characteristic importance score of the mineral element content using a random forest algorithm, screen out mineral elements whose importance score is greater than a preset score, and use the mineral elements as core features; S4. Establishing a random forest classification model based on the core features; S5. Obtain aquatic product samples that need to be traced to their origin, process the aquatic product samples that need to be traced to their origin based on the random forest classification model, obtain a source origin classification result of the aquatic product samples that need to be traced, and realize origin traceability for the aquatic product samples.
2. The method for tracing the origin of aquatic products according to claim 1, characterized in that: Before establishing the random forest classification model, the optimization of the random forest classification model is also included, including: According to the random forest algorithm, different numbers of decision trees are selected to import the random forest classification model; Train the model and predict the model accuracy score. The test model accuracy scores vary when different numbers of decision trees are selected. Combined with the model accuracy results output by Python using test data with different features, and taking into account the maintenance of the prediction performance of the random forest classification model, a suitable decision tree is selected to import the random forest classification model.
3. The method for tracing the origin of aquatic products according to claim 2, characterized in that: The specific steps for selecting the number of decision trees include: Select different numbers of decision trees and calculate the model accuracy score based on different numbers of decision trees; When the accuracy score of the model shows a gradually increasing trend, continue to increase the number of decision trees; After comprehensive consideration, select an appropriate number of decision trees to import into the random forest classification model.
4. The method for tracing the origin of aquatic products according to claim 1, characterized in that: The specific steps of obtaining the core features include: Rank all mineral element characteristics in order of importance; Remove features in the model whose importance scores are less than a first preset score, rebuild a classification model with a preset number of decision trees, and calculate the model accuracy score; Remove features in the model whose importance scores are less than a second preset score, re-establish a classification model with a preset number of decision trees, and calculate the model accuracy score; The matching model is selected by evaluating the accuracy score of the model, and the mineral element characteristics in the selected model are used as the core characteristics.
5. The method for tracing the origin of aquatic products according to claim 1, characterized in that: The specific steps of the mineral element detection include: S21, weighing the aquatic product sample and placing it in a digestion tank; S22, wetting the sample with ultrapure water, adding nitric acid, covering the sample, and letting it stand for a preset time; S23, heating and digesting the mixture after standing for a preset time; S24. After digestion, open the lid to remove the gas in the tank and transfer the digestion tank to the acid remover for heating; S25, allowing the heated digestion tank to stand and then adding ultrapure water to a fixed volume to obtain a solution to be tested; S26. Using a mineral element detection device to perform mineral element detection on the solution to be detected.
6. The method for tracing the origin of aquatic products according to claim 5, characterized in that: The mineral elements include: any one or more of calcium, copper, magnesium, sodium, lanthanum, zinc, potassium, gadolinium, barium, cerium, mercury, manganese, vanadium, thulium, praseodymium, dysprosium, samarium, arsenic, selenium, europium, cobalt, iron, cadmium, ytterbium, erbium, neodymium, strontium, chromium and titanium.
7. The method for tracing the origin of aquatic products according to claim 5, characterized in that: The mineral element detection equipment adopts: inductively coupled plasma emission spectrometer and inductively coupled plasma mass spectrometer.
8. The method for tracing the origin of aquatic products according to claim 7, characterized in that: The detection conditions of the inductively coupled plasma emission spectrometer are as follows: radio frequency power 1100-1200 W, atomizing gas flow rate 0.65-0.75 L / min, cooling gas flow rate 12.0-13.0 L / min, auxiliary gas flow rate 0.45-0.55 L / min, and peristaltic pump speed 40-50 r / min.
9. The method for tracing the origin of aquatic products according to claim 7, characterized in that: The detection conditions of the inductively coupled plasma mass spectrometer are as follows: RF power 1500-1600 W, cooling gas flow rate 12-16 L / min, auxiliary gas flow rate 0.7-0.9 L / min, nebulizing gas flow rate 0.8-1.2 L / min, helium flow rate 4.2-4.8 mL / min, dwell time 0.01-0.03 s, scan number 95-105 times, injection time 42-47 s, and peristaltic pump speed 40-50 r / min.
10. A system for tracing the origin of aquatic products, characterized in that: include: Detection module: used to obtain aquatic product samples after processing the collected aquatic product samples, and perform mineral element detection on the aquatic product samples to obtain the mineral element content of the aquatic product samples; Data processing module: used to calculate the characteristic importance score of the sample data of the detection module through the random forest algorithm, screen out the mineral elements whose importance score is greater than the preset score, and use the mineral elements as the core features; Classification model module: used to establish a random forest classification model based on the core features of the data processing module; Classification output module: used to import aquatic product samples that need to be traced into the classification model module, and output the source origin classification results of the aquatic product samples that need to be traced.
Citation Information
Patent Citations
Anti-counterfeit tracing method and device for hairy crab
CN108985410A
Methods and systems for tracing the origin of river crabs
CN114241248B
Hairy crab anti-counterfeiting traceability method and system and crab buckle
CN119313357A