Fruit and vegetable identification model, training method thereof and training data acquisition method

By collecting data that the fruit and vegetable identification model misidentifies and collecting correct data from the database, and combining feature extraction and similarity calculation, the problem of low data annotation efficiency in the optimization process of the fruit and vegetable identification model was solved, thus achieving efficient model optimization and improved recognition accuracy.

CN121789208APending Publication Date: 2026-04-03SHANGHAI SUMI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing fruit and vegetable recognition models suffer from low data labeling efficiency and resource waste during the optimization process. In particular, a large amount of correctly identified data is repeatedly labeled, resulting in low model optimization efficiency.

Method used

Data that is incorrectly identified by the fruit and vegetable recognition model and data that is correctly identified in the database are collected as training data. The actual labels are determined by feature extraction and similarity calculation, which reduces the amount of data and improves the labeling efficiency.

Benefits of technology

It improved data annotation efficiency, enhanced the targeting and accuracy of model optimization, and significantly improved the recognition accuracy of the fruit and vegetable identification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789208A_ABST
    Figure CN121789208A_ABST
Patent Text Reader

Abstract

The invention provides a fruit and vegetable identification model and a training method and a training data acquisition method thereof, and the training data acquisition method comprises the steps: collecting sample data which comprises at least one piece of first data and at least one piece of second data and is provided with a first label; performing feature extraction on each piece of first training data to obtain each first feature vector; performing feature extraction on each piece of sample data to obtain each second feature vector; for each second feature vector, performing similarity calculation with all the first feature vectors, determining the first N highest similarities, and determining the labels of the N first training data corresponding to the N highest similarities as N second labels of the sample data; and determining an actual label of the sample data based on the first label of the sample data and the N second labels, and taking the sample data and the actual label thereof as second training data. According to the method, the data labeling efficiency is improved, the model is effectively trained, and the training effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention mainly relates to the field of artificial intelligence model training technology, and in particular to a fruit and vegetable recognition model and its training method and training data acquisition method. Background Technology

[0002] With the rapid development of artificial intelligence (AI) technology, various application scenarios are constantly expanding, profoundly changing social production and lifestyles, from autonomous driving and intelligent medical diagnosis to personalized content recommendation. In the field of recognition technology, AI applications have covered multiple mature areas, from high-security facial recognition to product recognition in e-commerce scenarios. These technologies mostly focus on the recognition of objects with obvious features, standardized shapes, and relatively static characteristics.

[0003] The fruit and vegetable recognition model is a computer vision model specifically designed for identifying fruit and vegetable categories, falling under the category of image classification and object recognition. This model is trained using a large amount of labeled fruit and vegetable image data, employing convolutional neural networks (CNNs) or their lightweight variants, to transform the recognition process from one reliant on human experience to one that is automated and standardized. Its key advantage lies in its millisecond-level recognition speed, which significantly reduces customer waiting time, decreases reliance on skilled employees, and effectively saves labor costs.

[0004] However, after the fruit and vegetable recognition model is put into use, in order to continuously improve the recognition effect, the image data generated during the recognition process is usually uploaded to a data processing center, where algorithm engineers further clean, label, and use it for model iteration and optimization. However, since not all image data contributes positively to model optimization, this method often suffers from low efficiency in practical applications. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a fruit and vegetable identification model and its training method and training data collection method, which is conducive to improving the data annotation efficiency, effectively training the model and improving the training effect.

[0006] To address the aforementioned technical problems, in a first aspect, the present invention provides a method for collecting training data for a fruit and vegetable recognition model, comprising: collecting sample data, wherein the sample data includes at least one first data and at least one second data; the first data is data that was not correctly recognized after being identified by a first fruit and vegetable recognition model, and the second data is data from a database, with one data sample collected from each category; the sample data has a first label, wherein the first label is a label input by the user when the first fruit and vegetable recognition model is identified; performing feature extraction on each first training data to obtain each first feature vector, wherein the first training data is the training data used to train the first fruit and vegetable recognition model; performing feature extraction on each sample data to obtain each second feature vector; for each second feature vector, performing similarity calculation with all the first feature vectors, determining the top N highest similarities, and determining the labels of the N first training data corresponding to these N highest similarities as the N second labels of the sample data; determining the actual label of the sample data based on the first label and the N second labels, and using the sample data and its actual label as second training data.

[0007] Optionally, the training data acquisition method further includes: determining whether the first label of the sample data is one of the N second labels.

[0008] Optionally, when the first label of the sample data is one of the N second labels, determining the actual label of the sample data based on the first label of the sample data and the N second labels includes: selecting one of the N second labels as the actual label of the sample data.

[0009] Optionally, when the first label of the sample data is different from all N second labels, the training data acquisition method further includes: unifying the first label of the sample data and the N second labels into a third label; the step of determining the actual label of the sample data based on the first label of the sample data and the N second labels includes: selecting the third label as the actual label of the sample data, or having the user input a new label as the actual label of the sample data.

[0010] Optionally, the training data acquisition method further includes: in the step of determining the top N highest similarities, N=3.

[0011] Optionally, the training data acquisition method further includes uploading the sample data to the cloud.

[0012] Optionally, the training data acquisition method further includes: combining the first training dataset and the second training dataset to form a third training dataset, wherein the third training dataset is used to train the first fruit and vegetable recognition model, and wherein the first training dataset is a set of the first training data, and the second training dataset is a set of the second training data.

[0013] Optionally, in the step of collecting sample data, the collection of sample data is completed when the number of the first data collected reaches a set number.

[0014] Secondly, the present invention provides a training method for a fruit and vegetable recognition model, used to train a first fruit and vegetable recognition model as described in the first aspect, comprising: collecting training data to form a training dataset; training the first fruit and vegetable recognition model based on the training dataset to obtain a second fruit and vegetable recognition model, wherein the data in the training dataset includes the first training data and the second training data as described in the first aspect.

[0015] Secondly, the present invention provides a fruit and vegetable recognition model for classifying fruits or vegetables, which is trained by the training method of the fruit and vegetable recognition model as described in the second aspect.

[0016] Compared with existing technologies, this invention has the following advantages: The sample data used to optimize the fruit and vegetable recognition model comes from two sources. First, there is data that was not correctly identified during the model's recognition process; this data, because it was not correctly identified by the model, has significant value for optimizing the model. Second, there is existing data in the database, including data that the fruit and vegetable recognition model can correctly identify, which forms the foundation for the model's correct recognition. Furthermore, during model optimization, only one data point from each category is used, significantly reducing the number of sample data. Therefore, the training data collected in this invention retains the more valuable, incorrectly identified data, reducing the amount of data that needs labeling, improving data labeling efficiency, and enhancing the effectiveness of subsequent model training. Attached Figure Description

[0017] The accompanying drawings are included to provide a further understanding of this application; they are incorporated into and constitute a part of this application. The drawings illustrate embodiments of this application and, together with this specification, serve to explain the principles of this application. In the drawings:

[0018] Figure 1 This is a flowchart illustrating a method for collecting training data for a fruit and vegetable recognition model according to an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of the standard process for identification using a fruit and vegetable recognition model;

[0020] Figure 3This is a schematic diagram of the process of result recognition model recognition in one embodiment of the present invention;

[0021] Figure 4 This is a flowchart illustrating the training data acquisition method for a fruit and vegetable recognition model according to another embodiment of the present invention.

[0022] Figure 5 This is a flowchart illustrating the training data acquisition method for a fruit and vegetable recognition model according to another embodiment of the present invention. Detailed Implementation

[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this application. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0024] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0025] Furthermore, it should be noted that the use of terms such as "first" and "second" to define components is merely for the purpose of distinguishing the corresponding components. Unless otherwise stated, these terms have no special meaning and therefore should not be construed as limiting the scope of protection of this application. In addition, although the terminology used in this application is selected from commonly known and used terms, some terms mentioned in this application's specification may have been chosen by the applicant according to his or her judgment, and their detailed meanings are explained in the relevant sections of this description. Moreover, this application should be understood not only through the actual terms used, but also through the meaning implied by each term.

[0026] Flowcharts are used in this application to illustrate the operations performed according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more steps may be removed from them.

[0027] Currently, after fruit and vegetable recognition models are put into practical use, to improve their recognition performance, images / data collected during the application process are typically uploaded to the backend for data cleaning and labeling by algorithm engineers. The fruit and vegetable recognition model is then optimized based on this new data. However, this optimization process often lacks data selection and differentiation, using all collected data indiscriminately for model training. Since not all data contributes positively to model optimization, this approach not only increases unnecessary data labeling workload but also reduces the specificity and efficiency of model optimization, resulting in insignificant optimization effects.

[0028] As fruit and vegetable recognition models are used on a large scale in stores, the amount of data generated has increased dramatically, placing a huge burden on storage and manual annotation. For example, if a single cash register scale processes 100 recognitions per day, a store with 5 devices will generate 500 recognition images daily, accumulating to 15,000 images per month. This massive amount of data not only continuously consumes storage resources, but also, since approximately 80% of the images have already been correctly identified by the model, their value for model optimization is limited. Blindly uploading and annotating all images leads to a double waste of resources: on the one hand, it consumes a large amount of storage space, and on the other hand, it unnecessarily increases the workload of annotation personnel. Faced with massive amounts of data, annotation efficiency is far lower than the data growth rate. This not only squeezes labor costs, but also, due to the lack of a data filtering mechanism, many correctly identified samples are repeatedly annotated, potentially introducing inconsistencies in annotation. For example, differences in the names of the same fruit and vegetable category used by different annotators can easily introduce noisy data, ultimately weakening the effectiveness of model optimization. It is evident that the full upload and indiscriminate annotation modes not only exacerbate storage pressure, but also, due to the lack of data validity screening, result in weak targeting of model optimization, low resource utilization, and overall optimization results that are difficult to achieve as expected.

[0029] To address the shortcomings of existing technologies, this embodiment provides a method for collecting training data for a fruit and vegetable recognition model, referring to... Figure 1As shown, method 100 mainly includes: S110, collecting sample data, the sample data including at least one first data and at least one second data; the first data is data that was not correctly identified after being identified by the first fruit and vegetable identification model, and the second data is data in the database, with one data sample collected from each category; the sample data has a first label, the first label being the label input by the user when the first fruit and vegetable identification model identifies the data; S120, performing feature extraction on each first training data to obtain each first feature vector, wherein the first training data is the training data used to train the first fruit and vegetable identification model; S130, performing feature extraction on each sample data to obtain each second feature vector; S140, for each second feature vector, performing similarity calculation with all the first feature vectors, determining the top N highest similarities, and using the labels of the N first training data corresponding to these N highest similarities as the N second labels of the sample data; S150, based on the first label and the N second labels of the sample data, determining the actual label of the sample data, and using the sample data and its actual label as the second training data.

[0030] In the process of identifying fruit and vegetable categories using the fruit and vegetable recognition model, reference is made to... Figure 2 As shown, the fruit and vegetable recognition model is first used to identify the fruits and vegetables. Then, user input is provided, including confirmation of correct identification and correction of incorrect results. Next, the model learns from the identified fruit and vegetable images. It's important to note that during this learning process, it doesn't distinguish between correctly identified and incorrectly identified images (or images that were not correctly identified). Finally, all fruit and vegetable images are identified. To improve image annotation efficiency and fully utilize the important value of incorrectly identified images for model optimization, this embodiment requires collecting incorrectly identified images. (See reference...) Figure 3 As shown, the main difference between this model and conventional fruit and vegetable recognition models lies in the fact that after user input, it needs to determine whether the recognition result matches the user's input. That is, it determines whether the user confirms or corrects the input. A corrected input indicates that the fruit and vegetable recognition model has made an error in recognizing the image. For erroneous recognition results, these data (the first set of data) are collected and stored to form new training data later.

[0031] Furthermore, this embodiment also requires the collection of second data, which is data from a database, with one sample collected from each data category. This second data is existing data in the database, including data that the fruit and vegetable recognition model can correctly identify. It forms the foundation for the model's ability to correctly identify data. During model optimization, only one sample from each data category is used, significantly reducing the number of sample data. The data collected in this embodiment includes both first and second data. These data have corresponding labels to indicate their specific type; these labels are the labels entered by the user during the first fruit and vegetable recognition model's identification process. It can be understood that if the fruit and vegetable image is correctly identified, the label of the data is the label identified by the fruit and vegetable recognition model; if the fruit and vegetable image is incorrectly identified, the label of the data is the label entered by the user.

[0032] After data collection, it is also necessary to determine the corresponding labels for the data. Correct labels are more conducive to the optimization of the fruit and vegetable recognition model. In this embodiment, each collected sample data has its own label (first label). Then, by comparing the first training data with the sample data, several other labels with high similarity to its own label (second labels) are obtained. Based on these labels, the actual label of the sample data is determined, which makes the label corresponding to the sample data more accurate. In one implementation, the collection method can use cosine similarity calculation to calculate the similarity between the feature vector of the sample data and the feature vector of the first training data. That is, calculate the cosine value of the angle between the two feature vectors. The more similar the two feature vectors are, the closer the angle is to 0, and the closer the cosine value is to 1.

[0033] Finally, based on the sample data's own labels and the selected similar labels, the actual labels of the sample data are determined, and the sample data and its actual labels are used as the second training data. In this embodiment, the sample data can be labeled by algorithm engineers, selecting a label from the sample data's own labels and similar labels, or the actual labels of the sample data can be determined by a large language model. This process is automated, greatly saving manpower. Of course, when using a large language model to determine the actual labels of the sample data, algorithm engineers can also perform label verification. For algorithm engineers labeling the sample data, to further improve labeling efficiency, labeling tools can be used. When labeling, the labeler can directly choose to retain the user-input labels without performing any operations or select labels from the training set.

[0034] The training data acquisition method in this embodiment includes data that was not correctly identified during the fruit and vegetable identification model's recognition process. This data, because it was not correctly identified by the model, is of significant value for optimizing the model. It also includes existing data in the database, which includes data that the fruit and vegetable identification model can correctly identify; this data forms the foundation for the model's correct identification. Furthermore, only one data point from each category is used during model optimization, significantly reducing the number of sample data and improving data labeling efficiency. Simultaneously, the labels on the sample data are re-labeled to improve label accuracy, enabling more effective optimization and training of the fruit and vegetable identification model.

[0035] In another embodiment, reference Figure 4 As shown, method 200 mainly includes: S110, collecting sample data, the sample data including at least one first data and at least one second data; the first data is data that was not correctly identified after being identified by the first fruit and vegetable identification model, and the second data is data from the database, with one data sample collected from each category; the sample data has a first label, the first label being a label input by the user when the first fruit and vegetable identification model identifies the fruit and vegetable; S120, extracting features from each first training data to obtain each first feature vector, wherein the first training data is the training data used to train the first fruit and vegetable identification model; S130, S140. For each of the sample data, feature extraction is performed to obtain each second feature vector; S210. For each second feature vector, similarity calculation is performed with all the first feature vectors to determine the top N highest similarities, and the labels of the N first training data corresponding to these N highest similarities are determined as the N second labels of the sample data; S220. It is determined whether the first label of the sample data is one of the N second labels; S220. If the first label of the sample data is one of the N second labels, one of the N second labels is selected as the actual label of the sample data.

[0036] In this embodiment, if the first label of the sample data is one of N second labels, it indicates that the probability of the sample data being labeled with a label other than the selected second label is relatively small. Based on this, the method of this embodiment can use one of the labels obtainable from the first training data as the actual label of the sample data, resulting in high accuracy.

[0037] In other embodiments, reference is made to... Figure 5As shown, method 300 mainly includes: S110, collecting sample data, the sample data including at least one first data and at least one second data; the first data is data that was not correctly identified after being identified by the first fruit and vegetable identification model, and the second data is data from the database, with one data sample collected from each category; the sample data has a first label, the first label being the label input by the user when the first fruit and vegetable identification model identifies the fruit and vegetable; S120, performing feature extraction on each first training data to obtain each first feature vector, wherein the first training data is the training data used to train the first fruit and vegetable identification model; S130, performing feature extraction on each of the sample data to obtain each second feature vector; S 140. For each second feature vector, perform similarity calculation with all first feature vectors, determine the top N highest similarities, and determine the labels of the N first training data corresponding to these N highest similarities as the N second labels of the sample data; S210. Determine whether the first label of the sample data is one of the N second labels; S310. If the first label of the sample data is different from all N second labels, unify the first label of the sample data and the N second labels into a third label; S320. Select the third label as the actual label of the sample data, or have the user input a new label as the actual label of the sample data.

[0038] In this embodiment, when the first label of the sample data is different from all N second labels, it indicates that the label of the sample data may be other labels besides the first and second labels. Therefore, it is necessary to unify and merge the first and second labels and to accurately label them, making the labeling of the sample data more accurate. In this embodiment, a large language model can be used to deduplicate and merge the first and second labels, such as the Deepseek language model, which will not be elaborated further here. For example, for a sample data "yam," if its first label is "iron stick yam," but according to the method of this embodiment, its second label could be "rusty yam," "iron root yam," or "iron stem yam," these labels can be unified into one of them, such as "iron stick yam," or another label.

[0039] In one example, N=3 can be set in the step of determining the top N highest similarity scores. Generally, the number of candidate similar labels should not be too large to avoid the labels being too mixed, affecting the consistency of the data and the clarity of subsequent model training. For example, the currently collected sample data is "Red Fuji" apples. By calculating its feature vector and performing similarity calculation with the feature vector of the first training data, the three most similar categories are found to be "Red Fuji", "Gala" and "Red Delicious".

[0040] In one example, the training data acquisition method also includes uploading sample data to the cloud. The cloud, with its powerful computing capabilities, especially GPU cluster resources, can support the training and iterative optimization of more complex fruit and vegetable recognition models. While the amount of data collected from a single device is insufficient to train a high-performance recognition model, aggregating sample data from tens of thousands of devices enables the training of a more accurate and generalized fruit and vegetable recognition model.

[0041] In one example, the training data acquisition method further includes combining the first training dataset and the second training dataset to form a third training dataset. This third training dataset is used to train the first fruit and vegetable recognition model. The first training dataset is a collection of the first training data, and the second training dataset is a collection of the second training data. By expanding the training dataset, this expanded dataset not only contains the basic data to maintain the original correct recognition ability of the fruit and vegetable recognition model, but also adds incremental data for iterative optimization of the model, thereby giving the fruit and vegetable recognition model a stronger ability for continuous optimization.

[0042] In one example, during the sample data collection step, the current sample data collection is completed when the number of collected initial data reaches a set threshold. This means that by setting a threshold for the number of erroneous data to be identified, the fruit and vegetable recognition model can be updated promptly, and the number of data to be labeled can be effectively reduced when the data volume is small, thereby improving labeling efficiency. For example, the number of erroneous images to be collected each time can be set to 50. When collecting sample data, data whose recognition results from the fruit and vegetable recognition model do not match the user's actual input labels are selected as erroneous data. When the cumulative number of erroneous images reaches 50, an image upload and saving is triggered. Simultaneously, a representative image from each of the learned fruit and vegetable categories is selected and uploaded and saved together. Therefore, each uploaded and saved image will include 50 erroneous images, one image from each category of the user's actual input labels and the corresponding labels from the learning database.

[0043] The training data acquisition method provided in this embodiment aims to solve the data mining problem of fruit and vegetable recognition models. It addresses how to efficiently filter and label data that benefits model optimization from massive amounts of data, thereby improving model performance at the lowest cost and saving on data labeling investment. This embodiment can significantly improve labeling efficiency through pre-labeling by personnel and / or intelligent filtering of labels using a large language model. The method retains data with identification errors in the fruit and vegetable recognition model for subsequent model training, which not only reduces the workload of labeling but also enhances the data's relevance to model optimization, thus achieving significant improvements in both labeling efficiency and model optimization iteration effects.

[0044] The following shows the improvement in the recognition accuracy of the fruit and vegetable recognition model after the data collected using the training data acquisition method of this embodiment was applied to the optimization of the model. The obtained fruit and vegetable recognition model was tested on a test set, and the recognition accuracy improved by 1%-8% on multiple test sets. Specific results are shown in Table 1. "Test Group Data 1" includes 60 categories, with 4 images per category for learning and 16 images for testing; "Test Group Data 2" includes 30 categories, with 4 images per category for learning and 150 images for testing (100 unpackaged and 50 packaged); "Test Group Data 3" includes 64 categories, with 4 images per category for learning and 120 images for testing; and the "Self-collected Test Set" includes 91 categories, with 100 images per category for testing. A total of 216 non-fruit and vegetable images were included.

[0045] Table 1 Test results of the accuracy of the fruit and vegetable recognition model

[0046]

[0047] The training data acquisition method provided in this embodiment is not only applicable to fruit and vegetable recognition scenarios, but can also be extended to other model optimization tasks with similar technical architectures. Typical applications include product recognition and face recognition. Since the functional implementation of these tasks relies on the storage and comparison mechanism of feature vectors, their recognition logic is consistent with that of fruit and vegetable recognition. Therefore, the training data acquisition method provided in this embodiment can also effectively improve the annotation efficiency and optimization effect of similar models.

[0048] Another embodiment of the present invention provides a training method for a fruit and vegetable recognition model, used to train the first fruit and vegetable recognition model as described in the foregoing embodiment. The method mainly includes: collecting training data to form a training dataset; and training the first fruit and vegetable recognition model based on the training dataset to obtain a second fruit and vegetable recognition model. The data in the training dataset includes the first training data and the second training data as described in the foregoing embodiment. It is understood that the specific training and optimization process of the fruit and vegetable recognition model can be based on existing, conventional training and optimization processes. This embodiment does not impose any particular limitations on the training process and will not elaborate further here.

[0049] As can be seen from the descriptions in the foregoing embodiments, the training data used in the model training method in this embodiment includes not only the original training data, but also extended training data. This extended training data mainly includes data that was not correctly identified during the fruit and vegetable identification model identification process. Since the data was not correctly identified by the model, it has great value for optimizing the identification model, thereby effectively training the model and improving the training effect.

[0050] Another embodiment of the present invention provides a fruit and vegetable recognition model for classifying fruits or vegetables, which is trained by the training method of the fruit and vegetable recognition model mentioned in the foregoing embodiments.

[0051] The basic concepts have been described above. Obviously, for those skilled in the art, the above disclosure is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore remain within the spirit and scope of the exemplary embodiments of this application.

[0052] Some aspects of this application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The aforementioned hardware or software may be referred to as a "data block," "module," "engine," "unit," "component," or "system." The processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, or combinations thereof. Furthermore, aspects of this application may manifest as computer products residing in one or more computer-readable media, including computer-readable program code. For example, computer-readable media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tapes, etc.), optical discs (e.g., compressed CDs, digital multifunction DVDs, etc.), smart cards, and flash memory devices (e.g., cards, sticks, key drives, etc.).

[0053] A computer-readable medium may contain a propagated data signal containing computer program code, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and so on, or suitable combinations thereof. A computer-readable medium can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer-readable medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, radio frequency signals, or similar media, or any combination of the above media.

[0054] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of scope in some embodiments of this application are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0055] Although this application has been described with reference to specific embodiments, those skilled in the art should recognize that the above embodiments are only used to illustrate this application, and various equivalent changes or substitutions can be made without departing from the spirit of this application. Therefore, any changes or modifications to the above embodiments within the essential spirit of this application will fall within the scope of the claims of this application.

Claims

1. A method for collecting training data for a fruit and vegetable recognition model, characterized in that, include: Collect sample data, which includes at least one first data and at least one second data; the first data is data that was not correctly identified after being identified by the first fruit and vegetable identification model, and the second data is data from the database, with one data sample collected from each category; the sample data has a first label, which is a label input by the user when the first fruit and vegetable identification model identifies the fruit and vegetable. Feature extraction is performed on each first training data to obtain each first feature vector, wherein the first training data is the training data used to train the first fruit and vegetable recognition model. Feature extraction is performed on each of the sample data to obtain each second feature vector; For each second feature vector, a similarity calculation is performed with all the first feature vectors to determine the top N highest similarities. The labels of the N first training data corresponding to these N highest similarities are then determined as the N second labels of the sample data. Based on the first label and N second labels of the sample data, the actual label of the sample data is determined, and the sample data and its actual label are used as the second training data.

2. The method for collecting training data for the fruit and vegetable recognition model as described in claim 1, characterized in that, Also includes: Determine whether the first label of the sample data is one of the N second labels.

3. The method for collecting training data for the fruit and vegetable recognition model as described in claim 2, characterized in that, When the first label of the sample data is one of N second labels, The step of determining the actual label of the sample data based on the first label and N second labels includes: selecting one of the N second labels as the actual label of the sample data.

4. The training data acquisition method for the fruit and vegetable recognition model as described in claim 2, characterized in that, In the case where the first label of the sample data is different from all N second labels, the method further includes: The first label of the sample data and the N second labels are unified into a third label; The step of determining the actual label of the sample data based on the first label and N second labels includes: selecting the third label as the actual label of the sample data, or having the user input a new label as the actual label of the sample data.

5. The method for collecting training data for the fruit and vegetable recognition model as described in claim 1, characterized in that, include: In the step of determining the top N highest similarities, N=3.

6. The method for collecting training data for the fruit and vegetable recognition model as described in claim 1, characterized in that, Also includes: The sample data was uploaded to the cloud.

7. The method for collecting training data for the fruit and vegetable recognition model as described in claim 1, characterized in that, Also includes: The first training dataset and the second training dataset are combined to form a third training dataset, which is used to train the first fruit and vegetable recognition model. The first training dataset is a set of the first training data, and the second training dataset is a set of the second training data.

8. The method for collecting training data for the fruit and vegetable recognition model as described in claim 1, characterized in that, In the step of collecting sample data, the collection of sample data is completed when the number of the first data collected reaches a set number.

9. A method for training a fruit and vegetable recognition model, used to train a first fruit and vegetable recognition model as described in any one of claims 1-8, characterized in that, include: Collect training data to form a training dataset; The first fruit and vegetable recognition model is trained based on the training dataset to obtain the second fruit and vegetable recognition model, wherein the data in the training dataset includes the first training data and the second training data as described in claims 1-8.

10. A fruit and vegetable recognition model for classifying fruits or vegetables, characterized in that, It is trained using the training method of the fruit and vegetable recognition model as described in claim 9.