Industrialized System and Method for Rice Grain Identification
By adopting the combination of optical capture devices, digital platforms and machine learning structures in the rice particle recognition and classification system, the problem of low efficiency and accuracy of rice particle recognition and classification in the prior art is solved, and more efficient and accurate identification and classification effects are achieved.
Patent Information
- Application Number
- CN202080023318.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-19
- Filing Date
- 2020-03-19
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-03-19
AI Technical Summary
The prior art has problems with efficiency and accuracy in the identification and classification of rice particles, especially when dealing with non-breaking rice particles with small heterochromatic points and very short.
An industrial system is employed that includes an optical capture device, a digital platform and a machine learning structure. The optical capture device captures the optical image and transmits it to the digital platform through a data interface, which is processed by a segmentation module, a feature extractor and a classifier. The classifier uses selectors to select the best machine learning structure and optimizes the classification results through threshold triggers.
Efficient and accurate identification and classification of rice particles are achieved, and non-breaking rice particles with small heterochromatic points and very short, can be processed with higher efficiency and accuracy than the prior art.
Smart Images

Figure CN113826109B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial systems for rice and other grain identification and sorting. In particular, the present invention relates to optical sorting and optimized sorting devices for sorting spoiled rice or other seeds. Background Art
[0002] Optical recognition and sorting is an important technical area in the industrial production of rice, presenting many technical challenges. On the other hand, the importance of rice as a staple food cannot be overestimated. Rice is the most important food crop in developing countries and is the staple food for more than half of the world's population, making it the world's most important human food crop, directly supporting more people than any other crop. In 2012, almost half of the world's population - more than 3 billion people - depended on rice every day. It is also a staple food throughout Asia, where about half of the world's poorest people live, and is becoming increasingly important in Africa and Latin America. Rice also feeds more people for a longer period of time than any other crop. Rice is very diverse both in how it grows and how it is used by humans. Rice is unique because it can grow in humid environments where other crops cannot survive. Such humid environments are abundant throughout Asia. The domestication of rice was one of the most important developments in history, and thousands of rice varieties are currently cultivated on every continent except Antarctica.
[0003] Rice is grown in a wide variety of locations and climates, from the wettest regions of the world to the driest deserts. It is grown from the Rakhine Coast of Myanmar, where average rainfall recorded during the growing season is over 5,100 mm, to the Al-Hasa Oasis of Saudi Arabia, where annual rainfall is less than 100 mm. Temperatures also vary widely. In Upper Sindh, Pakistan, the rice season averages 33°C; in Otaru, Japan, the average temperature during the growing season is 17°C. The crop is grown at sea level in the coastal plains and deltaic regions of Asia, and at elevations of 2,600 m on the hillsides of Nepal. Rice also grows under a very wide range of solar radiation, from 25% potential during the main rice season in parts of Myanmar, Thailand, and Assam, India, to about 95% potential in southern Egypt and Sudan. Rice occupies a very high portion of the total cultivated area in South Asia, Southeast Asia, and East Asia. This region is subject to alternating wet and dry seasonal cycles, and also contains many of the world's major rivers, each with its own extensive delta. Here, large tracts of flat, low-lying farmland are flooded every year during and immediately after the rainy season. The only two staple crops, rice and taro, are easily grown under these saturated soils and high temperatures.
[0004] Two kinds of rice are important cereals for human nutrition: Asian cultivated rice (Oryza sativa) planted all over the world and African cultivated rice (O. glaberrima) planted in parts of West Africa. Only these two cultivars known to cultivated plants belong to a genus that includes about 25 other species, but taxonomy is still a matter of research and debate. Asian cultivated rice is believed to have originated in Malaysia about 14 million years ago. Since then, it has evolved, diversified and dispersed, and wild rice species are now distributed throughout the tropics. Their genomes can be divided into 11 groups labeled AA to LL, and most varieties can be grouped into 4 complexes of closely related varieties in the two main parts of the genus. Only the following two varieties, which are all diploid, have no relationship and are placed in the genus' own part: Australian rice (O. australiensis) and short medicine wild rice (O. brachyantha).
[0005] Throughout the growth process, rice goes through a series of processes before it finally reaches the final consumer. Its production can generally be divided into the following stages: seed selection, land preparation, crop seedling, water management, nutrient management, crop health, harvesting and post-harvest. Seed is a living product that must be grown, harvested and processed correctly to achieve the yield potential of any rice variety. Good quality seeds can increase yields by 5% to 20%. Using good seeds results in lower sowing rates, higher crop emergence rates, fewer replantings, more uniform plant spacing and more vigorous early crop growth. Early vigorous growth reduces weed problems and increases crop resistance to insect pests. All of these factors contribute to higher yields and increased rice farm yields. Good seeds are pure (selected varieties), plump and uniform in size, viable (more than 80% germination rate and good seedling vigor), and do not contain weed seeds, seed-borne diseases, pathogens, insects or other substances. Selecting seeds of the right variety for the environment in which rice is located and ensuring that the seeds of the selected variety are of the highest possible quality is an essential first step in rice production. For land preparation, before planting rice, the soil should be in the best physical condition for crop growth and the soil surface should be flat. Land preparation includes plowing and harrowing to "cultivate" or dig, mix and level the soil. Cultivation enables seeds to be planted at the right depth and also helps control weeds. Farmers can cultivate the land themselves using hoes and other equipment, or they can be assisted by draft animals such as buffaloes or tractors and other machinery. Next, the land is leveled to reduce the amount of water wasted due to uneven, deep puddles or exposed soil. Effective land leveling makes it easier to plant seedlings, reduces the amount of work required to manage crops, and increases grain quality and yield. For crop planting, the two main practices for planting rice plants are transplanting and direct seeding. Transplanting is the most common planting technique in Asia. Pre-sprouted seedlings are transferred from the seedbed to the wet field. It requires less seeds and is an effective way to control weeds, but requires more labor. Seedlings can be transplanted by machine or by hand. Direct seeding involves manually sowing dry seeds or pre-sprouted seeds and seedlings, or planting them by machine. In rainfed and deepwater ecosystems, dry seeds are manually broadcast onto the soil surface and then incorporated by ploughing or by harrowing while the soil is still dry. In irrigated areas, seeds are usually pre-germinated before sowing. With regard to water use and management, cultivated rice is extremely sensitive to water shortages. To ensure adequate moisture, most rice farmers aim to maintain waterlogged conditions in their fields. This is especially true for lowland rice. Good water management for lowland rice focuses on practices that conserve water while ensuring adequate moisture for the crop. In rainfed environments, when optimal water is not available for rice production, a range of options are available to help farmers cope with varying degrees and forms of water shortages. It includes sensible land preparation and pre-planting activities, followed by techniques such as saturated soil cultivation, alternating wetting and drying, seedbeds, mulching and the use of aerobic rice that can cope with dry conditions.Regarding nutrient management, rice plants have specific nutrient requirements at each growth stage. This makes nutrient management a critical aspect of rice cultivation. The unique properties of flooded soils make rice different from any other crop. Because rice fields are flooded for long periods of time, farmers are able to both conserve soil organic matter and receive free inputs of nitrogen from biological sources, which means they require little or no nitrogen fertilizer to maintain yields. However, farmers can tailor nutrient management to the specific conditions of their fields to increase yields. Regarding crop health, rice plants have a wide variety of "natural enemies" in the field. These "natural enemies" include rodents, harmful insects, viruses, diseases, and weeds. Farmers manage weeds through water management and land preparation, hand weeding, and in some cases, the application of herbicides. Understanding the interactions between pests, natural enemies, host plants, other organisms, and the environment allows farmers to determine if any pest management is necessary. Avoiding conditions that allow pests to adapt and thrive in a particular ecosystem helps identify weak links in the pest life cycle and, therefore, what factors can be manipulated to manage them. Maintaining natural ecosystems so that natural enemies and predators of pests and diseases remain abundant can also help reduce pest populations.
[0006] Regarding harvesting, harvesting is the process of collecting mature rice crops from the field. Depending on the variety, rice crops usually reach maturity about 105-150 days after the crop is seeded. Harvesting activities include cutting, stacking, handling, threshing, cleaning and transportation. Good harvesting methods help maximize grain yields and minimize grain damage and deterioration. Harvesting can be done manually or mechanically: (i) manual harvesting is common throughout Asia, and it involves using simple hand tools such as sickles and knives to cut rice crops. Manual harvesting is very effective when crops fall or fall, but it is labor-intensive. Manual harvesting requires 40 to 80 hours per hectare, and it takes additional labor to manually collect and transport the harvested crops; (ii) mechanical harvesting using a harvester or combine harvester is another option, but it is not so common due to the availability and cost of machinery. The rice after cutting must be threshed to separate the grain from the stalks and clean. These processes can also be done by hand or machine.
[0007] After harvesting, rice grains go through many processes depending on how they are used. These processes include drying, storage, grinding, processing and packaging - all before they are delivered to the market for sale. Drying is the process of reducing the moisture content of the grain to a safe storage level. It is the most critical operation after harvesting the rice crop. Delayed drying, incomplete drying or ineffective drying will reduce the quality of the grain and lead to losses. Storing grain is to reduce the loss of grain due to weather, moisture, rodents, birds, insects and microorganisms. Generally, rice should be stored in the form of paddy rice rather than husked rice because the husk provides some protection against insects. In the International Rice Gene Bank, which preserves rice seeds from more than 118,000 different types of rice, rice seeds are stored in vacuum-packed freezers at -18°C, where they can remain viable for 100 years. Rice storage facilities take many forms depending on the amount of grain to be stored, the purpose of the storage and the location of the storage. A good storage system should include: (i) protection from insects, rodents and birds by maintaining proper storage hygiene, (ii) easy loading and unloading, (iii) efficient use of space, (iv) easy maintenance and management, (v) prevention of moisture re-entry into the grains after drying, and (vi) specific solutions to meet the challenges of storing rice in humid tropical regions. Milling is a key step in the post-production of rice. The basic purpose of the rice milling system is to remove the husk and produce edible rice grains that are fully milled and free of impurities. If only the husk is removed, "brown rice" is the product. If the rice is further milled or polished, the bran layer is removed, revealing the "white" rice. Depending on the needs of the consumer, the rice should have minimal broken grains. The rice milling system can be a simple one-step or two-step process, or a multi-stage process. Depending on whether the rice is milled in the village for local consumption or for sale, the rice milling system can be divided into two categories: village rice mills and commercial rice mills. Once milled, the rice is packaged and transported to sales points, which can be local or international.
[0008] Between 1961 and 2010, global rice production more than tripled, with a compound annual growth rate of 2.24% (2.21% for Asian rice production). This increase was slightly higher than that for wheat (2.02% per year), but significantly lower than that for maize (2.71% per year). The main reason for the increase in rice production is higher rice yields, which have an average annual growth rate of 1.74%, while the average annual growth rate of harvested area is 0.49%. In absolute terms, rice production has increased at an average rate of 51.1 kilograms per hectare (kg / ha) per year, although the growth rate has declined in both percentage and absolute terms. Rice is grown by more people than any other crop in the world. There are more than 144 million rice fields worldwide, with an area of about 158 million hectares harvested. It is grown by hand or with large machinery by small families or large agricultural companies in a wide range of climates and terrains. The geographical, economic, and social conditions under which rice is produced vary greatly.
[0009] For rice to be sold on the market, it needs to be correctly sorted to contain as little unwanted material as possible. These unwanted materials can range from factors that affect quality (such as discolored or immature grains) to unacceptable foreign matter that affects food safety (such as sand, stones, glass or mechanical parts). To this end, optical sorting machines have been developed that can typically sort up to 20 metric tons of rice per hour. These machines pass the rice through a chute where it is analyzed as it falls using a high-resolution custom camera and a decision circuit that can detect unwanted materials within milliseconds. Once a defective particle or foreign matter is detected, a small jet of air ejects the defective object out of its track. As a result, only good particles continue to flow in the receiving output bin.
[0010] Of course, this process is not perfect; quality control must be performed regularly to ensure that an acceptable percentage of defective rice is filtered out to meet industry standards. In addition, broken kernels may be considered a second choice and can potentially be sold at a different price than whole kernels. In order to correctly estimate the content of the output from an optical sorter (or from any other processing stage), experts need to take into account the quality, size and state of breakage of each kernel. The sorting, classification and counting processes are still often done manually by experts who, based on their experience and training, can identify even minor kernel defects. This manual quality analysis is time consuming, requires expertise and is highly subjective. Different experts may have different opinions on the quality of rice kernels in terms of color and geometric defects. In addition, the consistency of the opinions of individual experts will be affected by various external factors, such as lighting conditions, current physiological and mental health status or the time available for quality control.
[0011] Frustrated by long and tedious work, quality control experts expressed the need to use smartphone cameras to quickly evaluate samples. This was the core motivation for starting the project to develop the present invention. In the first version of the project, which is currently in production, the user is given a mask in which he / she places a sample of the rice to be analyzed on a plate. The plate also contains a reference sample of good particles, as well as different defect categories to be detected. The user then places the smartphone on top of the mask and takes a photo of the sample. The resulting image is sent to a server to process and calculate the results. Currently, color analysis is used to classify good and defective particles. Broken analysis is performed by simply setting a threshold length to separate intact and broken particles. This method is well suited to detecting color changes in the overall color of the particles, such as heat-damaged particles (they appear more yellow than good particles). However, there are still challenges in classifying rice particles with small heterochromatic spots and very short non-broken rice particles.
[0012] Hai Vu et al., in the December 2018 publication “Inspecting rice seed species purity on a large dataset using geometrical and morphological features” of Information and Communication Technology, disclose a method for checking the purity of rice seed varieties. The method relies on unique feature extraction using both morphological and geometric features extracted from high-resolution RGB images. With the aid of appropriate preprocessing techniques, the collected seeds are normalized by their biological structure so that the geometric features at the local parts of the seeds can be measured accurately enough. The method uses a larger number of varieties, so that the dependence of the classification performance on the similarity of the varieties or the type of extracted features can be analyzed, which would otherwise be impossible using a smaller number of varieties by this method. For the recognition method, supervised classification techniques are used to train positive (correct varieties) and negative (wrong varieties) models. In particular, three recognition methods are used in combination, namely, decision trees (DT), random forests (RF), and adaptive boosting algorithms (Adaboost) that combine the decision trees used as weak learners to improve performance. The outputs of the weak classifiers are combined into a weighted sum representing the final output of the boosted classifier. The method does not wish to distinguish between individual learners based on their strength, since individual learners may be weak. But as long as each performs slightly better than random guessing, it is assumed in the disclosed method that the final model converges to the stronger learner. To train the different recognition methods, morphological features and geometric features are used separately and in combination in a second step to achieve the highest performance.
[0013] Ozabaki et al., November 2015, published a method for classifying rice grains using image processing and machine learning techniques in the document “Classification of Rice Grains Using Image Processing And Machine Learning Techniques” on www.reasearchgate. There are four types of grains for classification, among which broken grains are also considered as a type. Rice grain images are taken by a webcam. Six attributes are extracted for each grain image, where the grain attributes are associated with its shape geometry. The attributes of known types of rice grains are used to train and compare different machine learning algorithms.
[0014] Kevin Arvai's article "Fine tuning aclassifier in scikit-learn-Towards Data Science" on towarddatascience.com in January 2018 discloses a method for optimizing the sensitivity of a classifier, that is, minimizing the false negative rate. With recognition methods, accuracy is usually only one of four relevant metrics, where the other three metrics are accuracy, recall, and F1 score. Each metric measures a difference in the performance of the system. For this reason, it is usually also desirable to optimize one metric and therefore prioritize one metric over another. The metric to be optimized depends on the context and goals of the system. Precision is the ratio of correctly predicted positive observations (true positives) to the total predicted positive observations of the system (i.e. correctly predicted positives (true positives) and incorrectly predicted positives (false positives) generated by the system). Recall is the ratio of correctly predicted positive observations (true positives) to the system generated results of all actual observations (true positives) in that class, i.e. true positives together with false negatives. Accuracy is the ratio of correctly predicted classes (both true positives and true negatives) to the total used dataset. Finally, the F1 score is the weighted average (or harmonic mean) of precision and recall. ). Thus, this score takes into account both false positives and false negatives to achieve a balance between precision and recall. The accuracy % can be disproportionately skewed by a large number of actual negatives. This is true, for example, when there is high asymmetry in the data, so while the precision score, which assesses the ability to identify actual positives, is low, one can have a very high accuracy score of 99% (since accuracy assesses the ability to identify both actual positives and actual negatives). Considering the most relevant metrics strongly depends on the context. Arvai's method looks at "accuracy" and "Recall" is used as its best performance metric because the method is used to identify malignant cancer in the presence of a large number of people without cancer, so it has a very high accuracy score of 99% due to the high asymmetry of the data. Arvai's method discloses using a two-step "Scikit-Learning" method to adjust the classifier used for recall: (A) In the first step, the hyperparameters of the estimator used are adjusted by searching for the best hyperparameters and keeping the classifier with the highest recall score. In the Scikit learning method, hyperparameters represent parameters that are not directly learned within the estimator. In the proposed method, search candidates are sampled for a given value by exhaustively considering all possible parameter combinations through an exhaustive grid search. It should be noted that there are other strategies under the Scikit method to adjust hyperparameters, such as randomized parameter optimization using randomized search on parameters, where each setting is sampled from a distribution of possible parameter values; (B) The second step uses the precision-recall curve and the receiver operating characteristic curve (ROC curve) to adjust the decision threshold of the classifier.The ROC curve gives the diagnostic power of a binary classifier system as its discrimination threshold is varied, i.e. the sensitivity-specificity threshold in the classifier. The basic idea of Arvai's method disclosed for fine-tuning a classifier in Scikit-learn is that it uses the raw probability that a sample is predicted to be in a class. This is an important distinction from the absolute class predictions used by other methods.
[0015] Finally, Cheng Ju et al., “The Relative Performance of Ensemble Methods with Deep Convolutional Neural Networks for Image Classification” published in Cornell University Library in April 2017, discloses the use of an ensemble of artificial neural networks applied to different machine learning tasks including image recognition, semantic segmentation, and machine translation. The method involves multiple ensemble methods for image recognition tasks, including unweighted average, majority voting, Bayesian optimal classifier, and (discrete) super learner, with deep neural networks as candidate algorithms. The proposed algorithm uses the same network structure with different model checkpoints in a single training process, or uses a network with the same structure but randomly trained multiple times, or uses a network with different structures. Due to the overfitting phenomenon of neural networks and its impact on different ensemble methods, the method discloses the use of a super learner method to obtain the best results. Summary of the invention
[0016] It is an object of the present invention to provide an industrialized system and method for rice grain identification that is capable of handling the various effects of numerous sources of grain defects and defect variations as described above. In particular, the industrialized system and method for rice grain identification should overcome the shortcomings inherent in the prior art systems. More particularly, the system of the present invention should not have the shortcomings of the prior art systems that are based solely on the known classification process of good grains and defective grains using color analysis. Even more particularly, the system and method should be able to correctly classify rice grains with small heterochromatic spots and very short non-broken rice grains, operating with greater efficiency and accuracy than the known prior art systems.
[0017] According to the invention, these objects are achieved in particular by the features of the independent claims.In addition, further advantageous embodiments can be derived from the dependent claims and the associated description.
[0018] According to the present invention, the above object relates to an industrialized system for rice grain identification, the system comprising an optical capture device for capturing an optical image, the optical capture device comprising a data interface for transmitting the optical image to a digital platform via a data transmission network, wherein the digital platform analyzes the optical image and provides an appropriate response to the optical capture device and / or a user device, wherein the digital platform comprises a segmentation module for segmenting the captured optical image to obtain grains by identifying optical image segments capturing grains or grain parts, wherein the digital platform comprises a feature extractor for extracting measurable grain features from the identified grains or grain parts in the optical image, the feature extractor comprising a segmentation module for segmenting the captured optical image to obtain grains. Processing features describing different parameterized aspects of particles includes at least shape parameter values and / or color parameter values and / or spatial parameters and / or geometric parameters, wherein the digital platform includes a classifier having a selector for sequentially selecting multiple machine learning structures, the selector applies different machine learning structures to the extracted particle features for rice grain recognition, and selects the best machine learning structure among the applied machine learning structures, wherein the best machine learning structure selected among the applied machine learning structures is further optimized by changing an appropriate threshold with the aid of a threshold trigger, wherein a change in the discrimination threshold changes the diagnostic ability of the binary classifier system, and the diagnostic ability of the binary classifier system changes in correlation with the change in the discrimination threshold. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The invention will be explained in more detail by way of example with reference to the accompanying drawings, in which:
[0020] Figure 1 A block diagram schematically illustrating an exemplary industrialized system 1 for rice grain recognition and classification is shown, including segmentation 110, feature extraction 111, and classification 112, the classification 112 including selecting a preferred machine learning structure 11211, ..., 1121i, and also minimizing the false positive rate (FP).
[0021] Figure 2 A diagram schematically illustrates an exemplary light shield or light box 13 comprising a power supply 132 and a uniform light source 131 for the interior of the light shield 13. The light shield 13 has an aperture 134 that enables the application of an optical sensor / camera 121 / 122 of, for example, a smartphone or other optical image capture device 12.
[0022] Figure 3 A diagram schematically showing discs 14, in particular a training disc 141 and a sample disc 142, is shown, which allow capturing rice grains 4 to be applied to a mask 13 for optical image 123 capture.
[0023] Figure 4 A diagram schematically illustrates an exemplary industrialized system 1 for rice grain identification and classification, wherein a smartphone 12 and a photomask 13 are used to capture an optical image 123 which is transmitted via a data transmission network 2 to a digital platform 11 which responds with an appropriate rice identification analysis displayed on the smartphone 12 .
[0024] Figure 5 A diagram schematically illustrating an exemplary segmentation performed by the segmentation module 110 , which segments the captured optical image 123 to obtain particles 4 by identifying optical image 123 segments capturing particles 4 or particle parts.
[0025] Figure 6 A diagram is shown schematically illustrating an exemplary feature extraction process by a feature extractor 111 for extracting measurable particle features 41 from identified particles 4 or particle portions in an optical image 123, wherein features of different parameterized aspects of the particles 4 described by the feature extraction process 1111 include at least shape parameter values 411 and / or color parameter values 412 and / or spatial parameters 413 and / or geometric parameters 414.
[0026] Figures 7a to 7c , Figures 8a to 8d and Figures 9a to 9d An exemplary response screen of the mobile application of the smartphone 12 when the system 1 identifies a particle is schematically shown. Figure 8a Furthermore, a captured optical image 123 is shown before it is sent to the digital platform 11 via the data transmission network 2 for particle identification. Figure 9d Also shown is a screen that provides a feedback loop 1132 for the iterative retraining process 1131.
[0027] Figures 10 to 16 A diagram schematically illustrating the digital system 11 in an exemplary training mode 114 is shown. Fig.11 A more detailed pre-processing is shown. Fig.12 and Fig.13 Modification of the variable threshold by triggering a desired threshold probability (eg, a maximum of 1% FP) for a particle to be classified as good is shown in more detail. Fig.14A graph schematically illustrating an exemplary summary of experimental results of embodiments of the present invention is shown, wherein (a) the training set is retrained after each experiment; (b) threshold adjustments are applied after each retraining; and (c) the value of each point is: (i) experiment with previously modeled structure and threshold; (ii) accuracy (%) at equal error rate; (iii) defective particles are classified as good with 1% threshold (false positive rate, %); and (iv) error significance based on experiment size = false positive examples (%) * # of particles misclassified. Fig.15 A diagram schematically illustrating an exemplary effect of retraining a learning modeling structure is shown. The extremely positive trend shows the benefit of retraining a classification model with a corrected dataset through the active learning structure of the present invention.
[0028] Fig.16 Peck detection is shown. After retraining: 95% accuracy at equal error rate (see ROC curve).
[0029] Fig.17 A digital platform 11 operating in a forecast 115 for industrialization is schematically shown.
[0030] Fig.18 and Fig.19 A block diagram schematically illustrates an exemplary digital platform 11, including segmentation 110, feature extraction 111, and classification 112, the classification 112 including selecting a preferred machine learning structure 11211, ..., 1121i, and further minimizing the false positive rate (FP). Fig.19 In an embodiment variant, the digital platform 11 includes a processing separation structure, in which the cloud computing platform 11a provides rice grain recognition as software as a service (SaaS), and the machine learning / data mining system 11b provides machine learning and data mining processing through a modular data pipeline structure. In particular, further performance improvements can be achieved by using KNIME (Konstanz Information Miner) for the machine learning / data mining system 11b as a data analysis, reporting and integration structure. The KNIME structure allows the integration of various components for machine learning and data mining through its modular data pipeline concept. The use of JDBC and a graphical user interface allows the assembly of nodes that mix different data sources without programming or with only minimal programming, including pre-processing (ETL: Extraction, Transformation, Loading) for modeling, data analysis and visualization.
[0031] Fig. 20A diagram schematically illustrating an exemplary invention project context for the industrialization of rice grain recognition is shown. Sorting machine A is capable of sorting up to 20 tons / hour and is used for half of the world's primary food supply. Such sorting is usually based on conventional image processing operations, including (i) color discrimination; (ii) depending on the size of the grain; and (iii) involving a large amount of manual settings. The industrialization of rice grain recognition must be able to capture 120,000 rice varieties.
[0032] Fig.21 and Fig. 22 An exemplary light shield or light box 13 is schematically shown, comprising a power supply 132 and a uniform light source 131 for the interior of the light shield 13. The light shield 13 has an aperture 134 that enables the application of an optical sensor / camera 121 / 122 of, for example, a smartphone or other optical image capture device 12. DETAILED DESCRIPTION
[0033] Figures 1 to 22 The architecture of a possible implementation of an embodiment of the method and system 1 of the present invention for rice grain recognition and classification is schematically shown. The industrialized system 1 for rice grain recognition comprises an optical capture device 2 for capturing an optical image 123, the optical capture device 2 comprising a data interface 124 for transmitting the optical image 123 to a digital platform 11 via a data transmission network 2, wherein the digital platform 11 analyzes the optical image 123 and provides an appropriate response to the optical capture device 2 and / or a user device. The response can be transmitted, for example, in an http-request-response process or another request-response process between the digital platform 11 and the optical capture device 2 and / or the user device, wherein the response is transmitted in a manner similar to that of a wireless communication device. Figures 7a to 9d As shown, various results can be displayed to the user 5 on the optical capture device 2 and / or the user device, for example. The optical capture device 2 can be applied to the mask 13 including the uniform light source 131, for example. The optical capture device 2 can be applied to the mask 13 including the uniform light source 131, for example. The optical capture device 2 can be a mobile optical capture device 2, which is applied to the mask 13 from the outside via a camera hole in the mask 13. The optical capture device 2 can be a mobile smartphone, for example, wherein the optical image 123 is captured by the camera 121 of the mobile smartphone using a dedicated mobile application and the optical image 123 is transmitted to the digital platform 11 through the data transmission network 2 via one of the data transmission interfaces 124 of the mobile smartphone.
[0034] The digital platform 11 comprises a segmentation module 110 which segments the captured optical image 123 to obtain particles 4 by identifying segments of the optical image 123 that capture particles 4 or particle parts.
[0035] The digital platform 11 comprises a feature extractor 111 which extracts measurable particle features 41 from the identified particles 4 or particle parts in the optical image 123, the features of different parameterized aspects of the particles 4 described by the feature extraction process 1111 comprising at least shape parameter values 411 and / or color parameter values 412 and / or spatial parameters 413 and / or geometric parameters 414. In particular, the measurable particle features 41 may, for example, include physical properties such as length, width, translucency, degree and / or frequency and / or type of broken particles, color, age, particle weight, hardness, glassiness (particle glassiness is an optical property. Based on glassiness, particles can be divided into 3 major categories: glassy, powdery and mottled. Glassy particles differ from non-glassy particles in particle appearance (starchy and opaque); glassy particles are generally considered to be of better quality than non-glassy particles due to higher protein quality, darker color and coarser uniform particles), moisture content and / or protein content. The extracted features 41 are used as input parameters and input parameter values of the machine learning structure 11211, ..., 1121i (i.e., the selected artificial neural network or machine learning algorithm) during the training mode 114 and / or the prediction mode 115. The feature extractor 111 extracts up to 100 and more different particle features 41. In particular, optically extractable features may be necessary, such as physical particle properties of rice as described above, such as size, shape, color, uniformity and general appearance. Other factors that contribute to the overall appearance of rice, such as cleanliness, lack of other seeds, vitreous, translucent, chalky, color, broken and undesirable grains, may also be used. In order to extract, determine and / or measure particle size parameters, particles may be classified into three main groups, for example, (i) length (ii) shape and (iii) weight, where length is a measure of the rough, brown or crushed particle in its highest dimension, while shape is the ratio of length, width and thickness, and for the case of weight, determined by using 1,000 particle weights. For example, a long particle type can be classified by the feature extractor 111 as having a length of 6.61 mm to 7.5 mm, a shape (ratio) greater than 3, and a weight of 15 mg to 20 mg; a medium particle type can be classified as having a length of 5.51 mm to 6.6 mm, a shape (ratio) of 2.1 to 3, and a weight of 17 mg to 24 mg; and a short particle type can be classified as having a length of up to 5.5 mm, a shape (ratio) of up to 2.1, and a weight of 20 mg to 24 mg.
[0036] Test weight can be another feature 41 measured and extracted by the feature extractor 111. This feature 41 is related to bulk density and can be used as an indicator or measure of the relative amount of foreign matter or immature grains, for example. The average test weight per bushel of brown rice depends on the rice type and region and can be, for example, 45 pounds, etc. Impurities and damaged rice can also be extracted as features 41 of rice grains. In particular, for example, the presence of sand and stones will increase the weight of the grains and damage the rubber when sent to the grinder. Impurities and damaged rice include grain impurities, damaged grains, chalky grains, red rice, broken seeds or grains and odors. The main purpose of rice milling is to remove the outer layer (husk), bran and germ while minimizing the damage to the endosperm. The milling quality of rice can be associated with other extracted features 41 used as input parameters and input parameter values of machine learning structures 11211, ..., 1121i from the feature extractor 111 during the training mode 114 and / or the prediction mode 115. These features can also be measured by using milling yield. The variation in milling yield depends on several factors, such as grain type, variety, chalk, drying and storage conditions, others include environmental conditions and moisture content at harvest. The extracted features 41 associated with milling quality can be measured, for example, by two parameters (i) total yield and (ii) rice yield, and additional parameters such as degree of milling and broken rice can be used, for example, by the feature extractor 111 to measure the milling quality (expressed as a percentage). By definition, milling quality refers to the ability of rice grains to withstand milling and / or polishing without breakage and produce a higher recovery. Another extracted feature 41 can be, for example, the amylose content of rice grains. Similar to other cereals, rice is a good source of starch, especially amylose. It includes more than 80% starch, and at the molecular level, starch contains amylose (linear chain glucose with α(1-4) bonds) and amylopectin (branched glucose with α(1-6) bonds). Based on the amylose content, rice can be classified as waxy 0-2%, very low 2-10%, low 10-20%, medium 20-25% and high 25-32% (rice dry basis). For example, the starch content (amylose) of rice can also be a measure of grain yield, processing and palatability. A further extracted feature 41 can be, for example, the gelatinization temperature. This feature 41 is related to other possible extracted features 41, such as particle size, molecular size of the starch fraction. Like other features 41, it is usually also affected by environmental measurement parameters such as maturity temperature, genetics and rice variety and cooking time. The gelatinization temperature feature 41 is directly related to the amylose content; the higher the amylose, the higher the gelatinization temperature, so that high waxy rice has a higher gelatinization temperature compared to waxy or very low waxy rice. Another extracted feature 41 with the help of the feature extractor 11 can be, for example, the degree of grinding, appearance (color), damaged (broken) and the percentage of chalky grains.Thus, three important property features 41 may be, for example, size, color, and condition (kernel damage), since these properties directly allow for the measurement of quality, milling percentage, and other processing conditions. However, all extractable attribute features 41 may be important, for example, for input to the machine learning structures 11211, ..., 11211i during training mode 114 and / or prediction mode 115, e.g., kernels having a chalky texture are undesirable, since they produce lower milling yields after processing and are prone to breakage during processing.
[0037] The digital platform 11 includes a classifier 112 having a selector 112 for sequentially selecting a plurality of machine learning structures 11211, ..., 1121i, the selector 1122 applying different machine learning structures 11211, ..., 1121i to the extracted grain features 41 for rice grain recognition, and selecting the best machine learning structure from the applied machine learning structures 11211, ..., 1121i. The selectable machine learning structures 11211, ..., 11211i may, for example, include at least one neural network structure. The at least one neural network structure may, for example, include at least one convolutional neural network (CNN). The selectable machine learning structures 11211, ..., 1121i may also include other machine learning structures 11211, ..., 1121i as deep learning or deep structured learning or any other applicable machine learning that produces appropriate results during recognition, wherein the deep learning structure is based on an artificial neural network with representation learning. The learning in the learning mode 114 may, for example, be supervised, semi-supervised or unsupervised. However, the supervised learning mode 114 may be a preferred embodiment variant, in particular with respect to the possible active learning structure of the present invention 113. The deep learning structure 11211, ..., 1121i may, for example, include a deep neural network structure and / or a deep belief network structure and / or a recursive neural network structure and / or the already discussed convolutional neural network structure, etc. In the case of supervised learning, the proposed machine learning structure 11211, ..., 1121i may be trained by an optical data set annotated by a human expert (in particular in the active learning structure 113), which outperforms the overall classification results produced by human assessors and / or rule-based selection. This may, for example, be provided in particular using separated samples, such as using a training disk 141 and a sample disk 142. In the supervised learning mode 114, the extracted features are used as input parameters, and the output may, for example, be a binary classifier system (good / bad) or a more granular classification output involving a more granular classification feature, in particular involving an uncertainty feature regarding the classification performed.
[0038] In order to select the plurality of machine learning structures 11211, ..., 1121i, the selector 1122 of the classifier 112 may, for example, include a sampling process 11222 based on a random sampling process 112221, wherein the selector 112 randomly selects the machine learning structures 11211, ..., 1121i so that each machine learning structure 11211, ..., 1121i has the same probability of being selected during the sampling process. Each of the selected machine learning structures 11211, ..., 11211i may, for example, be trained based on a random forest structure that provides an overall learning process for classifying the particles 4 using the plurality of machine learning structures 11211, ..., 11211i. Each of the selected machine learning structures 11211...11211i can be trained, for example, with the aid of a random forest learner 112222 providing an appropriate random forest predictor 112223, wherein each random forest predictor 112223 is optimized by changing an appropriate threshold 11231 with the aid of a threshold trigger 1123, and wherein a change in the discrimination threshold 11231 changes the diagnostic capability of the binary classifier system 112, which changes in correlation with the change in the discrimination threshold 11231.
[0039] The selected best applied machine learning structure 11211, ..., 1121i is further optimized by changing the appropriate threshold 11231 by means of a threshold trigger 1123, wherein the change of the discrimination threshold 11231 changes the diagnostic ability of the binary classifier system 112, which changes in correlation with the change of the discrimination threshold 11231. The threshold trigger 1123 can, for example, trigger the optimal threshold parameter value 112311 in a two-dimensional optimization process, which measures the true positive rate 112331 and the false positive rate 112332 at various threshold settings, thereby optimizing the sensitivity of the classifier 112 according to its false alarm rate (fallout). The threshold parameter value 112311 can, for example, be optimized by the threshold trigger 1123 in the following way: the false positive rate 112332 is limited to a predefined trigger value. The predetermined trigger value for the false positive rate 112332 can, for example, be >= 1%.
[0040] In addition, the active learning structure 113 can, for example, provide a feedback loop to the user or human expert 5 as an iterative retraining process 1131 based on a confusion matrix 11233 including values of true positives (TP), false negatives (FN), false positives (FP) and true negatives (TN) for the classified rice grains, wherein the system 1 is retrained based on feedback parameters of the feedback loop 1131.
[0041] In an embodiment variant, the digital platform 11 includes a separate processing structure, in which the cloud computing platform 11a provides rice grain recognition as software as a service (SaaS), and the machine learning / data mining system 11b provides machine learning and data mining processing through a modular data pipeline structure. In particular, further performance improvements can be achieved by using KNIME (Constance Information Miner) for the machine learning / data mining system 11b as a data analysis, reporting and integration structure. The KNIME structure allows the integration of various components for machine learning and data mining through its modular data pipeline concept. The use of JDBC and a graphical user interface allows the assembly of nodes mixing different data sources without programming or with only minimal programming, including pre-processing (ETL: Extract, Transform, Load) for modeling, data analysis and visualization.
[0042] As described above, in order to handle the purposes of the present invention, the system of the present invention uses a more advanced classification structure, such as using KNIME (Constance Information Miner). KNIME is a data analysis, reporting and integration structure. KNIME integrates various components for machine learning and data mining through its modular data pipeline concept. The use of JDBC and a graphical user interface allow the assembly of nodes that mix different data sources without programming or with only minimal programming, including preprocessing (ETL: extraction, transformation, loading) for modeling, data analysis and visualization. To some extent, the advanced analysis possibilities of KNIME can be considered as a SAS alternative.
[0043] The use of more advanced classification structures means that more complex structures and processing algorithms need to be developed to determine whether a given particle 4 is good or bad, broken or intact. In the first approach, a neural network specially designed for image recognition is used to determine the quality of the particle. Convolutional neural networks (CNNs) can provide very accurate results in image recognition, but require extremely large data sets to be properly trained. However, the available data sets are generally not large enough to produce promising results with only 85% accuracy. In order to converge to the required accuracy level faster, different machine learning algorithms were tested and their efficiency was compared on the available data sets. In order to do this, features must be extracted from each image to complete the classification. Features 41 are extracted from the image to describe different aspects of the particle: shape 411, color 412, special features 413, geometric features 414, etc.
[0044] An appropriate accuracy metric needs to be defined to accurately compare algorithms. As in many machine learning defect detection applications, missing a defective data point is a more costly error than misclassifying a good data point as a defect. A tool used to understand this misclassification is often called a confusion matrix. 4 numbers are shown: True Positives (TP) 112331, False Negatives (FN) 112334, False Positives (FP) 112332, and True Negatives (TN) 112333. TP represents the number of good grains that were correctly classified as good, while FN represents the number of good grains that were classified as defective. Conversely, TN represents the number of defective grains that were correctly classified as defective, while FP represents the number of defective grains that were classified as good. This last metric is chosen to be as low as possible while keeping the overall accuracy as high as possible.
[0045] Once the algorithm is selected, the FP rate 112332 needs to be further minimized. To achieve this, it is necessary to observe how the FP rate 112332 behaves relative to modifications of the model's threshold. For any given particle 4, the algorithm outputs a pair of values corresponding to the probability that the particle 4 belongs to the good or defective class. It is then up to the system and / or operator to determine the threshold probability for classifying the particle as good. A higher threshold means more selective classification and lower FP 112332. However, due to the increase in the threshold, many false negative classifications occur; some new FN 112334 particles still have a relatively high but below-threshold probability of being good. Therefore, defective particles with a probability of being good exceeding 50% are considered "acceptable" and are not classified in any final class. If all "acceptable" particles are removed from the final classification statistics, the overall FN rate is significantly reduced.
[0046] Under the situation of this classification uncertainty, the next step is to include the possibility of human correction in the present invention's active learning process in the system of the present invention.In fact, although the model has a considerable performance, new problems will soon appear, which will sharply reduce its performance: new varieties of defective rice, different tones of rice in different regions, etc. In order to improve the overall classification, adapt to more variable data sets and reduce the acceptable rate of the model, an iterative retraining process is defined.For example, a web application is developed so that users may upload new images, observe the classification of the model, correct them when needed and share them for cross validation.Once a series of reports have been corrected by a rice expert, they can be parsed together with new rice grain images and the classification model can be retrained using the new data of this update.
[0047] Under the platform of development, rice experts can test the model, correct the output and provide feedback to the behavior of the system. The first few experiments quickly revealed some defects in the preliminary training of the model. In the training set used, several serious performances in various defect types were insufficient, and therefore were not identified as defective particles, causing a large number of FP results. Once the model is retrained, this is quickly corrected. In fact, with more data from these rare categories, the type of rice defects that the model can identify becomes more diverse. As the model becomes more and more familiar with various defect types after each retraining, this experiment with many FPs becomes less frequent.
[0048] The proposed solution provides relevant examples of how machine learning-based computer vision techniques can be applied innovatively and help improve the consistency of industrial quality control. It also reveals how collecting a quality initial dataset (which can be challenging) and how closing the feedback loop with domain experts is necessary to achieve highly accurate predictions.
[0049] Reference List
[0050] 1Industrial system for rice grain identification
[0051] 11Digital Platforms
[0052] 11a Cloud Computing Platform
[0053] 11b Machine learning / data mining platform through modular data pipeline
[0054] 110 Segmentation module (preprocessing)
[0055] 1101 Segmentation Processing
[0056] 111 Feature Extractor (Preprocessing)
[0057] 1111 Feature extraction processing
[0058] 112 Classification module / classifier system (processing)
[0059] 1120 Classification Processing
[0060] 1121Machine-based recognition intelligence
[0061] 11211, ..., 1121i Machine Learning Structure
[0062] 1122 selector
[0063] 11221 Confusion Matrix
[0064] 11222 Sampling process
[0065] 112221 Random Sampling Process
[0066] 112222 Random Forest Learner
[0067] 112223 Random Forest Predictor
[0068] 1123 Threshold Trigger
[0069] 11231 Variable threshold parameters
[0070] 112311threshold(tv 1,…,i )
[0071] 112312Probability of a positive example at a specific tvi11232Threshold probability of a positive example
[0072] 11233 confusion matrix
[0073] 112331True Probability (TP)
[0074] (Probability of detection)
[0075] 112332 False Positive Rate (FP)
[0076] (Probability of false warning)
[0077] 112333 True Negative Rate (TN)
[0078] 112334 False Negative Rate (TN)
[0079] 113 Active Learning Structure
[0080] 1131 Iterative retraining process
[0081] 11311Corrective Adjustment of Optical Recognition
[0082] 11312 Retraining Request
[0083] 1132 Feedback Loop
[0084] 114 Training Mode
[0085] 115 Prediction Model
[0086] 12Optical image capture device
[0087] 121 Camera Device
[0088] 122 optical sensor
[0089] 123 Optical Image
[0090] 124 data interface
[0091] 13 Lampshade
[0092] 131 Uniform light source
[0093] 132 Power Supply
[0094] 133 Optical sensor / camera hole
[0095] 14 plates
[0096] 141 training disc
[0097] 142 sample tray
[0098] 2Data transmission network
[0099] 3. Sorter
[0100] 31 Optical Sensor
[0101] 32 sorting units
[0102] 33 Sorting and cleaning process
[0103] 331 Color
[0104] 332 Damage
[0105] 333 Other materials (sand, rock, glass, ...) 4 Rice grains
[0106] 41 Characteristic parameters of particles
[0107] 411 shape parameters
[0108] 412 Color Parameters
[0109] 413 Space Parameters
[0110] 414Geometric parameters 5User / Human expert.
Claims
1. An industrialized system for rice grain identification, the industrialized system comprising an optical capture device for capturing an optical image, the optical capture device comprising a data interface for transmitting the optical image to a digital platform via a data transmission network, wherein: The digital platform analyzes the optical image and provides a response of rice identification to the optical capture device and / or a user device, wherein the digital platform includes a segmentation module that segments the captured optical image to obtain grains by identifying optical image segments that capture grains or grain portions, and wherein the digital platform includes a feature extractor that extracts measurable grain features from the identified grains or grain portions in the optical image, characterized in that The features processed by feature extraction describe different parameterized aspects of the particles, including at least shape parameter values and color parameter values and / or spatial parameters and / or geometric parameters, In a learning mode, the digital platform includes a classifier having a selector for sequentially selecting a plurality of machine learning structures, the selector applying different machine learning structures to the extracted grain features for rice grain recognition, and selecting an applied machine learning structure among the applied machine learning structures that minimizes the number of grains classified as false positives while maintaining the highest overall accuracy, and The machine learning structure selected in the applied machine learning structure is further optimized by changing the discrimination threshold with the help of a threshold trigger, wherein the false positive examples of classification are further minimized to a certain threshold probability that the particles are classified as good, so that the change of the discrimination threshold changes the diagnostic ability of the binary classifier system, and the diagnostic ability of the binary classifier system is related to the change of the discrimination threshold.
2. The industrial system according to claim 1, characterized in that: The threshold trigger triggers an optimal threshold parameter value in a two-dimensional optimization process that measures true positive rates and false positive rates at various threshold settings to optimize the sensitivity of the classifier based on the false alarm rate of the classifier.
3. The industrial system according to claim 2, characterized in that: The threshold parameter value is optimized by the threshold trigger such that the false positive rate is limited to a predefined trigger value.
4. The industrial system according to claim 3, characterized in that: The predefined trigger value for the false positive rate is ≤ 1%.
5. Industrial system according to one of claims 1 to 4, characterized in that The threshold trigger triggers the best threshold parameter value in a two-dimensional optimization process, which measures the true positive rate and the false positive rate at various threshold settings, thereby optimizing the sensitivity of the classifier according to the false alarm rate of the classifier with the help of a confusion matrix.
6. Industrial system according to one of claims 1 to 4, characterized in that The active learning structure provides a feedback loop to a user or a human expert as an iterative retraining process based on a confusion matrix including values of true positives, false negatives, false positives and true negatives for the classified rice grains, wherein the industrialized system is retrained based on feedback parameters of the feedback loop using a segmentation process and / or a feature extraction process and / or a classification process in a training mode.
7. Industrial system according to one of claims 1 to 4, characterized in that The selectable machine learning structures include at least one neural network structure.
8. The industrial system according to claim 7, characterized in that: The at least one neural network structure includes at least one convolutional neural network.
9. Industrial system according to one of claims 1 to 4, characterized in that The optical capture device is applied to a photomask including a uniform light source.
10. The industrial system according to claim 9, characterized in that: The optical capture device is a mobile optical capture device which is applied to the reticle from the outside via a camera hole in the reticle.
11. Industrial system according to one of claims 1 to 4, characterized in that The optical capture device is a mobile smartphone, wherein the optical image is captured by a camera of the mobile smartphone using a dedicated mobile application, and the optical image is transmitted to the digital platform through the data transmission network via one of the data transmission interfaces of the mobile smartphone.
12. Industrial system according to one of claims 1 to 4, characterized in that The digital platform includes a separate processing structure, in which a cloud computing platform provides software operation services for rice grain recognition, and a machine learning / data mining system provides machine learning and data mining processing through a modular data pipeline structure.
13. Industrial system according to one of claims 1 to 4, characterized in that To select the plurality of machine learning structures, the selector of the classifier includes a sampling process based on a random sampling process, wherein the selector randomly selects the machine learning structures such that each machine learning structure has the same probability of being selected during the sampling process.
14. The industrial system according to claim 13, characterized in that: Each of the selected machine learning structures is trained based on a random forest structure, providing an overall learning process for classifying the particles using the plurality of machine learning structures.
15. The industrial system according to claim 14, characterized in that: Each of the selected machine learning structures is trained by means of a random forest learner providing an appropriate random forest predictor, wherein each random forest predictor is optimized by changing an appropriate threshold by means of the threshold trigger, and wherein a change in the discrimination threshold changes the diagnostic power of the binary classifier system, and the diagnostic power of the binary classifier system changes in correlation with the change in the discrimination threshold.
16. An industrial method for a system for rice grain identification, wherein: An optical image is captured by means of an optical capture device, wherein the optical image is transmitted to a digital platform via a data interface over a data transmission network, and wherein the digital platform analyzes the optical image and provides a response to rice identification to the optical capture device and / or a user device, wherein the captured and transmitted optical image is automatically segmented to obtain grains by identifying optical image segments capturing grains or grain parts by means of a segmentation module, and wherein measurable grain features are extracted from the identified grains or grain parts in the optical image by means of a feature extractor, characterized in that The features processed by feature extraction describe different parameterized aspects of the particles, including at least shape parameter values and color parameter values and / or spatial parameters and / or geometric parameters, In a learning mode, a plurality of machine learning structures are sequentially selected by a selector of a classifier, the selector applies different machine learning structures to the extracted grain features for rice grain recognition, and selects the applied machine learning structure that minimizes the number of grains classified as false positives while maintaining the highest overall accuracy among the applied machine learning structures, and The best machine learning structure selected among the applied machine learning structures is further optimized by changing the appropriate threshold with the aid of a threshold trigger, wherein the false positives of the classification are further minimized to a certain threshold probability that the particles are classified as good, so that the change in the discrimination threshold changes the diagnostic ability of the binary classifier system, and the diagnostic ability of the binary classifier system is related to the change in the discrimination threshold.
Citation Information
Patent Citations
Nondestructive detecting and screening method based on near-infrared for crop single-grain components
CN102179375A
Method for detecting rice transparency
CN103411929A