A grain defective product detection method combining fingerprint and finite element analysis
By combining fingerprint mapping with finite element analysis, a grain quality knowledge graph was constructed, which solved the problem of the single existing detection method, realized intelligent detection and early control of defective grain products, and improved detection efficiency and accuracy.
Patent Information
- Application Number
- CN202510934010.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-08
AI Technical Summary
The existing grain quality testing methods are single and cannot fully reveal the dynamic correlation and coupling mechanism of internal and external quality changes of grain, resulting in delayed test results and inability to achieve early detection and control of the defective product formation process.
Combining fingerprint mapping with finite element analysis, a grain quality knowledge graph is constructed through feature extraction, dimensionality reduction, cross-modal mutual information analysis and feature sparsity evaluation to achieve intelligent detection of defective grains.
It has achieved accurate identification and dynamic analysis of the formation process of defective grain products, improved the intelligence level and efficiency of detection, provided a basis for scientific decision-making, reduced grain losses, and improved the quality assurance level of the grain industry chain.
Smart Images

Figure CN120448799B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of grain quality detection, and more particularly to a grain defective product detection method combining fingerprint spectrum and finite element analysis. BACKGROUND
[0002] With the development of large-scale and intensive grain storage and transportation system, grain is affected by various environmental factors in the process of harvesting, storage and transportation, which easily leads to quality deterioration and defective product generation. The formation of defective grain not only affects the effective use of grain resources, but also may bring food safety risks and threaten the sustainable development of the grain industry chain. How to effectively identify and control the generation of defective grain has become an important technical problem in grain storage and transportation management and quality control.
[0003] Existing grain quality detection methods mostly rely on single modal detection technology, such as chemical quality detection based on spectral composition analysis or appearance screening based on visual detection and image processing. The single detection means lacks comprehensive analysis ability for internal and external quality changes of grain, cannot fully reveal the dynamic correlation and coupling mechanism between different quality indicators, cannot realize early detection and control of the formation process of defective products, and the detection results often lag behind the grain quality change process, leading to delayed decision-making and increased risk. Therefore, the present application proposes a grain defective product detection method combining fingerprint spectrum and finite element analysis to solve the above problems. SUMMARY
[0004] To achieve the above purpose, the present application provides the following technical scheme:
[0005] A grain defective product detection method combining fingerprint spectrum and finite element analysis, comprising the following steps:
[0006] Performing fingerprint spectrum feature extraction to obtain a first feature set, and extracting features of a finite element model to obtain a second feature set;
[0007] Performing feature dimension reduction on the first feature set and the second feature set respectively, and eliminating redundant variables to obtain a chemical modal feature set and a physical modal set;
[0008] Taking defective grain grade as a target label, calculating the mutual information between each feature in the chemical modal feature set and the target label to perform first layer dimension reduction and obtain a chemical modal vector, and calculating the mutual information between each feature in the physical modal set and the target label to perform first layer dimension reduction and obtain a physical modal vector;
[0009] Under the condition that the target label is known, cross-modal mutual information analysis is performed to obtain a strong dependency feature group, while feature sparsity evaluation and feature stability evaluation are performed, and based on the evaluation results, a way of automatically screening key features of substandard grains is selected and the strong dependency feature group is processed, and finally a grain quality knowledge graph with grain quality and physical-chemical indicators as core nodes is constructed for detecting substandard grains.
[0010] In a preferred embodiment, principal component analysis is used for feature dimension reduction.
[0011] In a preferred embodiment, the first layer dimension reduction refers to:
[0012] A standard threshold is set, and the mutual information between each feature in the chemical modal feature set and the target label, and the mutual information between each feature in the physical modal set and the target label, are compared with the standard threshold respectively, and features below the standard threshold are removed, and features not below the standard threshold are retained.
[0013] In a preferred embodiment, cross-modal mutual information analysis refers to:
[0014] The mutual information value between each feature in the chemical modal vector and each feature in the physical modal vector is calculated to form an nchemx nphys MI matrix, nchem and nphys respectively refer to the number of features in the chemical modal vector and the physical modal vector, then the feature pairs corresponding to the preset percentile value are retained by identifying and screening the MI values in descending order, to obtain a strong dependency feature group.
[0015] In a preferred embodiment, the evaluation results refer to:
[0016] The feature sparsity index obtained by feature sparsity evaluation and the feature stability index obtained by feature stability evaluation.
[0017] In a preferred embodiment, the way of automatically screening key features of substandard grains refers to:
[0018] Fuzzy logic is used to infer the way of automatically screening key features of substandard grains based on the feature sparsity index and the feature stability index.
[0019] In a preferred embodiment, the way of automatically screening key features of substandard grains includes:
[0020] L1 regularization feature selection method and recursive feature elimination feature selection method.
[0021] In a preferred embodiment, the logic for obtaining the feature sparsity index is:
[0022] From the chemical and physical modalities feature set retained after preliminary screening by mutual information, identify the features with mutual information values lower than the average of all feature mutual information values, and mark them as low importance features;
[0023] Count the number of all low importance features as the total amount of invalid or weakly related features in the current feature set, and at the same time obtain the number of all retained features as the total number of features in this round of feature evaluation;
[0024] Calculate the dispersion degree of mutual information values of all retained features in different features, measure the overall information volatility in the form of standard deviation, as an important parameter to measure the uniformity of feature importance distribution;
[0025] The ratio between the number of low importance features and the total number of features is subjected to nonlinear enhancement processing to highlight the influence of the proportion of redundant features on the quality of the model; At the same time, the mutual information volatility parameter is introduced into the calculation system to comprehensively reflect the influence of feature contribution difference on feature selection strategy selection, and the product of the two constitutes the final sparsity index. In the calculation process, in order to adapt to the different model requirements and feature structure differences, two constant factors are set to adjust the sensitivity, which control the amplification weight of the proportion of redundancy and volatility respectively.
[0026] In a preferred embodiment, the logic for obtaining the feature stability index is:
[0027] From the chemical and physical modalities feature set screened by mutual information, carry out multiple cross-validation or use different data division strategies for repeated feature selection, and record all features retained into the final feature set in each selection;
[0028] Statistical the number of features that are always retained in all screening rounds as the "stable selected feature number", that is, the core features that always have discriminant ability in all repeated experiments; At the same time, calculate the average value of the number of features selected in each screening process as the "total selected feature number";
[0029] For stable selected features, respectively count the importance scores corresponding to them in multiple screenings, and calculate the ratio between the standard deviation and the average of these score values as the "score variation ratio" to measure the volatility degree of feature importance;
[0030] The ratio between the "stably selected feature number" and the "total selected feature number" is taken as the first part index, representing the consistency degree of the feature being repeatedly selected under multiple experimental conditions; at the same time, the "one minus the score variation ratio multiplied by the fluctuation adjustment factor" is taken as the second part index, and the two parts are multiplied to obtain the final feature stability index, and in the calculation process, an index enhancement factor is set to adjust the response strength of the stability index to the proportion part.
[0031] Technical effects and advantages of the present application:
[0032] The present application first systematically integrates grain quality chemical information and physical structure information, breaks through the traditional single detection mode, and comprehensively reveals the formation mechanism and dynamic evolution law of defective grain in the whole process of collection, storage and transportation. Through cross-modal mutual information analysis method, the dependence and coupling relationship between chemical modal features and physical modal features are accurately captured, the system analysis of the linkage effect between different influencing factors is realized, and the internal and external causal chain of defective product formation is clarified. Effectively make up for the defects of "information island" and "data fault" in the existing technology for the formation process of defective grain, significantly improve the explainability and prediction ability of the quality change process of grain.
[0033] The present application dynamically determines the automatic feature screening mode through feature sparsity evaluation and feature stability evaluation, ensures that the key features screened have high correlation and high stability, and improves the generalization ability and robustness of the detection model, i.e. grain quality knowledge graph. Adopt fuzzy logic decision mechanism, automatically select regularization method based on sparsity optimization or recursive feature elimination method, ensure that the model can realize adaptive optimization under different grain quality data environment. Realize the dimension reduction and optimization of high-dimensional complex feature data, reduce the calculation complexity of the detection system, improve the efficiency of model training and reasoning, speed up the response speed of grain defective product detection, meet the detection demand of large-scale grain collection, storage and transportation scene.
[0034] The application utilizes a multi-modal information fusion method of fingerprint spectrum and finite element analysis, combines mutual information analysis and dynamic characteristic screening algorithm, and constructs an automatic grain defective product detection and identification mechanism, thereby reducing the dependence on artificial intervention and experience judgment. By constructing a knowledge graph with grain quality and physical-chemical indexes as core nodes, the grain quality evolution law and defective product formation path are systematically expressed, and the knowledge-driven and intelligent reasoning functions of the grain defective product detection system are realized. The knowledge graph supports the whole-process automatic operation of data input, causal analysis and defective grade determination, and significantly improves the intelligent level and business processing efficiency of grain quality monitoring. By constructing high-value nodes and strong dependence paths in the knowledge graph, grain quality dynamic detection and early detection of defective products are formed, and the whole-process quality control in the grain storage and transportation process is realized. According to the weight values of different nodes and paths in the knowledge graph, the main influencing factors of defective grain formation are analyzed, scientific decision-making basis is provided for grain storage and transportation management, the proportion of defective grain is reduced, grain loss is reduced, and the quality guarantee level of the grain industry chain is improved. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to facilitate the understanding of those skilled in the art, the application will be further described below in conjunction with the drawings.
[0036] Fig. 1 A principle diagram of a grain defective product detection method combining fingerprint spectrum and finite element analysis in the application.
[0037] Fig. 2 A recursive feature screening experimental data diagram in the application.
[0038] Fig. 3 A model accuracy comparison diagram before and after recursive screening in the application. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the application will be described clearly and completely below in conjunction with the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0040] Reference Figs. 1-3 The following embodiments are obtained:
[0041] Embodiment 1:
[0042] In China and globally, food safety issues have become the focus of social attention. In the post-harvest links of major food crops such as wheat, rice, corn, and soybeans, due to the complexity of environmental factors and the imperfection of technical means, the quality of food is prone to deterioration, forming substandard grains (such as moldy grains, worm-eaten grains, broken grains, and deteriorated grains). These substandard grains not only reduce the utilization efficiency of food, but also may cause food safety risks and economic losses. The current problems are:
[0043] Single detection technology, limited information: Most existing grain quality detection methods rely on sensory detection and single physicochemical index detection, which are traditional and have low accuracy, and cannot realize multi-dimensional comprehensive analysis of the formation mechanism of substandard grains. Data isolation, model inconsistency: The current detection system cannot realize cross-modal integrated analysis of chemical quality information (such as changes in nutritional components) and physical structure information (such as grain stress deformation and crack propagation).
[0044] The present application proposes a grain substandard product detection method combining fingerprint spectrum and finite element analysis to realize accurate identification and dynamic analysis of the formation process of substandard grains, aiming to: reveal the formation law of substandard grains: through multi-modal information analysis (chemical + physical), systematically clarify the mechanism of grain quality deterioration under different environmental conditions, and fill the gap in the theoretical research of grain storage and transportation process quality control. Build a dynamic model of substandard grains: establish a dynamic model of different substandard product formation processes, and clarify the causal relationship between "quality index - substandard product state", which will help the scientific management of grain.
[0045] Improve the intelligent detection ability of grain quality: adopt fingerprint spectrum and finite element analysis, integrate multi-dimensional features, break through the limitations of traditional detection methods, and realize intelligent and automatic identification of grain quality deterioration.
[0046] Provide technical support for accurate control of substandard grains: through knowledge graph, provide decision-making basis for grain depot management, transportation scheduling, storage control, etc., reduce grain loss, and ensure food safety.
[0047] Promote the intelligent upgrading of grain industry chain; promote the transformation of grain storage and transportation links to intelligent, digital, and fine, and drive the industrial development of related equipment and software systems.
[0048] The present application proposes a grain substandard product detection method combining fingerprint spectrum and finite element analysis, including the following steps:
[0049] The first feature set is obtained by performing fingerprint feature extraction, and the second feature set is obtained by performing finite element model feature extraction; the grain samples (such as wheat, rice, corn, soybean, etc.) are detected by using fingerprint technology (such as infrared spectrum, mass spectrum, fingerprint method), and the chemical component information (such as moisture, starch, protein, fat, cellulose, etc.) of the grain and the distribution characteristics of the metabolic products generated under different storage and transportation conditions are obtained. The first feature set Xchem is formed, which contains a large amount of chemical characteristic data describing the quality change of the grain, and provides detailed chemical information basis for the analysis of the formation mechanism of the defective grain, captures the component fluctuation in the deterioration process of the grain, and the fingerprint can reveal the chemical trajectory of the evolution of the grain from “high quality” to “defective”. It is a necessary prerequisite for analyzing the formation of defective grain.
[0050] Based on finite element analysis (FEA), a physical structure model of grain particles is constructed, the stress behavior, crack generation and expansion, internal damage and deformation of the grain particles under different storage and transportation conditions are simulated, the finite element simulation output data are extracted, and the second feature set Xphys is formed, including stress distribution, crack initiation and expansion rate, crushing probability, displacement, and other physical and mechanical parameters, which reveals the physical change law in the generation process of the defective grain, and provides a basis for understanding the damage mode of the particles in the transportation, stacking and other processes. Physical characteristics are an important source of characterization of particle damage, structure degradation and other characteristics, and are the key to determining the deterioration degree.
[0051] The first feature set and the second feature set are respectively subjected to feature dimension reduction, and redundant variables are removed to obtain a chemical modal feature set and a physical modal set; the principal component analysis (PCA) dimension reduction algorithm is applied to the first feature set Xchem and the second feature set Xphys, respectively, to remove redundant, highly correlated or noisy feature variables, compress the data dimension, retain the main features that best represent the data structure and change trend, and finally form two modal feature subsets, which are respectively referred to as the chemical modal feature set and the physical modal feature set. Reducing data complexity, improving the efficiency and accuracy of subsequent analysis, avoiding “dimension disaster”, removing irrelevant variables, highlighting the essential features of grain quality change, and improving the robustness of the model. This step prepares effective features at the “single modal” level, ensuring that the next feature evaluation is more accurate.
[0052] With the substandard grain grade as the target label, the mutual information of each feature in the chemical modal feature set and the target label is calculated to perform the first layer of dimension reduction to obtain a chemical modal vector. The mutual information of each feature in the physical modal set feature set and the target label is also calculated to perform the first layer of dimension reduction to obtain a physical modal vector. The mutual information (MI) value of each feature in the two modal feature sets (chemical and physical) and the target label (substandard grain grade) is calculated. By setting a standard threshold (such as the average MI value or a specific quantile), features with a mutual information value higher than the threshold are screened out, and variables that do not contribute or weakly contribute to the target are removed. Finally, two optimized feature vectors are formed: the chemical modal vector and the physical modal vector. The core features that most affect the prediction of substandard grain grading are accurately identified, further improving data effectiveness, laying a solid foundation for subsequent cross-modal analysis, reducing the influence of invalid features, improving analysis accuracy, and ensuring that the subsequent "cross-modal mutual information analysis" is performed between the most valuable feature sets.
[0053] Under the condition that the target label is known, cross-modal mutual information analysis is performed to obtain a strong dependency feature group. At the same time, feature sparsity evaluation and feature stability evaluation are performed. Based on the evaluation results, the way to automatically screen the key features of substandard grain is selected and the strong dependency feature group is processed. Finally, a grain quality knowledge graph with grain quality and physical-chemical indicators as core nodes is constructed for detecting substandard grain.
[0054] The mutual information values between the features in the chemical modal vector and the physical modal vector are calculated to form an nchem x nphys MI matrix. Under the condition that the target label is known, strong dependency combinations between the two types of features are identified. According to the descending order of mutual information values, features exceeding a predetermined percentile (such as 95%) are screened out to form a strong dependency feature group. This reveals the potential coupling mechanism between chemical features and physical features, such as how chemical degradation affects physical damage. This forms the basis for multi-modal data fusion and provides high-value data input for further comprehensive analysis of the dynamic mechanism of substandard grain formation. The strong dependency feature group is the key basis for constructing a joint model, i.e., a knowledge graph.
[0055] The strong dependency feature group is evaluated for feature sparsity (calculating a sparsity index) and feature stability (calculating a stability index). Sparsity evaluation is based on the distribution and volatility analysis of mutual information values, and stability evaluation is based on the consistency of features in different data splits and cross-validation. Sparsity evaluation identifies the proportion of sparse and invalid features in the feature data distribution, guiding the selection of optimization algorithms. Stability evaluation ensures that the selected features remain robust in different scenarios and data sampling, avoiding overfitting or instability factors. This provides a quantitative basis for selecting automatic feature screening algorithms, ensuring the scientificity and dynamic adaptability of the screening strategy.
[0056] Based on the sparsity index and stability index, the automatic feature screening method is selected by fuzzy logic reasoning: high sparsity and low stability→L1 regularization (LASSO) is adopted; low sparsity and high stability→recursive feature elimination (RFE) is adopted; further automatic screening and optimization of the strong dependency feature group are carried out. The dynamic intelligent decision of the feature screening method is realized, different data structures and distributions are adapted, the screening effect is improved, the features of the final input model are efficient, refined, have strong interpretability and stability, the dependence on human decision is reduced, and the automation and intelligence level of the analysis system is improved.
[0057] Based on the strong dependency feature group and the features screened and optimized automatically, a grain quality knowledge graph is constructed, the knowledge graph nodes cover grain quality indexes, chemical compositions, physical structure characteristics, and defective grade information, the edges represent the causal relationship and influence mechanism between features, the knowledge graph is used for defective grain identification, classification and quality monitoring, the transition from single parameter detection to multi-source information comprehensive analysis is realized in defective grain detection, the knowledge graph provides knowledge support for the intelligent decision system, realizes dynamic early warning, quality control and traceability analysis, promotes intelligent management of grain storage and transportation, and improves the grain safety guarantee capability and storage and transportation efficiency.
[0058] When the feature dimensionality is reduced, the principal component analysis method is used. Taking the first feature set, i.e., the fingerprint spectrum data, as an example, in the present application, the first feature set is a plurality of grain chemical composition features extracted by fingerprint spectrum technology (such as infrared spectrum, mass spectrum, etc.). For example: water content, protein content, starch content, fat content, cellulose content, and other metabolite characteristic peak values, which constitute the “original indexes” in the fingerprint spectrum.
[0059] Sample data standardization: each feature collected by the fingerprint spectrum needs to be uniformly processed due to different dimensions (such as water content expressed in percentage, protein expressed in milligrams per gram). Through the standardization method, the influence of different units and dimensions is eliminated, and each feature is ensured to participate in calculation under the same conditions. After processing, the data of each index is centralized to the same order of magnitude, which is convenient for subsequent analysis.
[0060] Obtain correlation: analyze whether there is correlation between each index in the first feature set, for example, find that the change trend of water content and starch content in different samples is similar, and there is a high correlation between them, at this time, through the correlation analysis means in the prior art, it can be known which indexes have redundant information.
[0061] Determining the principal component directions: After analyzing the correlations between indicators, mathematical methods are used to identify the directions that describe the greatest data variance. Each such direction represents a new composite indicator, known as a "principal component." The first principal component represents the most informative composite indicator in the data and can explain the vast majority of data variation. Building on the first, the second principal component continues to describe minor information in the remaining data, and so on.
[0062] Obtain the contribution of the principal components: Each principal component has an "explanatory power ratio" for how much information it can explain about the original data. For example, the first principal component may explain 50% of the overall variation in the data, the second principal component another 30%, and the third principal component 10%. Based on the size of the contribution, decide how many principal components to retain. Usually, the principal components with a cumulative explanation rate of more than 85% are selected.
[0063] Forming new comprehensive features: The original multiple indicators are recombined according to the direction of the principal components to form a new set of comprehensive features. For example, principal component one may combine the changes in moisture, starch, and protein; principal component two may combine the changes in fat and cellulose. These two principal components can roughly describe the main information of grain quality.
[0064] Screening for effective features and eliminating redundant variables: Based on principal component contributions and actual business needs, the most representative principal components are selected for subsequent analysis. Indicators with low weights that are grouped within the principal components can be eliminated to avoid interference from redundant data. The resulting chemical modal feature set, derived from principal component analysis, serves as crucial input for subsequent modeling.
[0065] Significance of the analysis results: Through the principal component analysis method, a large amount of original feature data in the first feature set was streamlined into a few representative information comprehensive indicators, which reduced the complexity of model calculation, eliminated redundancy and noise in the data, ensured the efficiency and effectiveness of data input, and laid a data foundation for cross-modal analysis and modeling of the formation patterns of defective grain.
[0066] Suppose the following grain indicators are extracted from a fingerprint: moisture content, protein content, starch content, fat content, and cellulose content. Principal component analysis reveals that principal component 1, which contains comprehensive information about the three main indicators (moisture, starch, and protein), explains 60% of the total information; principal component 2, which contains information about fat and cellulose, explains 25% of the total information; and principal components 3 and beyond explain less information and can be ignored. Therefore, retaining only principal components 1 and 2 can represent 85% of the original feature information. These two principal components serve as the final input features for the "chemical modal vector."
[0067] The second feature set is derived from finite element model simulation analysis of grain particles. Under different storage, transportation, and stacking conditions, grain particles (such as wheat, rice, corn, and soybeans) can be simulated through finite element simulation to obtain a large amount of physical response characteristic data. This includes but is not limited to: maximum stress, maximum strain, crack initiation time, crack propagation rate, particle displacement, damage energy accumulation, and breakage probability, etc. These constitute the second feature set (physical modal characteristics).
[0068] An example of the principal component analysis step is as follows:
[0069] Standardization: Convert different dimensional characteristic values to dimensionless standard values (for example, mean value is 0, standard deviation is 1).
[0070] Correlation matrix construction: It is found that the maximum stress, crack propagation rate, and breakage probability are highly correlated, and the particle displacement and crack initiation time are weakly correlated.
[0071] Extract principal components: The first principal component integrates "maximum stress + crack propagation rate + breakage probability + damage energy" to describe the overall damage trend, and the second principal component integrates "particle displacement + crack initiation time" to describe the initial physical state influence.
[0072] Contribution rate: The first principal component explains 60% of the total information, and the second principal component explains 30% of the total information, with a cumulative explanation rate of 90%, meeting the analysis requirements.
[0073] New feature formation: The original seven physical characteristics are simplified into two principal component integrated characteristics, and only these two principal components need to be considered in subsequent analysis and modeling, rather than all original variables.
[0074] The first layer of dimensionality reduction refers to:
[0075] Set a standard threshold, compare the mutual information between each feature in the chemical modal feature set and the target label, and the mutual information between each feature in the physical modal set and the target label, respectively, with the standard threshold. Features below the standard threshold are removed, and features not below the standard threshold are retained.
[0076] In the process of grain defect detection and analysis, the system obtains a large amount of data features from different modalities, including: a set of chemical modal features (from fingerprint technology), reflecting the change information of moisture, starch, protein, fat and other components in grain. A set of physical modal features (from finite element analysis technology), reflecting the stress condition, crack propagation, damage probability and other state information of the physical structure of grain particles during storage, transportation and other processes. Due to the large number of these feature data, there must be some redundant feature information, noise interference or weak relationship with grain defect determination. Therefore, it is necessary to effectively filter these data and retain important features meaningful for analysis. The first layer of dimensionality reduction is based on this demand, and the key step for single modal internal feature selection is designed. The specific method of the first layer of dimensionality reduction is:
[0077] Step one: Take the grain defect grade as the target label: First, determine the quality state of the grain sample by grading, such as high-quality grain, light defective grain, medium defective grain, and heavy defective grain. This grade is considered as the target of classification or prediction, and is used to judge the closeness of each feature to it.
[0078] Step two: Calculate the information dependence between each feature and the target label: Based on each feature, analyze its influence on the grain defect grade. Through information theory methods (such as information gain or mutual information calculation), quantify the correlation between each feature and the defect grade. The analysis result is reflected in the "information value" corresponding to each feature. The higher the value, the closer the relationship between the feature and the target label, and vice versa, indicating that the feature has less impact on determining the grain grade.
[0079] Step three: Set a standard threshold: According to system requirements, a filtering threshold is set in advance, usually based on the following method. The threshold is used as the standard to judge whether to retain the feature.
[0080] Method one: Set the threshold based on the statistical distribution of mutual information value. Principle: Calculate the mutual information value between all features and the target label to form a sequence of mutual information. Determine the threshold by statistical method. Specific operation steps: Calculate the average value and standard deviation of the mutual information value sequence; set the threshold as the average value plus the standard deviation of the preset multiple, for example: threshold = average value + one times standard deviation (moderate filtering); threshold = average value + two times standard deviation (conservative retention); only retain features with mutual information value higher than the threshold. Advantages: simple calculation, strong adaptability; can dynamically adjust the standard deviation multiple to adjust the retention strength.
[0081] Method two: Set threshold based on percentile method, principle: sort according to the size of mutual information value, set a certain percentile as the screening line. Specific operation steps: sort all feature mutual information values from high to low; Select for example the 80th percentile (or the 90th percentile) corresponding value as the threshold; Keep the features with mutual information value higher than the value. Advantage: easy to control the screening ratio; Especially suitable for high-dimensional feature set.
[0082] Method three: Based on model performance change feedback method, principle: through multiple experiments to observe the response of model performance to mutual information threshold change, find the optimal threshold point. Specific operation steps: Set a group of different candidate thresholds (such as: average value, average value plus one standard deviation, 80% percentile, etc.); At each threshold, use the retained features to train the model and evaluate the accuracy, recall rate and other performance; Select the threshold corresponding to the optimal performance as the final screening standard. Advantage: The threshold is highly consistent with the final model effect; Especially suitable for tasks requiring high precision.
[0083] Method four: Set threshold combined with domain knowledge and expert score, principle: combine expert experience judgment of feature importance with mutual information value to set threshold. Specific operation steps: Domain experts score part of the key indicators and provide feature priority suggestions; Match with mutual information value to find the area where both are highly evaluated; Set a moderate retention threshold according to expert opinion (such as retaining the top 30% indicators with high label correlation). Advantage: Suitable for data with actual business background; Strengthen the interpretability between features and business indicators.
[0084] The above four methods are commonly used and well-known threshold setting strategies in the prior art. In the implementation of the present application, one method can be selected according to different requirements of precision, stability and interpretability of the task.
[0085] Step four: Feature screening: Compare the information value of each feature in the chemical and physical modal feature sets with the pre-set standard threshold one by one: if the information value of the feature is lower than the standard threshold, it means that the feature has limited effect on the judgment of grain defective grade, and is discarded. If the information value is higher than or equal to the standard threshold, it means that the feature has strong discrimination ability and is retained.
[0086] Step five: Output the first layer dimension reduction result: After the above screening process, two new feature sets are obtained: chemical modal vector, containing all the key chemical component features that pass the screening; Physical modal vector, containing all the key physical structure features that pass the screening. These two feature sets will be used as the core input data for subsequent cross-modal analysis and model construction.
[0087] The significance of the first layer of dimension reduction: remove irrelevant features, optimize data structure: by removing features with weak correlation with grain defect grade, reduce data redundancy, improve data quality, avoid invalid information interference, improve the accuracy and efficiency of subsequent modeling and analysis. Reduce computational complexity and improve system efficiency: after removing invalid or low information contribution features, the data dimension is significantly reduced, reducing the occupation of computing resources and improving the analysis processing speed, especially suitable for processing large-scale and diverse grain detection data. Enhance the generalization ability and robustness of the system: after the first layer of dimension reduction screening, the retained features are more representative and can reflect the essential laws of grain quality changes, enhancing the adaptability of the analysis system to different samples and environmental conditions. Lay the foundation for cross-modal feature analysis: the first layer of dimension reduction effectively compresses the data dimension, ensuring that the processing object is a core feature with rich information and clear meaning when performing cross-modal feature mutual information analysis, improving the scientificity and reliability of subsequent feature fusion and comprehensive analysis.
[0088] Cross-modal mutual information analysis refers to:
[0089] Calculate the mutual information value between each feature in the chemical modal vector and each feature in the physical modal vector to form an nchem x nphys MI matrix, where nchem and nphys represent the number of features in the chemical modal vector and the physical modal vector, respectively. Then, identify the features corresponding to the MI values exceeding the preset percentile value in descending order, and select the strong dependency feature group.
[0090] In the present invention, "cross-modal" refers to two different sources of data types: chemical modal features and physical modal features. Chemical modal features are derived from fingerprint technology and reflect the chemical composition changes of grain quality, such as fluctuations in moisture, starch, and protein content. Physical modal features are derived from finite element model simulation analysis and reflect the physical changes of grain particles during storage and transportation, such as stress conditions, crack generation, and fragmentation degree. Since grain quality changes are often caused by both chemical and physical factors, analyzing the dependency between different modal features can help reveal the mechanism of grain defect formation.
[0091] Cross-modal mutual information analysis measures the degree of information correlation between different modal features to determine whether they have significant dependency or synergistic change trends. The higher the mutual information value, the stronger the correlation between the two features, indicating that there may be an internal mechanism that jointly affects grain quality changes. The specific analysis steps are as follows:
[0092] Step one: determine the input data: input data is the two feature sets after the first layer of dimension reduction: chemical modal vector, including all screened and retained chemical features. Physical modal vector, including all screened and retained physical features.
[0093] Step 2: Calculate the mutual information between different modal features: Each feature in the chemical modal vector is paired with each feature in the physical modal vector. Information theory analysis is performed between each pair of features, and their mutual information value is calculated. Mutual information measures the dependency between two different modal features, reflecting the degree of information sharing. For example, the mutual information between moisture content (a chemical feature) and crack growth rate (a physical feature) indicates whether moisture significantly affects the degree of grain breakage.
[0094] Step 3: Form the Mutual Information Matrix: Integrate the mutual information values of all paired features to form a matrix. Rows represent the serial numbers or names of chemical modal features; columns represent the serial numbers or names of physical modal features. Each element in the matrix corresponds to the mutual information value of a feature pair. This matrix comprehensively reflects the distribution of dependencies between all chemical and physical features.
[0095] Step 4: Identify strongly dependent feature combinations: Sort all mutual information values in the mutual information matrix, sorting from highest to lowest. Set a screening criterion, typically based on a percentile, such as selecting feature pairs with mutual information values in the top 5 percentile. All feature combinations above the pre-set percentile are considered to have significant dependencies and are retained. The resulting feature pairs form the strongly dependent feature group.
[0096] Step 5: Output strongly dependent feature groups: The retained strongly dependent feature groups contain both chemical modal features and physical modal features. These feature groups will play a key role in subsequent model training, analysis, and testing.
[0097] The analysis aims to: reveal synergistic effects between data from different modalities: By analyzing mutual information, we can identify synergistic patterns between chemical and physical changes. For example, detecting an increase in moisture may also lead to an increase in cracks. This synergistic relationship reveals a joint factor in the formation of defective grain. It also promotes effective data fusion: Strongly dependent feature pairs serve as key connections between different modalities, providing a reliable basis for subsequent data fusion and model construction. It also avoids information redundancy and the introduction of invalid data, improving the model's analytical capabilities and predictive accuracy. It also provides a scientific basis for the formation mechanism of defective grain: Strongly dependent feature pairs reveal the causal chain in the process of grain quality deterioration, clarifying the internal and external factors that contribute to the formation of defective grain. It also facilitates the construction of a dynamic model of defective grain formation, guiding quality management and risk control during grain storage and transportation. It also improves the accuracy and efficiency of the defective grain detection system: By selecting feature combinations with high mutual information, the number of features can be reduced, improving the efficiency of the detection system. This ensures that all retained features have a high explanatory power for grain quality assessment, thereby enhancing the overall performance of the intelligent defective grain detection system.
[0098] The evaluation result refers to a feature sparsity index obtained by performing feature sparsity evaluation and a feature stability index obtained by performing feature stability evaluation. The selected automatic screening of substandard grain key feature mode refers to: using fuzzy logic, based on the feature sparsity index and the feature stability index, to infer the automatic screening of substandard grain key feature mode. The automatic screening of substandard grain key feature mode includes: an L1 regularization feature screening method and a recursive feature elimination feature screening method.
[0099] The feature sparsity index acquisition logic is:
[0100] From the chemical modal vector and the physical modal vector obtained after preliminary mutual information screening, the number of features judged to be "low importance" is obtained and denoted as "low importance" refers to features with MI values lower than the average value, the total number of features in the chemical modal vector and the physical modal vector is obtained and denoted as , and the standard deviation of the MI values corresponding to each feature in the chemical modal vector and the physical modal vector is obtained and denoted as , and then substituted into the following formula:
[0101] ;
[0102] , represents a preset sparsity sensitivity factor, , represents a preset mutual information fluctuation factor, , represents a feature sparsity index.
[0103] In actual grain substandard product detection, cross-modal feature data has multiple sources and high dimensions, but only part of the features actually contribute to the recognition result. In order to avoid the calculation burden and model performance decline caused by "feature redundancy", a quantitative index needs to be introduced to judge the "sparsity" of the feature set, and according to the index, it is determined whether to use an algorithm with strong sparsity (such as L1 regularization). The degree of redundancy is judged by the proportion of low importance features, the uncertainty of features is measured by mutual information fluctuation, and the sensitivity and adaptability of the algorithm are adjusted by the sparsity sensitivity factor. The number of low importance features These features basically have no substantial impact on substandard grain recognition, and may even become a model disturbance term. The number of all features in the cross-modal feature set is the denominator for calculating the proportion of low importance features. The higher the ratio, the more redundant data, and the stronger the sparsity. The mutual information standard deviation (mutual information fluctuation) refers to the standard deviation of the mutual information value between all retained features and the target label. The greater the fluctuation, the less stable the importance of the features. A high standard deviation means that the data is highly uncertain, and the sensitivity to sparse features needs to be improved. The sparsity sensitivity factor is used to control the response degree of the sparsity index to the proportion of low importance features. When When larger (greater than 1), it means that the influence of amplifying the proportion of low importance features makes the model more sensitive to sparsity, and tends to activate stronger sparse feature screening mechanism (such as L1 regularization). The mutual information fluctuation factor is used to measure and adjust the influence strength of mutual information fluctuation. When When larger, the change of mutual information standard deviation intensifies the influence of the final sparsity index, making the model pay more attention to the instability problem of feature information. The sparsity index (FSI) is the final output value of the formula, representing the sparsity level of the current feature set. The higher the value, the greater the feature redundancy and instability, and it is recommended to use strong sparse automatic screening method (such as L1 regularization). The lower the value, the more intensive and stable the feature distribution, and it is recommended to use recursive feature elimination, a step-by-step fine screening method.
[0104] The logic for obtaining the feature stability index is as follows:
[0105] From the preliminary mutual information screening of the chemical and physical modal vectors, obtain the number of features that are always selected into the chemical and physical modal vectors under multiple cross-validation or different data partitioning, and record it as The average total number of features selected in all screenings is recorded as The ratio of the standard deviation to the mean of the total weight of the features that are always selected into the chemical and physical modal vectors in all screenings is recorded as Then substitute into the following formula:
[0106] ;
[0107] represents the preset stability enhancement factor, represents the preset feature change fluctuation factor, represents the feature stability index.
[0108] The feature stability index is a comprehensive index for measuring the consistency and reliability of a certain feature in the results under different data partitions, different sample training or multiple cross-validation conditions. In the grain defective product detection method, the cross-modal feature data sources are complex, and the sample data has diversity and non-stationarity. Therefore, the feature set obtained under different batches and different conditions often has inconsistent screening results, which affects the generalization ability and robustness of the subsequent model. In order to scientifically evaluate the reliability and stability of the feature screening result, the feature stability index is constructed to dynamically reflect the consistency and stability in the feature selection process, and to guide the automatic feature selection algorithm selection. The feature stability index determines whether the cross-modal feature has stable representativeness under different sample conditions, avoids the model relying on volatile and unstable features, improves the generalization ability of the detection model, and is one of the important inputs of the fuzzy logic rule, which determines which automatic screening method to use.
[0109] The number of stable selected features refers to the number of features that are always selected into the chemical modal vector and the physical modal vector under multiple cross-validation or different data partitions. The larger the number is, the higher the stability of the feature under different conditions is, and it is a reliable feature. The average value of the total number of selected features in all screening refers to the average value of the number of selected features in each feature screening, which is a sample number basis and is used as a denominator to measure the proportion of the number of stable selected features, reflect the consistency level of the feature, and the feature weight change coefficient The feature weight change coefficient represents the weight fluctuation degree of the feature in the multiple screening process. The larger the fluctuation is, the more unstable the feature is under different conditions, which may be a pseudo-important feature caused by data disturbance. The smaller the fluctuation is, the more consistent the importance score of the feature is in different experiments, and the stronger the reliability is. Mathematically, it is the ratio of the standard deviation to the average of the total feature weight of the features that are always selected into the chemical modal vector and the physical modal vector in all screening. The mutual information value reflects the strength of the information dependence between a certain feature and the target label (such as the grade of defective grain). The stronger the information dependence is, the higher the importance of the feature in the model is. The mutual information value can be directly used as the feature weight, or it can be simply transformed (normalized) as the weight score.
[0110] Stability enhancement factor for controlling the sensitivity of the formula to the proportion of stable selected features, The larger the value is, the more attention is paid to the high-stability feature, and the influence of the proportion on the overall index is further amplified. The feature change fluctuation factor is used to control the sensitivity of the formula to the feature weight volatility. The larger the value is, the more attention is paid to the stability of the feature weight, and the larger the fluctuation is, which leads to a significant decrease in the index. The feature stability index is the final output comprehensive index, which quantitatively evaluates the overall stability of the feature set. The higher the value is, the more stable the feature set is under different conditions, and the stronger the reliability is.
[0111] The feature sparsity index is used to measure the proportion of invalid or low-contribution features in the current feature dataset. When there are a large number of features in the dataset that have limited effect on the identification of substandard grain grades, such features are referred to as sparse features. The higher the sparsity index, the more irrelevant information in the feature data, and the more sparse and scattered the data. This index can be used as an important basis for determining whether to use an efficient compressed feature selection method. High sparsity can cause an increase in model complexity, a decrease in computational efficiency, and a decrease in result stability, so effective means must be used to compress the feature space and eliminate redundant interference information.
[0112] The feature stability index is used to measure the consistency of a feature in different sample groupings, data divisions, or cross-validation processes. If a feature consistently exhibits good classification and identification ability in different situations and is frequently selected, it indicates that the feature has strong stability. The higher the stability index, the more consistent and reliable the feature remains under different data distributions and model training conditions. A feature with high stability is beneficial to improving the generalization ability of the model and reducing identification errors caused by data fluctuations.
[0113] Steps and principles of fuzzy logic:
[0114] Background of fuzzy logic reasoning: In the process of detecting substandard grain, the values of the feature sparsity index and the feature stability index are not simply "high" or "low", and there is continuity and fuzziness between them. Therefore, by using fuzzy logic to comprehensively evaluate these two indices, it can be more scientific and dynamic to determine which automatic selection method to use.
[0115] Steps of fuzzy logic reasoning:
[0116] First step, input fuzzification: The feature sparsity index and the feature stability index are taken as input variables and divided into multiple fuzzy levels: "low sparsity", "medium sparsity", and "high sparsity" for the sparsity index, and "low stability", "medium stability", and "high stability" for the stability index. These levels are quantified by membership functions. The membership function is defined using a triangular membership function, which is symmetrical, making it easy to calculate and interpret.
[0117] Second step, the definition of membership function of triangular function: triangular membership function is a basic membership function, the form is: the function image presents "peak" shape, the middle point membership is the maximum value, left and right gradually weaken. The function is determined by three points, respectively, the starting point, the peak point and the end point. Each input value is calculated according to its position in the triangular function interval. The membership value represents the degree of belonging to a certain level, ranging from zero to one. For example: if the sparsity index is a certain value, the value is located at the peak position of the "moderate sparsity" membership function, and the membership is one; if the value deviates from the peak position, the membership decreases linearly.
[0118] Third step, rule base design: according to different combinations of sparsity index and stability index, set fuzzy rule base: for example: if the sparsity is "high" and the stability is "low", select the first screening method (i.e. regularization method). If the sparsity is "low" and the stability is "high", select the second screening method (i.e. recursive screening method). If the sparsity is "moderate" and the stability is "moderate", you can adopt a mixed strategy or bias towards regularization method.
[0119] Fourth step, fuzzy reasoning calculation: according to the membership degree corresponding to the input value and the preset fuzzy rule, the comprehensive operation is carried out. The maximum membership principle or weighted average method is used to synthesize the output result. The reasoning process is equivalent to selecting the conclusion with the highest degree of conformity from different fuzzy conditions.
[0120] Fifth step, output defuzzification: convert the fuzzy reasoning result into clear decision output and determine which automatic screening method to choose. For example: the output result is "biased towards regularization" → execute regularization screening. The output result is "biased towards recursive screening" → execute recursive feature screening. If both are close, you can weigh the pros and cons and prefer to use regularization, and then use recursive method for fine tuning.
[0121] The regularization screening method has strong sparsity processing ability, which can automatically compress the feature coefficients with low importance to zero, realizing the effective compression of feature space. When the feature sparsity index is high, it means that the invalid information in the feature accounts for a large proportion, and the regularization method can quickly screen out irrelevant variables, improving the efficiency and stability of the model. The regularization method has fast calculation speed, which is suitable for large-scale data processing, especially for high-dimensional sparse data environment.
[0122] The recursive feature screening method can gradually eliminate the features with small influence by repeatedly training and evaluating the model, and retain the features that contribute most to the target prediction. When the feature stability index is high, it means that the feature performs consistently under different sample and model conditions, which is suitable for using fine screening method to further optimize the model performance. Recursive screening method can fully utilize the discriminant ability of stable features to improve the interpretability and accuracy of the final model.
[0123] In the implementation of the present application, the experimental process of the recursive feature screening method is as follows: first, a preliminary candidate feature set is determined from the modal feature set after mutual information screening and principal component dimensionality reduction processing. This set is respectively from the chemical modal extracted by the fingerprint spectrum and the physical modal extracted by the finite element analysis, and represents all effective variables currently available for the grade discrimination of substandard grain. Next, in the evaluation of the feature stability index, the present application sets a multi-round evaluation mechanism based on cross-validation. Specifically, in the total data set, the five-fold cross-validation method is used to randomly divide the data into five parts, and each time four parts are used as the training set and one part is used as the validation set, for a total of five rounds. In each round, whether each feature is retained by the feature selector (such as the mutual information filter or the L1 regular selector) is recorded, and the features selected into the model are assigned an importance score value. For each feature, the number of times selected in all rounds divided by the total number of rounds is the selection frequency. At the same time, the standard deviation and the mean of the score of each feature in different rounds are counted, and the importance variation degree of the feature is calculated. On this basis, the threshold of the feature stability index is set to 0.75. That is, when the selection frequency of a feature is greater than or equal to 75%, and the variation degree of its score is less than 80% of the average level, it is determined to be a stable feature. When the number of stable features in the entire candidate feature set accounts for more than 50%, it is considered that the overall feature stability index is high.
[0124] When the feature stability index is high, the recursive feature elimination algorithm is started. The operation steps are as follows: initial training: use all stable features to train the classification model, and calculate the overall performance indicators such as accuracy and F1 value. Feature scoring: use the feature importance calculation method inside the model (such as the Gini coefficient based on tree model or the weight coefficient based on support vector machine) to sort the current features. Feature elimination: one to two lowest-scored features are eliminated in each round, and the model performance is recorded. Convergence judgment: if the model performance does not significantly improve or shows a downward trend after three consecutive rounds of elimination, the screening is stopped, and the feature set of the last round is retained as the final core feature set.
[0125] In the experimental process, five groups of data sets under different modalities (containing different batches of samples of wheat, rice, corn, and soybeans) corresponding to samples A to sample E are used for comparison and verification on the training set and the test set. For example, the recursive feature screening experimental data in Fig. 2 , and Fig. 3The comparison of model accuracy before and after recursive screening in the figure, the y-axis represents the model accuracy, and the x-axis represents the corresponding group before and after recursion. From the recursive feature screening experiment data and the corresponding figure, the following key information can be seen: stability verification: the number of stable features finally selected in the five different experimental samples is basically consistent (between thirty and forty), indicating that the feature screening has good stability and repeatability. Accuracy improvement: after using recursive feature screening, the model accuracy is significantly improved in all samples, with an average improvement of more than 4%. Misjudgment rate decreases: the misjudgment rate is significantly reduced after recursive screening, indicating that the method effectively eliminates pseudo-features without distinguishing power. After using recursive screening, the accuracy of each group of experiments increases significantly; the precision improvement is about 3% to 5%; experiments B and E show the highest subsequent accuracy, more than 90%; the figure directly verifies that the feature screening strategy has a positive promoting effect on the improvement of model performance. This shows that the recursive screening strategy based on the feature stability index used in the invention has strong adaptability and repeatability under different sample and experimental conditions.
[0126] By introducing a fuzzy logic system, the feature quality of the analysis object is dynamically evaluated and analyzed based on the feature sparsity index and the feature stability index, and the most suitable automatic feature screening method is selected scientifically and reasonably: it can ensure the efficiency of data compression and model training, and ensure that the defective grain detection model has excellent generalization ability and robustness, and finally significantly improve the accuracy and practicality of the grain defective product detection system.
[0127] In the method of the invention, through cross-modal mutual information analysis and feature screening, key features of defective grain with high dependence and high stability are obtained. These features include chemical indicators and physical structure information of grain quality, which constitute the core feature system describing the change of grain quality and the formation rule of defective grain.
[0128] In order to realize the systematic and structured management of grain defective product detection, it is necessary to construct a grain quality knowledge graph based on the above features. The knowledge graph can reveal the internal relationship and mechanism between different quality characteristics, and support intelligent detection and classification of grain defective products.
[0129] The specific steps of constructing the knowledge graph are:
[0130] Determine core nodes: Core nodes are the most basic units in the knowledge graph, representing key entities related to grain quality. In this invention, core nodes are divided into three categories: grain quality indicator nodes, including moisture, starch, protein, fat, cellulose, and mineral elements; grain physical structure indicator nodes, including crack length, crack propagation rate, particle displacement, stress concentration coefficient, and breakage probability; and grain quality state nodes, including high-quality grain, light residual grain, medium residual grain, and heavy residual grain.
[0131] Define node attributes: Each node has different attribute information defined according to its characteristics. Grain quality indicator node attributes include detection values, fluctuation intervals, and change trends. Physical structure indicator node attributes include simulation results and damage assessment values. Quality state node attributes include classification standards, reference level descriptions, and detection labels.
[0132] Establish associated edges between nodes: Establish an association between nodes, including cause-and-effect relationships, such as increased moisture leading to faster crack propagation; interaction relationships, such as the coupling relationship between starch content changes and breakage probability; and conditional dependency relationships, such as water exceeding standards being a key factor in grain quality decline under high temperature and humidity conditions.
[0133] Determine relationship strength and weight: Each associated edge can be assigned different weight values using expert assignment methods based on cross-modal mutual information analysis and feature stability evaluation results, reflecting the influence degree between different indicators. The greater the weight value, the more significant the relationship's impact on grain quality judgment. This approach strengthens the connection between high-value nodes in the knowledge graph and highlights important information.
[0134] Form the knowledge graph network structure: All nodes and edges are structured and combined to form a complete knowledge graph network. The knowledge graph presents a tree or network structure, with different nodes organized hierarchically and classified according to their relationships, forming a clear and compact graph system.
[0135] Visualize the knowledge graph: Display the knowledge graph network graphically, including nodes, edges, and weight information. Different types of nodes are distinguished by shape and color, and relationship strength is represented by edge thickness or color depth, enhancing the graph's intuitiveness and readability.
[0136] Method for using the knowledge graph:
[0137] Defective Grain Detection: Grain sample test data, including chemical composition data and physical property simulation results, is input into the system. Matching nodes and paths are automatically retrieved from the knowledge graph. Based on the distribution of sample characteristics within the graph, the corresponding quality status node is determined. The system then determines whether the sample is high-quality or defective, as well as the defect grade. Based on the input sample test data, reasoning and analysis are performed along the edges between nodes within the grain quality knowledge graph. By searching for the sum of weights along different paths, the path with the largest sum of weights is prioritized. This path indicates the most significant causal chain of factors leading to changes in grain quality for the current sample. The path with the largest sum of weights reflects the most significant chain of factors influencing quality. The final grain quality grade is output based on the grain quality status node to which the path ends. For example, if the path ends at the "Severely Defective Grain" node, the system outputs a "Severely Defective Grain" judgment. If the path ends at the "High-Quality Grain" node, the system outputs a "High-Quality Grain" judgment.
[0138] Quality Status Analysis and Traceability: Based on the causal relationships within the knowledge graph, we analyze the main factors and causal pathways that lead to defective grain. For example, the analysis indicates that the current sample has an increased probability of breakage due to excessive moisture and rapid crack propagation, ultimately resulting in severe defective grain.
[0139] Risk Warning and Quality Control: The knowledge graph system identifies potential risk points during grain storage and transportation based on historical data and monitoring results. For example, when indicators such as temperature, humidity, and moisture approach critical values, an early warning message is issued, prompting relevant managers to take intervention measures to prevent grain quality deterioration.
[0140] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0141] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0142] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0143] Those skilled in the art can clearly understand the specific working process of the system, device and unit described above for the convenience and brevity of description, which can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0144] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for detecting defective grains by combining fingerprint analysis with finite element analysis, characterized in that: The following steps are involved: Extracting fingerprint features to obtain a first feature set, and extracting finite element model features to obtain a second feature set; Performing feature dimensionality reduction on the first feature set and the second feature set respectively, eliminating redundant variables, and obtaining the chemical modal feature set and the physical modal set; Taking the defective grain grade as the target label, the mutual information between each feature in the chemical modal feature set and the target label is calculated, and the first layer of dimensionality reduction is performed to obtain the chemical modal vector. The mutual information between each feature in the physical modal feature set and the target label is calculated, and the first layer of dimensionality reduction is also performed to obtain the physical modal vector. Under the condition that the target label is known, cross-modal mutual information analysis is performed to obtain a group of strongly dependent features. Feature sparsity and feature stability assessments are also performed. Based on the assessment results, a method for automatically screening key features of defective grain is selected and the group of strongly dependent features is processed. Finally, a grain quality knowledge graph with grain quality and physical-chemical indicators as core nodes is constructed for the detection of defective grain. Performing cross-modal mutual information analysis means: The mutual information value between each feature in the chemical modal vector and each feature in the physical modal vector is calculated to form an nchem×nphys MI matrix, where nchem and nphys refer to the number of features in the chemical modal vector and the physical modal vector, respectively. Then, the feature pairs corresponding to the MI values exceeding the preset percentile values are identified and retained to obtain a strongly dependent feature group.
2. The method for detecting defective grains by combining fingerprint analysis and finite element analysis according to claim 1, characterized in that: When reducing feature dimensionality, principal component analysis is used.
3. The method for detecting defective grains by combining fingerprint and finite element analysis according to claim 2, characterized in that: The first level of dimensionality reduction refers to: A standard threshold is set, and the mutual information between each feature in the chemical modal feature set and the target label, as well as the mutual information between each feature in the physical modal feature set and the target label, are compared with the standard threshold respectively. Features below the standard threshold are eliminated, and features not below the standard threshold are retained.
4. The method for detecting defective grains by combining fingerprint analysis and finite element analysis according to claim 3, characterized in that: The assessment results refer to: The feature sparsity index obtained by feature sparsity evaluation and the feature stability index obtained by feature stability evaluation are performed.
5. The method for detecting defective grains by combining fingerprint and finite element analysis according to claim 4, characterized in that: The method of selecting the key features of automatic screening of defective grains refers to: Fuzzy logic is used to infer the key features of automatic screening of defective grain based on feature sparsity index and feature stability index.
6. The method for detecting defective grains by combining fingerprint and finite element analysis according to claim 5, characterized in that: Methods for automatically screening key characteristics of defective grain include: L1 regularization feature screening method and recursive feature elimination feature screening method.
7. The method for detecting defective grains by combining fingerprint and finite element analysis according to claim 6, characterized in that: The logic for obtaining the feature sparsity index is: From the set of chemical and physical modal features retained after the initial mutual information screening, identify the features whose mutual information value with the target label is lower than the average mutual information value of all features and record them as low-importance features; Count the number of all low-importance features as the total number of invalid or weakly relevant features in the current feature set, and at the same time obtain the number of all retained features as the total number of features for this round of feature evaluation; Calculate the degree of dispersion of the mutual information values of all currently retained features among different features, and measure the overall information volatility in the form of standard deviation, which is used as an important parameter to measure the uniformity of feature importance distribution; The ratio between the number of low-importance features and the total number of features is nonlinearly enhanced to highlight the impact of the proportion of redundant features on the model quality. At the same time, the mutual information volatility parameter is introduced into the calculation system to comprehensively reflect the impact of feature contribution differences on the selection of feature screening strategies. The two are multiplied together to form the final sparsity index. In the calculation process, in order to adapt to different model requirements and feature structure differences, two constant factors for adjusting sensitivity are set to control the amplification weights of the redundant ratio and volatility respectively.
8. The method for detecting defective grains by combining fingerprint and finite element analysis according to claim 7, characterized in that: The logic for obtaining the characteristic stability index is: From the chemical mode and physical mode feature sets that have been screened by mutual information, conduct multiple cross-validations or use different data partitioning strategies to perform repeated feature screening, and record all the features that are retained in the final feature set in each screening; Count the number of features that are consistently retained across all screening rounds as the "number of stable selected features," i.e., the core features that consistently maintain discriminative power across all repeated experiments. Also, calculate the average number of features selected in each screening process as the "total number of selected features." For features that are stably selected, their corresponding importance scores in multiple screenings are counted, and the ratio between the standard deviation of these scores and the average is calculated as the "score variation ratio" to measure the degree of fluctuation in feature importance. The ratio between the "number of stable selected features" and the "total number of selected features" is used as the first indicator, indicating the consistency of the feature being repeatedly selected under multiple experimental conditions. At the same time, "one minus the score variation ratio multiplied by the fluctuation adjustment factor" is used as the second indicator. The two parts are multiplied together to obtain the final feature stability index. During the calculation process, an exponential enhancement factor is set to adjust the response strength of the stability index to the proportional part.
Citation Information
Patent Citations
Citrus internal and external quality detection and grading method and system based on deep learning
CN119339135A
Key information extraction method and system based on multi-modal model
CN119892215A