Image analysis-based method and system for evaluating quality of wood-based forage silage
By decoupling the characteristics of woody forage silage through causal inference and optimal transport theory, the problem of feature coupling caused by differences in raw material types was solved, achieving high accuracy and reliability in the assessment of new raw material types and improving the generalization ability of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUIZHOU UNIV
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-10
AI Technical Summary
Existing computer vision technology for evaluating the quality of woody silage suffers from feature coupling due to differences in raw material types, leading to a decline in the model's performance in evaluating new, unseen raw material types and failing to meet general requirements.
The causal relationship between fermentation-related initial visual features and common visual base features is analyzed by causal inference method. Causal decoupling guidance parameters are generated, a feature mask is constructed using spatial weighting matrix to shield the base signal, and the feature distribution is aligned to the common reference distribution using optimal transmission theory. The decoupled fermentation features are then output for quality assessment.
It significantly improved the accuracy and reliability of the evaluation of silage samples of new raw material types, reduced the risk of overfitting, and established a quality evaluation model with high generalization performance.
Smart Images

Figure CN121582255B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital image processing technology, and more specifically, to a method and system for evaluating the quality of woody silage based on image analysis. Background Technology
[0002] Analyzing the external visual information of materials using computer vision technology to infer their internal quality or state is an important research direction in the field of image processing and pattern recognition. The typical practical approach is as follows: acquire digital images of the object under visible light or a specific wavelength using image acquisition equipment; then, use image processing algorithms to extract its global or local visual features such as color, texture, and shape; finally, build a statistical model or data-driven model based on these features to predict one or more quality-related indicators. This technical approach has been applied in the quality inspection process of some standardized industrial products due to its advantages of being non-contact and quick to implement, and is regarded as a potential auxiliary or alternative to many traditional destructive and time-consuming laboratory analysis methods.
[0003] When applied to the quality assessment of woody forage silage, the significant interspecific differences in leaf texture, stem structure, and inherent color among various woody plants used as raw materials (such as paper mulberry and mulberry) result in distinctly different visual bases even without processing. Under current technological frameworks, visual features extracted from images of silage materials form a mixed feature profile where strong signals of raw material species identity are coupled with weak signals of fermentation process status. This deep feature confusion and coupling causes the decision boundary of the assessment model trained on it to heavily depend on the distribution of raw material species in the training data. The direct consequence is a significant decrease in the model's assessment performance for new raw material species samples that are unseen in the training set or have significantly different visual bases. This makes it impossible to meet the universality requirement of woody forage silage quality assessment for different raw material species, thus limiting the applicability of existing computer vision technology in the field of woody forage silage quality assessment. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of the prior art, the present invention provides a method and system for evaluating the quality of woody forage silage based on image analysis to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] Image analysis-based methods for evaluating the quality of woody forage silage include:
[0007] S1. Obtain digital images of the woody silage material to be evaluated and the corresponding raw material type labeling information;
[0008] S2. Cluster the raw material type labeling information to obtain several raw material type clusters, and extract the common visual basis features and fermentation-related initial visual features of the digital images based on each raw material type cluster.
[0009] S3. Analyze the causal relationship between fermentation-related initial visual features and common visual base features based on the causal inference method, and generate causal decoupling guidance parameters according to the direction and strength of the causal relationship;
[0010] S4. Analyze the heterogeneity of the causal decoupling guidance parameters in the image space and generate a spatial weighting matrix;
[0011] S5. Use a spatial weighting matrix to weight the common visual basis features to construct a feature mask. Apply the feature mask to mask the basis signals of the fermentation-related initial visual features to obtain the features after the basis signals are masked.
[0012] S6. The source distribution is constructed by shielding the base signal characteristics under each raw material cluster. Based on the optimal transmission theory, each source distribution is aligned to a common reference distribution. The aligned characteristics are output as decoupled fermentation characteristics, and the quality assessment is completed based on the decoupled fermentation characteristics.
[0013] Furthermore, S1 includes:
[0014] Digital images of woody silage were acquired using image acquisition equipment under controlled lighting conditions.
[0015] By identifying the raw materials or querying production records, the material type labeling information corresponding to the digital image can be determined.
[0016] Furthermore, S2 includes:
[0017] Clustering algorithms are used to process the raw material type labeling information, grouping digital images with the same raw material type labeling information into the same set to form raw material type clusters;
[0018] For each raw material cluster, an image feature extraction algorithm is used to analyze all digital images within the raw material cluster. Common color and texture features among the images within the raw material cluster are extracted as common visual base features. At the same time, color changes and texture structure features related to silage fermentation quality are extracted as fermentation-related initial visual features.
[0019] Furthermore, S3 includes:
[0020] Based on the conditional independence test method, we determine whether the change of common visual base features is independent of the change of fermentation-related initial visual features, thereby identifying the causal direction from common visual base features to fermentation-related initial visual features.
[0021] Meanwhile, regression analysis was used to quantify the contribution of shared visual base features to changes in fermentation-related initial visual features, which was then used as a measure of the strength of causal effects.
[0022] The direction and intensity of causal action are combined to generate causal decoupling guidance parameters to guide subsequent decoupling.
[0023] Furthermore, identifying the causal direction through the conditional independence test method includes: calculating the joint conditional probability distribution of the shared visual base features and the fermentation-related initial visual features while controlling for other image feature variables; performing hypothesis testing by comparing the joint conditional probability distribution with the expected distribution under the independence assumption; and determining whether there is a causal direction from the shared visual base features to the fermentation-related initial visual features based on the test results.
[0024] Furthermore, S4 includes:
[0025] The causal decoupling guide parameters are mapped to the pixel coordinate space of the digital image to form a parameter space distribution map;
[0026] By traversing the parameter space distribution map through a sliding window, the local statistics of the causal decoupling guidance parameters within each window are calculated to quantify local heterogeneity.
[0027] The local heterogeneity values of each window are normalized, and the normalized local heterogeneity values are then filled into the corresponding positions of a matrix with the same size as the digital image to generate a spatial weighted matrix.
[0028] Furthermore, S5 includes:
[0029] The spatial weighted matrix is multiplied element-wise with the common visual basis features at the corresponding spatial locations to generate a spatially adaptive feature mask.
[0030] The spatially adaptive feature mask is applied to the fermentation-related initial visual features. The base signal component represented by the spatially adaptive feature mask is subtracted from the fermentation-related initial visual features to obtain the base signal masked features.
[0031] Furthermore, S6 includes:
[0032] The feature set of all base signals within each raw material cluster after masking is regarded as a source distribution, and the mean distribution of all source distributions is defined as a common reference distribution;
[0033] Calculate the optimal transport plan from each source distribution to the common reference distribution. The optimal transport plan defines the mapping relationship in which each feature vector in the source distribution should move to the corresponding position in the common reference distribution.
[0034] Based on the optimal transmission plan, spatial transformation correction is performed on the features of all base signals after shielding, and the spatially transformed and corrected features are output as decoupled fermentation features.
[0035] By inputting decoupled fermentation characteristics into pre-trained quality assessment rules, the quality assessment results of woody forage silage are obtained and output.
[0036] Furthermore, the pre-trained quality assessment rules are obtained by: collecting woody forage silage samples of known quality grades and their corresponding decoupled fermentation characteristics as a training set; learning the mapping rules between decoupled fermentation characteristics and quality grades through machine learning algorithms; and storing the learned mapping rules as quality assessment rules.
[0037] On the other hand, the present invention provides an image analysis-based system for evaluating the quality of woody forage silage, comprising:
[0038] The image acquisition module is used to acquire digital images of the woody silage material to be evaluated and the corresponding raw material type labeling information;
[0039] The feature extraction module is used to cluster the raw material type labeling information to obtain several raw material type clusters, and extract the common visual basis features and fermentation-related initial visual features of the digital images based on each raw material type cluster.
[0040] The parameter generation module is used to analyze the causal relationship between fermentation-related initial visual features and common visual base features based on the causal inference method, and generate causal decoupling guidance parameters according to the direction and strength of the causal relationship.
[0041] The matrix generation module is used to analyze the heterogeneity of causal decoupling guidance parameters in the image space and generate a spatial weighted matrix.
[0042] The signal masking module is used to weight the common visual basis features using a spatial weighting matrix to construct a feature mask. The feature mask is then applied to the fermentation-related initial visual features to mask the basis signals, thus obtaining the features after the basis signals are masked.
[0043] The quality assessment module is used to construct source distributions based on the masked features of the base signals under each raw material cluster. Based on the optimal transmission theory, each source distribution is aligned to a common reference distribution. The aligned features are output as decoupled fermentation features, and the quality assessment is completed based on the decoupled fermentation features.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] 1. By introducing causal inference and optimal transport theory into the image feature processing workflow, common visual basis features of raw material clusters are extracted, and causal inference is used to analyze their causal relationship with fermentation-related initial visual features. Then, a spatial weighted matrix is generated and a feature mask is constructed. In the image feature space, signals from different raw materials' inherent visual basis are identified and suppressed in a targeted and adaptive manner. This effectively weakens the interference components dominated by raw material differences in the signal composition of the decoupled fermentation features used for quality assessment, while highlighting the information components that truly reflect the fermentation process and state. It can focus more on visual patterns that are universally related to fermentation quality, thus significantly improving the accuracy and reliability of the assessment of new raw material silage samples that have never appeared in the training data or have very different visual basis. This solves the problem of insufficient model generalization ability caused by feature coupling in traditional methods.
[0046] 2. By utilizing optimal transport theory to align the feature distributions of different raw material clusters to a common reference distribution, this operation is mathematically equivalent to performing a normalized geometric transformation in the feature space. It systematically corrects the overall offset of feature distributions caused by different raw material types. This not only ensures the consistency of input features for subsequent quality assessment rules at the algorithm level, but also reduces the risk of overfitting the assessment model to training samples of specific raw material types from the root of data distribution. Based on the analysis and transformation of digital images, its final output of decoupled fermentation features is a purer and more standardized image-derived data representation, which provides crucial high-quality input for constructing a quality assessment model with high generalization performance. Attached Figure Description
[0047] Figure 1 This is a flowchart of the image analysis-based method for evaluating the quality of woody forage silage according to the present invention;
[0048] Figure 2 This is a schematic diagram of the image analysis-based silage quality assessment system for woody forage according to the present invention. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Example 1: Figure 1 This invention presents a method for evaluating the quality of woody forage silage based on image analysis, comprising:
[0051] S1. Obtain digital images of the woody silage material to be evaluated and the corresponding raw material type labeling information;
[0052] S2. Cluster the raw material type labeling information to obtain several raw material type clusters, and extract the common visual basis features and fermentation-related initial visual features of the digital images based on each raw material type cluster.
[0053] S3. Analyze the causal relationship between fermentation-related initial visual features and common visual base features based on the causal inference method, and generate causal decoupling guidance parameters according to the direction and strength of the causal relationship;
[0054] S4. Analyze the heterogeneity of the causal decoupling guidance parameters in the image space and generate a spatial weighting matrix;
[0055] S5. Use a spatial weighting matrix to weight the common visual basis features to construct a feature mask. Apply the feature mask to mask the basis signals of the fermentation-related initial visual features to obtain the features after the basis signals are masked.
[0056] S6. The source distribution is constructed by shielding the base signal characteristics under each raw material cluster. Based on the optimal transmission theory, each source distribution is aligned to a common reference distribution. The aligned characteristics are output as decoupled fermentation characteristics, and the quality assessment is completed based on the decoupled fermentation characteristics.
[0057] S1. Obtain digital images of the woody silage material to be evaluated and the corresponding raw material type labeling information. The specific implementation is as follows:
[0058] Digital images of woody silage are acquired using image acquisition equipment under controlled lighting conditions. This controlled lighting environment refers to stable, uniform, and repeatable lighting conditions created by artificial light sources within a closed or semi-closed photography box or specific shooting area. Artificial light sources may include, for example, LED flat panel lights or diffused light sources with a color temperature of 5000K to 6500K. Light intensity is controlled, for example, between 500 lux and 1500 lux. This range ensures image clarity and avoids overexposure or underexposure, and is verified using a lux meter at multiple points on the material placement plane to ensure that the lighting uniformity error across the entire shooting area is less than 15%. Image acquisition equipment may include, for example, an industrial camera or a high-definition digital camera with adjustable manual mode. The camera is fixed on a bracket directly above the material, with the lens optical axis perpendicular to the material placement plane. During shooting, camera parameters are set to fixed values, such as an aperture of F8 to ensure sufficient depth of field, a shutter speed of 1 / 125 second to avoid motion blur caused by manual operation, and an ISO of 100 to reduce image noise. The material is placed on a neutral gray background, the color of which contrasts sharply with the material color; for example, a light gray background with the Munsell color chart number N7 is used. Each time data is collected, ensure that each portion of woody silage is spread evenly on the background with no significant overlap, then trigger the camera to capture a color digital image with a resolution of at least 1920 pixels by 1080 pixels. This digital image is stored in a standard RGB color space format such as JPEG or PNG, and a unique filename containing a timestamp and device number is automatically generated. If the lighting conditions do not meet the preset range or the image is blurry, the light source or camera parameters are readjusted before collecting data again.
[0059] By identifying the raw materials or querying production records, the corresponding raw material type labeling information is determined for each digital image, and this labeling information is then stored in association with the digital image. Specifically, identifying the raw materials involves experienced feeders or agronomists identifying the same batch of woody silage material photographed during the same time period when the digital images are collected. Identification is based on experiential knowledge, including the residual morphology of leaves, stem thickness and internode characteristics, and the overall color tendency of the material. For example, identifying the material as primarily composed of paper mulberry leaves and tender branches, or primarily composed of mulberry branches, etc. The identification result is recorded as the raw material type labeling information. This information can be a textual description, such as "paper mulberry," or a predefined type code, such as 01 representing paper mulberry and 02 representing mulberry. If the on-site identification cannot determine the type, it is marked as unknown and recorded for later review. Querying production records refers to the detailed agricultural records available for each batch of silage material in large-scale silage production before the raw materials are added to the fermentation container. The archive exists in the form of an electronic database or paper record sheet, recording information such as the harvest date, plot location, and composition ratio of raw material plant species for that batch of materials. By tracing the batch number to which the material to be evaluated belongs, the corresponding raw material composition can be retrieved from the production record database and simplified into a primary or comprehensive raw material category label. Whether obtained through identification or querying, the raw material category label must be uniquely and unambiguously associated with the corresponding digital image. Practices for associated storage include directly writing the raw material category label as metadata into the EXIF information field of the digital image file when storing it; or recording the digital image filename and the raw material category label in the same row of a structured relational database table or a spreadsheet, where the database table fields include at least an image file storage path field and a raw material category label field. This explicit associated storage ensures that in any subsequent processing step, the corresponding raw material category label can be accurately indexed from the image file, providing an accurate data foundation for subsequent cluster analysis based on raw material category. The entire acquisition process ensures a strict correspondence between visual data and semantic annotations, forming a reliable input starting point for all subsequent image analysis logic.
[0060] S2. Cluster the raw material type labeling information to obtain several raw material type clusters. Extract the common visual basis features and fermentation-related initial visual features of the digital images based on each raw material type cluster. The specific implementation is as follows:
[0061] Clustering algorithms are used to process the raw material type labeling information, grouping digital images with the same raw material type labeling information into the same set, forming raw material type clusters. The clustering algorithm here refers to a direct grouping method based on existing semantic labels. Its input data consists of all digital images acquired and associated with the data in step S1, along with their corresponding raw material type labeling information. The raw material type labeling information is in text or encoded form, such as tree or 01. During processing, the system reads the raw material type labeling information string associated with each digital image file and assigns all digital images with identical raw material type labeling information strings to the same group. For example, all digital images with the raw material type labeling information string "tree" are grouped into one set, and all digital images with the raw material type labeling information string "mulberry" are grouped into another set. Each such set constitutes a raw material type cluster. If the raw material type labeling information of a digital image is missing or marked as unknown, the image is temporarily placed in a separate unprocessed cluster, to be supplemented and reclassified later through manual verification of the production records in step S1 or re-identification. This grouping process does not involve calculating inter-sample distances or iterative cluster centers; its grouping is based on the explicitly provided semantic labels. Ultimately, the output consists of several raw material clusters. Each raw material cluster is a data set that logically contains multiple digital images and their common raw material labeling information. This set is organized in computer memory as a list data structure or stored in a database through a field that identifies the raw material cluster number.
[0062] For each raw material cluster, an image feature extraction algorithm is used to analyze all digital images within the cluster. Common color and texture features shared among images within the cluster are extracted as shared visual base features. Simultaneously, color changes and texture structure features related to silage fermentation quality are extracted as initial fermentation-related visual features. In practice, for each established raw material cluster, each digital image file within that cluster is read sequentially. The image feature extraction algorithm consists of two parts: color feature extraction and texture feature extraction. Color feature extraction first converts each digital image from the RGB color space to the HSV color space to separate luminance and color information. Then, the global color histogram of each digital image in the H, S, and V channels in the HSV space is calculated. For example, the numerical range of each channel is evenly divided into 16 intervals, and the number of pixels in each interval is counted and normalized, resulting in a 48-dimensional color histogram feature vector for each image. The 16 intervals are chosen to control computational complexity while ensuring feature discrimination capability; the number of intervals can be adjusted, for example, to 8 or 32, depending on experimental results. For all digital images within a raw material cluster, the 48-dimensional color histogram feature vector corresponding to each image is averaged element-wise. This means the average of all vectors in the first dimension is used as the first dimension of the resulting vector. This process is repeated until a 48-dimensional average color histogram vector is obtained. This average color histogram vector represents the common color statistical distribution characteristics of all images within the raw material cluster and is defined as the color basis part of the common visual basis features. Texture feature extraction involves converting each digital image to grayscale and then calculating the gray-level co-occurrence matrix (GLCM). Calculating the GLCM requires pre-setting parameters for the calculation direction and pixel distance pair, such as four directions (0°, 45°, 90°, and 135°) and a pixel distance (e.g., 1 pixel). The direction selection is based on the isotropic nature of the texture, and the pixel distance of 1 is used to capture the most subtle texture structures. For each direction and distance pair, a GLCM is calculated, and then four statistical measures—contrast, energy, homogeneity, and correlation—are extracted from each matrix. For a single digital image, the average of the same statistic calculated in four directions is taken. For example, the contrast values in the four directions are added together and divided by 4, resulting in a texture feature vector containing four statistics. For a material category cluster, the element-wise average of the 4D texture feature vectors of all digital images within the cluster is taken to obtain a 4D average texture feature vector. This average texture feature vector represents the shared texture statistical characteristics among images within the material category cluster and is defined as the texture basis part of the shared visual basis features. Finally, the 48-dimensional average color histogram vector and the 4-dimensional average texture feature vector are concatenated sequentially to form a 52-dimensional vector, which is the shared visual basis feature of the material category cluster.
[0063] The extraction of initial visual features related to fermentation is also performed for each raw material cluster, but the features of interest are directly related to changes in material quality caused by the fermentation process. Color change feature extraction focuses on indicators reflecting typical color changes during silage fermentation. Specifically, for each digital image within a raw material cluster, the ratio of the G channel pixel value to the R channel pixel value in the RGB color space is first calculated (G divided by R). The number of pixels in the entire image whose G divided by R ratio is lower than a specific color change threshold is counted, and this number is divided by the total number of pixels in the image to obtain a proportion of pixels with a low G divided by R ratio. The method for obtaining and setting the color change judgment threshold is as follows: A dataset consisting of images of woody silage samples of known quality grades is collected. The known quality grades are determined through chemical analysis methods, such as measuring pH value, the proportion of ammonia nitrogen to total nitrogen, and the content of organic acids. The material areas in each sample image that have visually turned noticeably yellowish-brown are manually labeled. The statistical distribution of the ratio of green channel values to red channel values of pixels within all labeled areas is calculated. The statistical distribution, such as the 10th percentile or the category boundary value determined through cluster analysis, is set as the color change judgment threshold. Simultaneously, the images are converted to the Lab color space, and the arithmetic mean of all pixel values in the b-channel is calculated to obtain a b-channel mean, which is used to quantify the yellowness of the image. For a raw material cluster, instead of averaging the above values of each image within the cluster, the low-ratio pixel proportions (G / R) and the b-channel mean of each image are recorded separately, forming a set of low-ratio pixel proportions (G / R) and a set of b-channel means for that cluster. The extraction of texture structure features focuses on the manifestation of material tissue softening or structural decomposition caused by fermentation in the texture. When calculating the texture statistics for each image using the aforementioned gray-level co-occurrence matrix, an additional entropy statistic is calculated. An increase in entropy typically indicates an increase in texture complexity. The gray-level parameter used for entropy calculation is set to, for example, 16 levels. Simultaneously, a local binary mode algorithm is applied to the gray-level image to calculate its histogram in uniform mode, and the variance of this histogram is calculated as a measure of texture uniformity; a decrease in variance indicates that the texture is becoming more uniform. For each image within a raw material cluster, its entropy value and texture uniformity variance value are recorded. Similarly, for a raw material cluster, the entropy values of all images within the cluster are collected to form an entropy set, and the texture uniformity variance values are collected to form a texture uniformity variance set. Finally, the set of low-ratio pixel proportions (G divided by R), the set of b-channel mean values, the entropy set, and the texture uniformity variance set corresponding to a raw material cluster are collectively defined as the initial visual features related to fermentation for that raw material cluster. These feature sets serve as direct inputs for subsequent causal inference analysis. The entire feature extraction process separated stable common base components and variable components that are dynamically related to the fermentation state from visual data of the same raw material type, providing clear feature inputs for subsequent causal decoupling.
[0064] S3. Analyze the causal relationship between fermentation-related initial visual features and shared visual base features based on causal inference methods, and generate causal decoupling guidance parameters according to the direction and strength of the causal relationship. Specifically, the implementation is as follows:
[0065] Based on the conditional independence test method, this study determines whether changes in the shared visual basis features are independent of changes in the fermentation-related initial visual features, thereby identifying the causal direction from the shared visual basis features to the fermentation-related initial visual features. Specifically, the input data consists of the shared visual basis features and fermentation-related initial visual features obtained from step S2 for each raw material cluster. For a specific raw material cluster, its shared visual basis features are a 52-dimensional vector, while its fermentation-related initial visual features are a set of multiple values, such as the set of low-ratio pixel proportions of G divided by R, the set of b-channel mean values, the set of entropy values, and the set of texture uniformity variance. To perform the conditional independence test, the analysis needs to focus on the causal relationship between each specific feature dimension of the shared visual basis features and the fermentation-related initial visual features. For example, the dimension of low-ratio pixel proportions of G divided by R in the fermentation-related initial visual features is selected for analysis. In practice, the continuously valued features are first discretized to facilitate probability estimation. For each dimension of the 52-dimensional common visual basis features, based on its numerical distribution across all samples in the raw material category cluster, it is divided into, for example, three intervals: a low-value interval, a medium-value interval, and a high-value interval. This division can be based on equal-frequency binning or data quantiles. Similarly, all values in the set of low-ratio pixel proportions (G / R) are also divided into three intervals. Then, conditional probability distributions are calculated while controlling for other image feature variables. Controlling other image feature variables here means that, during statistical analysis, other visual feature values of each sample image within the raw material category cluster are used as grouping conditions, for example, grouping according to the interval where the mean of the b-channel of the sample lies. Within the same group, the relationship between the common visual basis features and the target fermentation features is calculated. Next, the joint conditional probability is calculated when, under specific grouping conditions, the common visual basis features are in a certain interval and the low-ratio pixel proportions (G / R) are simultaneously in a certain interval. Simultaneously, the expected joint probability under the independence assumption, i.e., the product of their respective marginal conditional probabilities, is calculated under the same grouping conditions. The chi-square test is used as a hypothesis testing method to compare the observed joint conditional probability with the expected joint probability and calculate the chi-square test statistic. The chi-square test requires setting a significance level threshold, for example, 0.05. A significance level threshold of 0.05 is a standard in statistical hypothesis testing, used to balance the risks of Type I and Type II errors. If the p-value corresponding to the calculated chi-square test statistic is less than 0.05, the null hypothesis of independence is rejected, indicating a statistical dependency between the shared visual base feature dimension and the fermentation-related initial visual feature dimension under this grouping condition. If the p-value is greater than or equal to 0.05, the null hypothesis of independence cannot be rejected, indicating that no statistical dependency was detected under this condition.To determine the causal direction, further analysis of the temporal or logical sequence is needed. In this method, based on the fact that the basic features reflect the inherent properties of the raw materials, and their changes may precede and affect the fermentation process features, the statistical dependency is interpreted as a causal direction from shared visual basic features to fermentation-related initial visual features. The above test is performed on each dimension of the shared visual basic features and each dimension of the fermentation-related initial visual features, and all relationship pairs determined to have a causal direction are recorded.
[0066] Simultaneously, regression analysis is used to quantify the contribution of shared visual base features to changes in fermentation-related initial visual features, thus representing the strength of the causal effect. For each pair of features determined to have a causal direction in the conditional independence test, a univariate linear regression model is established for quantification. A certain dimension of the shared visual base features is used as the independent variable X, and the corresponding value of the fermentation-related initial visual feature dimension in all samples within the cluster is used as the dependent variable Y. The goal of the regression analysis is to fit a linear equation Y = β × X + C, where β is the regression coefficient and C is a constant term. The least squares method is used to estimate the value of the regression coefficient β, which solves for the optimal β by minimizing the sum of squares of the differences between the predicted and actual values of all sample points. The absolute value of the obtained regression coefficient β directly reflects the average change in the dependent variable Y for every unit change in the independent variable X, representing the strength of the causal effect. To eliminate the influence of differences in dimensions between different features on the comparability of strength, the regression coefficients need to be standardized. The standardization method involves first standardizing all sample values for both the independent variable X and the dependent variable Y. Specifically, for each feature dimension, the mean μ and standard deviation σ of all sample values are calculated. Then, each sample value is subtracted from μ and divided by σ to obtain the standardized value. The standardized data is then used to perform the regression analysis again, and the resulting standardized regression coefficient is the final causal strength value. This value is dimensionless; a larger absolute value indicates a stronger causal effect. For each pair of co-occurring feature dimensions with a causal direction, such a standardized regression coefficient is calculated as the strength of their causal effect.
[0067] The causal direction and causal strength are combined to generate causal decoupling guidance parameters to guide subsequent decoupling. Specifically, this merging operation involves creating a structured parameter table. This table contains multiple entries, each corresponding to a pair of co-causal feature dimensions whose causal direction was confirmed in the conditional independence test. Each entry records four pieces of information: the dimension number of the shared visual base feature, the dimension name of the fermentation-related initial visual feature, the causal direction identifier, and the causal strength value. The causal direction identifier is a binary variable; for example, 1 indicates a direction from the shared visual base feature to the fermentation-related initial visual feature. The causal strength value is the standardized regression coefficient mentioned above. Furthermore, when generating the final guidance parameters, a contribution level screening threshold is set to filter out feature pairs that, while having a causal relationship, have a weak effect. The contribution level screening threshold is set based on the distribution of all calculated causal strength values. For example, setting the threshold to the median value (the middle value after sorting all strength values by absolute value) is an empirical screening strategy that retains at least half of the important causal relationships. This can also be adjusted according to actual needs. The method is based on retaining feature relationships with a contribution level above the medium level. Only when the absolute value of the causal effect strength of a feature pair is greater than the contribution level screening threshold will its entry be retained in the final causal decoupling guidance parameter table. If the strength values of all feature pairs are very close, all entries are retained by default. This parameter table is the generated causal decoupling guidance parameter, which will be passed as a key input to the subsequent step S4 to analyze heterogeneity in the image space and guide the construction of the spatial weighting matrix. Through the above process, the causal cognition obtained from statistical inference is transformed into specific, computable decoupling guidance parameters, thereby applying causal theory to practical image feature decoupling tasks.
[0068] S4. Analyze the heterogeneity of the causal decoupling guidance parameters in the image space and generate a spatial weighting matrix. The specific implementation is as follows:
[0069] The causal decoupling guidance parameters are mapped to the pixel coordinate space of the digital image, forming a parameter space distribution map. The input data consists of the causal decoupling guidance parameter table generated in step S3 and the size information of the original digital image corresponding to the feature extraction in step S2. Each entry in the causal decoupling guidance parameter table records the feature dimension pairs with causal relationships and their intensity values. The key to the mapping is to establish the correspondence between each feature dimension and the pixel region of the image space. For spatial features in the common visual base features and fermentation-related initial visual features, such as texture features calculated from local regions of the image, they are inherently associated with specific blocks in the image. In specific implementation, the original spatial location or region of the image reflected by each feature dimension is determined according to the specific implementation of the feature extraction algorithm in step S2. For example, if step S2 uses a method of uniformly dividing the image into, for example, 16 grids of 4 rows by 4 columns and extracting texture features within each grid, then each texture feature dimension uniquely corresponds to an image grid. The feature dimensions involved in each entry in the causal decoupling guidance parameter table are filled with their causal effect intensity values into their corresponding image pixel coordinate regions according to their predefined spatial correspondence. If a feature dimension corresponds to an image block rather than a single pixel, then the causal effect strength value is assigned to all pixels within that block. After processing all entries, a two-dimensional image with the same dimensions as the original digital image is obtained, where each pixel location has a causal effect strength value; this image is defined as the parameter space distribution map. For global statistical features that do not directly correspond to spatial locations, such as the color histogram feature of the entire image, their causal effect strength values are uniformly assigned to all pixels in the image.
[0070] By traversing the parameter space distribution map through a sliding window, local statistics of causal decoupling-guided parameters within each window are calculated to quantify local heterogeneity. The sliding window is a fixed-size rectangular region that moves across the parameter space distribution map. The window size is determined by considering the scale of the heterogeneity analysis; for example, the window side length can be set to 1 / 10 of the image width and 1 / 10 of the image height, but a minimum size limit is set, such as a minimum side length of 16 pixels. The window side length of 1 / 10 of the image width and height is determined through a grid search on a training image set containing multiple raw materials, ensuring that the local heterogeneity calculated at this scale has the strongest correlation with subsequent quality assessment results. The minimum size limit of 16 pixels is an empirical value based on image resolution and ensuring that the window contains enough pixels for effective statistical calculations. The sliding step size is set to half the window side length; for example, if the window side length is 32 pixels, the sliding step size is 16 pixels to ensure sufficient coverage and analysis of the image space. Local statistics are used to quantify the dispersion of causal effect strength values within the window, i.e., local heterogeneity. In the specific calculation, for each position covered by the sliding window, the causal effect intensity values of all pixels within the window are extracted, and the variance of these intensity values is calculated as a measure of local heterogeneity. The variance calculation process is as follows: first, the arithmetic mean of all intensity values within the window is calculated; then, the square of the difference between each intensity value and the mean is calculated; finally, all squared differences are summed and divided by the total number of pixels within the window minus 1. The larger the variance value, the more uneven the distribution of causal relationship intensity in that local region, and the higher the heterogeneity. After traversing the entire parameter space distribution map, each window position will generate a local heterogeneity value, which represents the degree of spatial non-uniformity of causal influence in the local image region centered on that window.
[0071] The local heterogeneity values of each window are normalized, and the normalized local heterogeneity values are then filled into the corresponding positions of a matrix with the same size as the digital image to generate a spatial weighting matrix. Since sliding windows overlap, the same pixel may be covered by multiple windows. Therefore, a representative local heterogeneity value needs to be determined for each pixel location. In practice, for each pixel in the image, all sliding windows covering that pixel are identified, and the arithmetic mean of the local heterogeneity values calculated for these windows is taken as the initial spatial weight value for that pixel location. Next, the initial spatial weight values obtained for all pixel locations are normalized so that their values fall between 0 and 1. The normalization uses a minimum-maximum normalization method: first, the maximum and minimum values of the initial spatial weight values for all pixels are found; then, for each pixel's weight value, the minimum value is subtracted, and the result is divided by the difference between the maximum and minimum values. To prevent the normalization result from being too unstable due to the maximum and minimum values being too close, a normalization range stability threshold is set. For example, when the difference between the maximum and minimum values is less than 0.01, the weight value of all pixels is directly set to 0.5. The normalization range stability threshold of 0.01 is a very small positive number set to prevent the numerical calculation from being unstable due to an excessively small denominator. The basis for setting the normalization range stability threshold is to ensure that the denominator has sufficient numerical stability and avoid unstable calculations such as division by zero or near division by zero. After normalization, each pixel position corresponds to a final weight value between 0 and 1. A two-dimensional matrix with the same number of rows and columns as the original digital image is created, and the final weight value of each pixel position is filled into the corresponding row and column index positions in the matrix. This two-dimensional matrix filled with normalized weight values is the generated spatial weighting matrix. In this matrix, regions with weight values closer to 1 indicate higher heterogeneity in causal analysis and require greater attention or adjustment in subsequent decoupling operations; regions with weight values closer to 0 indicate lower heterogeneity and more homogeneous characteristics. This spatial weighting matrix will be passed as a direct input parameter to the subsequent step S5 to guide the construction of the feature mask.
[0072] S5. A feature mask is constructed by weighting the common visual basis features using a spatial weighting matrix. This feature mask is then applied to mask the basis signals of the initial visual features related to fermentation, yielding the masked features. The specific implementation is as follows:
[0073] The spatially weighted matrix and the shared visual basis features are multiplied element-wise at their corresponding spatial locations to generate a spatially adaptive feature mask. The input data consists of the spatially weighted matrix generated in step S4 and the shared visual basis features for each material cluster obtained in step S2. The spatially weighted matrix is a two-dimensional matrix with, for example, H rows and W columns, exactly the same as the pixel height and width of the original digital image. The value of each element in the matrix is a normalized weight value, ranging from 0 to 1. The shared visual basis features are a 52-dimensional feature vector that represents the color and texture statistical characteristics shared by all images within the entire material cluster. This feature vector itself does not possess explicit spatial location information. To perform spatial location mapping, this 52-dimensional shared visual basis feature vector needs to be spatially expanded. Specifically, a three-dimensional data structure with H rows, W columns, and 52 channels is created. This structure consists of 52 stacked two-dimensional planar layers, each corresponding to one dimension of the shared visual basis feature vector. For each layer, its values at all row and column positions are initialized to the scalar values of the corresponding feature dimension in the original 52-dimensional vector. This transforms a global statistical feature vector into a three-dimensional feature map with spatial dimensions, its width and height matching the image size, and its depth being 52. Next, element-wise multiplication is performed. For each spatial location in the three-dimensional feature map—that is, for each specific row index i, column index j, and feature channel index k—the feature value Fijk at that location is extracted, and the corresponding weight values Wij at the same row index i and column index j are extracted from the spatial weighting matrix. The feature value Fijk and the weight value Wij are multiplied to obtain the result value Mijk = Fijk × Wij. This operation iterates through all locations in the three-dimensional feature map, ultimately generating a new three-dimensional data structure with the same dimensions (H rows, W columns, 52 channels), which is the spatially adaptive feature mask. In this feature mask, the values in each dimension of the original common visual basis feature vector are scaled to different degrees at different spatial locations according to the weight values of the spatial weighting matrix. In regions with high weight values, the feature values are largely preserved; in regions with low weight values, the feature values are significantly weakened. If the size of the spatial weighting matrix is inconsistent with the size of the 3D feature map after the expansion of the shared visual basis features, a size verification error is triggered, and the feature extraction region definition in step S2 and the matrix generation process in step S4 are rechecked.
[0074] A spatially adaptive feature mask is applied to the initial visual features related to fermentation. The base signal components represented by the spatially adaptive feature mask are subtracted from the initial visual features related to fermentation to obtain the base signal masked features. The input data consists of the spatially adaptive feature mask generated above and the initial visual features related to fermentation of the same raw material cluster obtained in step S2. The initial visual features related to fermentation are a set of multiple values, such as the set of low-ratio pixel proportions of G divided by R, the set of b-channel mean values, the set of entropy values, and the set of texture uniformity variance. These values originate from each specific digital image sample within the cluster. To perform the signal masking operation, the base signal components represented by the spatially adaptive feature mask need to be aligned and matched with the initial visual features related to fermentation of each sample. The specific implementation is divided into two stages. The first stage is the generation and normalization of the base signal components. The spatially adaptive feature mask is converted back to a representation comparable to the initial visual features related to fermentation. Since the initial visual features related to fermentation are specific to a particular image sample, while the feature mask is spatialized, it is necessary to extract the corresponding signal from the feature mask based on the specific spatial region involved in feature extraction for each sample image in step S2. For example, for the entropy value of texture features in the initial visual features related to fermentation, if the entropy value is calculated from a specific grid region in the image in step S2, then the values of all feature channels are extracted from the spatial location corresponding to the grid region in the feature mask, and these values are used to calculate a scalar value through an aggregation function, such as calculating the arithmetic mean, as the estimated value of the basis signal component for that feature of the sample. For color features, such as the mean of the b channel, since it is a global feature, the values of the corresponding color feature channels are extracted from all spatial locations in the feature mask and their global average value is calculated as the basis signal component. The second stage is to perform signal masking. For each sample within the raw material cluster, for each feature dimension in its initial visual features related to fermentation, the following calculation is performed: the original observation value of the sample in that feature dimension is subtracted from the basis signal component value estimated by the above method corresponding to that feature dimension, and the difference is the value of the feature in that dimension of the sample after basis signal masking. In mathematical terms, if the original value of a sample on feature A is Vorig, and the estimated basis signal component is BA, then the shielded feature value Vshielded = Vorig - BA. This subtraction operation removes the signal component contributed by the common basis characteristics of the raw materials in the estimation, thus making the remaining features more purely reflect the variation related to the specific fermentation state of the sample. After performing the above operation on all fermentation-related initial visual feature dimensions of all samples within the cluster, the basis signal shielded features of that raw material cluster are obtained. It is a new set of features with the same structure as the original fermentation-related initial visual features but with changed values, while its dimensions and number remain unchanged.The masked features of the substrate signal will serve as input data for distribution alignment in subsequent step S6. The entire masking process, through spatially adaptive weighting and subtraction operations, differentiates and weakens the influence of the raw material substrate on fermentation features at different spatial locations, preparing cleaner feature data for subsequent decoupling analysis.
[0075] S6. The source distribution is constructed using the masked features of the base signals under each raw material cluster. Based on optimal transmission theory, each source distribution is aligned to a common reference distribution. The aligned features are output as decoupled fermentation features, and quality assessment is completed based on these features. The specific implementation is as follows:
[0076] The feature set of all base signal masked within each raw material cluster is considered as a source distribution, and the mean distribution of all source distributions is defined as the common reference distribution. The input data is the base signal masked features of all raw material clusters output from step S5. For a specific raw material cluster, its base signal masked features are the feature set obtained after masking all digital image samples within that cluster. Each sample corresponds to a feature vector, which consists of multiple feature dimensions, such as the proportion of low-ratio pixels divided by G / R, the mean of the b-channel, the entropy value, and the variance of texture uniformity. These sample feature vectors are aggregated to form the source distribution of that raw material cluster, which reflects the variation of its fermentation-related characteristics among different samples after removing the common base influence. Assume there are several raw material clusters, and the source distribution of each cluster contains a number of sample feature vectors. Next, the common reference distribution is calculated. Specifically, all sample feature vectors in the source distributions of all raw material clusters are merged together to form a global set containing all samples. Calculate the arithmetic mean of each feature dimension in the global set, which is achieved by summing the values of all samples in that dimension and dividing by the total number of samples. This yields a mean vector with the same dimensions as the feature vector of a single sample. The probability distribution represented by this mean vector is the common reference distribution, which mathematically represents a virtual, centralized feature distribution as a common target for alignment of all source distributions. In actual computation, the common reference distribution is not directly represented as a specific set of samples, but rather as the parameters of a multivariate probability distribution with a specific mean.
[0077] The optimal transport plan is calculated from each source distribution to the common reference distribution. The optimal transport plan defines the mapping relationship where each feature vector in the source distribution should move to its corresponding position in the common reference distribution. The calculation of the optimal transport plan aims to find a way to move sample points from the source distribution to the common reference distribution with the minimum cumulative movement cost. To achieve this goal, it is first necessary to define the distance metric between sample points in the two distributions. Euclidean distance is used as the distance metric; that is, for a sample feature vector in the source distribution and a virtual point in the common reference distribution, the distance between them is defined as the square root of the sum of the squares of the differences in the corresponding dimensions of the two vectors. Since the common reference distribution does not have concrete sample points but is represented as a distribution centered on the mean vector, a calculation method based on discrete approximation and iterative optimization, such as the Sinkhorn algorithm, is used. The specific implementation includes the following steps. The first step is initialization. For a source distribution, it contains a certain number of sample points. It is assumed that the common reference distribution is approximated by a certain number of support points, which can be set as a proportion of the total number of samples in all source distributions, for example, ten percent. Initially, the locations of these support points can be set by randomly selecting a corresponding number of points from the samples merged from all source distributions. The second step is to construct a cost matrix. The Euclidean distance between each sample point in the source distribution and each support point in the common reference distribution is calculated, forming a cost matrix with the number of rows equal to the number of sample points and the number of columns equal to the number of support points. The value at each position in the matrix represents the distance from a corresponding sample point to a support point. The third step is to perform Sinkhorn iteration to solve for the optimal transport plan. The transport plan is a non-negative matrix with the number of rows equal to the number of sample points and the number of columns equal to the number of support points. The value at each position in the matrix represents the amount of mass transferred from the corresponding sample point to the corresponding support point. The goal of the iteration is to minimize the total transport cost while satisfying two marginal constraints: the total mass emitted from each sample point is equal to the weight of that sample point, usually set to 1 divided by the total number of sample points in the source distribution; the total mass transferred to each support point is equal to the weight of that support point, initially set to 1 divided by the total number of support points. A regularization parameter is introduced during the iteration process, for example, this parameter can be set to 0.1, to smooth the solution and accelerate the calculation. The iterative formula alternately normalizes the rows and columns of the transfer matrix to satisfy edge constraints, multiplying each normalization by an exponential decay factor calculated based on the cost matrix and regularization parameters. An iterative convergence threshold is set; for example, convergence is determined when the maximum absolute value of the difference between corresponding values in the transfer matrix obtained from two consecutive iterations is less than 0.001. This convergence threshold of 0.001 was selected experimentally. When the maximum change in the transfer matrix between two iterations is less than this value, the calculation result is considered sufficiently stable, and further iterations do not significantly improve the result.The convergence threshold is set to ensure that the change in the transmission plan matrix is small enough and the calculation results tend to stabilize. The final transmission matrix is the optimal transmission plan, which indicates how to redistribute the sample points of the source distribution to the support points of the common reference distribution with the minimum cumulative distance cost.
[0078] Based on the optimal transmission plan, spatial transformation correction is performed on the features of all base signal shielded samples. The spatially transformed and corrected features are then output as decoupled fermentation features. The purpose of spatial transformation correction is to adjust the feature vector of each sample in the source distribution according to the mapping relationship defined in the optimal transmission plan, so that its position in the feature space moves towards the direction of the common reference distribution, thereby reducing the systematic offset between different source distributions. Specifically, for a sample feature vector in the source distribution, the indices of the support points to which the sample point is mainly transmitted are found according to the optimal transmission plan matrix. For example, the support points pointed to by the top few data items with the largest values in the row corresponding to the sample in the transmission plan matrix are identified. Then, the weighted average of the coordinates of these support points is calculated, with the weights being the values of the corresponding positions in the transmission plan. The new coordinate vector calculated by the weighted average is the new feature vector of the sample point after spatial transformation correction. This process is equivalent to moving the original sample point towards its target support point, with the magnitude and direction of the movement precisely controlled by the transmission plan. After performing the above correction operation on each sample point in each source distribution, the decoupled fermentation features of all samples are obtained. These features have been aligned. Theoretically, samples from different raw material clusters have more consistent distribution of decoupled fermentation features in the feature space. The distribution offset caused by differences in raw material types has been weakened, thus making the features more focused on information reflecting the fermentation quality itself.
[0079] By inputting decoupled fermentation features into pre-trained quality assessment rules, the quality assessment results of woody silage are obtained and output. The quality assessment rules are pre-trained decision functions or classifiers. The training process is as follows: A historical sample set is collected, containing a large number of woody silage samples with known quality grades, such as excellent, good, medium, and poor, or using specific comprehensive scores. For each sample in the historical sample set, all steps S1 to S6 of this method are processed to obtain its decoupled fermentation features. These decoupled fermentation features and known quality grade labels constitute the training dataset. Machine learning algorithms, such as random forest algorithms, are used to learn the mapping pattern from decoupled fermentation features to quality grades. The random forest algorithm constructs multiple decision trees, each trained based on different subsets of features and samples, and makes the final decision through voting or averaging. During model training, parameters for the random forest need to be set. For example, the number of decision trees to be constructed is set to 100, and the maximum depth of each tree is set to 10. The number of decision trees (100) and the maximum depth (10) are initial empirical parameters. In actual training, cross-validation is used to search within a pre-defined parameter grid, aiming to achieve the highest prediction accuracy on the validation set, to determine the final number of trees and maximum depth parameters. These parameters can be set based on empirical values or optimized using cross-validation. The optimization algorithm adjusts these parameters to maximize the model's prediction accuracy on the validation set. After training, the obtained model parameters, tree structure, and other mapping rules are stored, forming the quality assessment rules. In the application phase, for an unknown woody silage material to be evaluated, its decoupled fermentation feature vector is first obtained according to steps S1 to S6. This feature vector is then input into the stored quality assessment rules. The quality assessment rules calculate the most likely quality grade or a predicted score corresponding to the feature vector based on the internally learned mapping relationship, and output this result as the final quality assessment result for the material. If the dimension of the input feature vector is inconsistent with the dimension used during training, a feature dimension verification error is triggered, and the consistency of the aforementioned processing steps is checked. In this way, an automated assessment of the quality of woody forage silage is achieved based on image analysis, decoupled from the influence of the raw material substrate, and objective.
[0080] Example 2: Figure 2 A schematic diagram of the image analysis-based woody forage silage quality assessment system of the present invention is provided. The image analysis-based woody forage silage quality assessment system includes:
[0081] The image acquisition module is used to acquire digital images of the woody silage material to be evaluated and the corresponding raw material type labeling information;
[0082] The feature extraction module is used to cluster the raw material type labeling information to obtain several raw material type clusters, and extract the common visual basis features and fermentation-related initial visual features of the digital images based on each raw material type cluster.
[0083] The parameter generation module is used to analyze the causal relationship between fermentation-related initial visual features and common visual base features based on the causal inference method, and generate causal decoupling guidance parameters according to the direction and strength of the causal relationship.
[0084] The matrix generation module is used to analyze the heterogeneity of causal decoupling guidance parameters in the image space and generate a spatial weighted matrix.
[0085] The signal masking module is used to weight the common visual basis features using a spatial weighting matrix to construct a feature mask. The feature mask is then applied to the fermentation-related initial visual features to mask the basis signals, thus obtaining the features after the basis signals are masked.
[0086] The quality assessment module is used to construct source distributions based on the masked features of the base signals under each raw material cluster. Based on the optimal transmission theory, each source distribution is aligned to a common reference distribution. The aligned features are output as decoupled fermentation features, and the quality assessment is completed based on the decoupled fermentation features.
[0087] All calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.
[0088] It should be noted that this invention can be deployed on the device itself to realize embedded applications, or it can run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.
[0089] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wireless or wired transmission; wired transmission methods include optical fiber, twisted pair, coaxial cable, etc.; wireless transmission includes infrared, microwave, etc. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center containing one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0090] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0091] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0092] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0093] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0094] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0095] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0096] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for evaluating the quality of a wood-based forage silage based on image analysis, characterized in that, The method comprises the following steps: S1, obtaining a digital image of a wood-based feed silage material to be evaluated and corresponding raw material category annotation information; S2, clustering and dividing the raw material category annotation information to obtain a plurality of raw material category clusters, and extracting common visual base features and fermentation-related initial visual features of the digital image based on each raw material category cluster, comprising: using a clustering algorithm to process the raw material category annotation information, grouping digital images with the same raw material category annotation information into the same set to form a raw material category cluster; for each raw material category cluster, analyzing all digital images in the raw material category cluster through an image feature extraction algorithm to extract color and texture features common among the images in the raw material category cluster as common visual base features, and extract color change and texture structure features related to silage fermentation quality as fermentation-related initial visual features; S3, analyzing the causal relationship between the fermentation-related initial visual features and the common visual base features based on a causal inference method, and generating a causal decoupling guide parameter according to the direction and strength of the causal relationship, comprising: based on a conditional independence test method, determining whether the change of the common visual base features is independent of the change of the fermentation-related initial visual features, to identify the causal action direction from the common visual base features to the fermentation-related initial visual features; at the same time, quantifying the contribution of the common visual base features to the change of the fermentation-related initial visual features through regression analysis as the causal action strength; combining the causal action direction and the causal action strength to generate a causal decoupling guide parameter for guiding subsequent decoupling; identifying the causal action direction through the conditional independence test method includes: calculating the joint conditional probability distribution of the common visual base features and the fermentation-related initial visual features under the condition of controlling other image feature variables; performing hypothesis testing by comparing the joint conditional probability distribution with the expected distribution under the independence hypothesis; and determining whether there is a causal direction from the common visual base features to the fermentation-related initial visual features according to the test result; S4, analyzing the heterogeneity of the causal decoupling guide parameter in the image space to generate a spatial weighting matrix; S5, weighting the common visual base features using the spatial weighting matrix to construct a feature mask, applying the feature mask to the fermentation-related initial visual features to shield the base signal, and obtaining the base signal shielded features; S6, constructing a source distribution with the base signal shielded features under each raw material category cluster, aligning each source distribution to a common reference distribution based on optimal transport theory, outputting the aligned features as decoupled fermentation features, and completing quality evaluation according to the decoupled fermentation features.
2. The image analysis-based evaluation method of the quality of a woody forage silage according to claim 1, characterized in that, S1 comprises: using an image acquisition device to collect digital images of wood-based feed silage material in a controlled lighting environment; determining the raw material category annotation information corresponding to the digital image by identifying the material raw material or querying the production record.
3. The image analysis-based evaluation method of the quality of woody forage silage according to claim 1, characterized in that, S4 comprises: mapping the causal decoupling guide parameter to the pixel coordinate space of the digital image to form a parameter space distribution map; traversing the parameter space distribution map through a sliding window to calculate the local statistics of the causal decoupling guide parameter in each window to quantify the local heterogeneity; The local heterogeneity values of each window are normalized, and the normalized local heterogeneity values are filled into the corresponding positions of a matrix with the same size as the digital image to generate a spatial weighting matrix.
4. The image analysis-based quality evaluation method of woody forage silage according to claim 1, characterized in that S5 The method comprises the steps of: The spatial weighting matrix and the common visual basis features are element-wise multiplied at the corresponding spatial positions to generate a spatially adaptive feature mask. The spatially adaptive feature mask is applied to the fermentation-related initial visual features to obtain the basis signal shielded features by subtracting the basis signal components represented by the spatially adaptive feature mask from the fermentation-related initial visual features.
5. The image analysis-based quality evaluation method of woody forage silage according to claim 1, characterized in that S6 The method comprises the steps of: All basis signal shielded features in each raw material category cluster are regarded as a source distribution, and the mean distribution of all source distributions is defined as a common reference distribution. An optimal transport plan from each source distribution to the common reference distribution is calculated, and the optimal transport plan defines the mapping relationship that each feature vector in the source distribution should move to the corresponding position in the common reference distribution. The spatially transformed and corrected features are output as decoupled fermentation features. The quality evaluation result of the woody forage silage material is obtained and output by inputting the decoupled fermentation features into the pre-trained quality evaluation rule.
6. The image analysis-based evaluation method of the quality of a woody forage silage according to claim 5, characterized in that, The pre-trained quality evaluation rule is obtained by the following method: collecting woody forage silage samples with known quality grades and their corresponding decoupled fermentation features as a training set; learning the mapping rule between the decoupled fermentation features and the quality grades by a machine learning algorithm; and storing the learned mapping rule as the quality evaluation rule.
7. An image analysis-based forage quality evaluation system for silage of woody plants for implementing the image analysis-based forage quality evaluation method according to any one of claims 1 to 6, characterized by, The method comprises the steps of: An image acquisition module is configured to acquire a digital image of a woody forage silage material to be evaluated and corresponding raw material category annotation information; A feature extraction module is configured to cluster and divide the raw material category annotation information to obtain a plurality of raw material category clusters, and extract common visual basis features and fermentation-related initial visual features of the digital image based on each raw material category cluster; A parameter generation module is configured to analyze the causal relationship between the fermentation-related initial visual features and the common visual basis features based on a causal inference method, and generate causal decoupling guidance parameters according to the direction and strength of the causal relationship; A matrix generation module is configured to analyze the heterogeneity of the causal decoupling guidance parameters in the image space to generate a spatial weighting matrix; A signal shielding module is configured to weight the common visual basis features using the spatial weighting matrix to construct a feature mask, and shield the basis signals of the fermentation-related initial visual features using the feature mask to obtain basis signal shielded features. A quality evaluation module is configured to construct source distributions with the basis signal shielded features under each raw material category cluster, align each source distribution to a common reference distribution based on optimal transport theory, output the aligned features as decoupled fermentation features, and complete quality evaluation according to the decoupled fermentation features.
Citation Information
Patent Citations
Remanufacturing raw material evaluation method and device, storage medium and electronic equipment
CN115984843A
Chemical raw material detection method and system based on adversarial network, medium and equipment
CN119541706A