Method and system for predicting intramuscular fat content of yak meat based on near infrared spectroscopy
By combining wavelet scattering networks and hypergraph structures, and utilizing the thermal diffusion process to synergistically diffuse label information, the problems of feature extraction and utilization of high-order correlation structures in the prediction of intramuscular fat content in yak meat are solved, thereby improving the accuracy and stability of the prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies struggle to automate and deeply mine spectral features in predicting intramuscular fat content in yak meat, and cannot effectively utilize high-order correlation structures between samples, resulting in insufficient prediction accuracy and robustness.
A wavelet scattering network is used for multi-level feature extraction. A hypergraph structure is constructed by combining metadata such as pasture and location. Label information is then collaboratively diffused on the hypergraph through a thermal diffusion process to build a prediction model.
It significantly improves the accuracy and stability of predicting intramuscular fat content in yak meat without the need for complex manual feature engineering, and enhances its generalization ability across different scenarios.
Smart Images

Figure CN121502234B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and more specifically, to a method and system for predicting intramuscular fat content in yak meat based on near-infrared spectroscopy. Background Technology
[0002] Near-infrared spectroscopy, due to its rapid and non-destructive nature, has become an important tool for meat quality testing. It reflects chemical composition information by detecting the absorption of specific wavelengths of light by molecules, making it possible to rapidly estimate the intramuscular fat content in meat products. In the field of specialty livestock products such as yak meat, achieving accurate and rapid prediction of intramuscular fat content is of great significance for early live breeding, post-slaughter grading and pricing, and quality standardization, and has been a long-term goal in this field.
[0003] Currently, common technical approaches to this goal typically rely on chemometric methods. However, these conventional methods have significant limitations: First, their model performance heavily depends on manually selected bands and feature engineering, making it difficult to automatically and deeply mine stable and robust feature patterns related to fat in the spectrum. Second, their modeling process usually treats each sample as an independent and identically distributed data point, completely ignoring the metadata information naturally carried by the sample itself, such as pasture origin, anatomical location, and physiological age, as well as the complex and high-order correlation structure between samples caused by these factors. This makes it difficult for the model to effectively learn and utilize the hidden and valuable topological constraint information in the sample population. Ultimately, the model faces significant challenges in prediction accuracy, robustness, and generalization ability when dealing with real yak meat samples from diverse sources and with high heterogeneity. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for predicting intramuscular fat content in yak meat based on near-infrared spectroscopy, in order to improve the aforementioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows:
[0005] In a first aspect, this application provides a method for predicting intramuscular fat content in yak meat based on near-infrared spectroscopy, characterized by comprising:
[0006] Obtain raw near-infrared spectral curves and intramuscular fat content values of yak meat samples from multiple pastures, multiple cuts, and multiple age groups;
[0007] Multi-level feature extraction is performed based on the original near-infrared spectral curves. Each spectral curve is processed by applying a wavelet scattering network. The wavelet scattering network performs stacked convolution and modulus operations through a preset wavelet filter bank to obtain multi-level scattering feature representations.
[0008] Based on the multi-level scattering feature representation, an implicit manifold structure between samples is constructed. By treating the scattering feature of each sample as a point in a high-dimensional space and using the preset sample metadata as a geometric prior, a hypergraph structure is constructed with samples as vertices and similarity measured by both diffusion distance in the feature space and metadata consistency as edges.
[0009] Based on the hypergraph structure, feature and label information are diffused in a coordinated manner. By defining a heat diffusion equation based on its Laplacian operator on the hypergraph, the known intramuscular fat content is used as the initial heat source to simulate the transmission process of its evolution over time on the hypergraph structure. By solving the equation, the steady-state information value of each sample node when it reaches a stable state is obtained, which integrates the topological structure and the initial label information.
[0010] A prediction model is constructed based on the multi-level scattering feature representation and the steady-state information value. The intramuscular fat content prediction model is then used to predict the near-infrared spectral curve of the sample to be predicted, thereby determining the intramuscular fat content value corresponding to the near-infrared spectral curve of the sample to be predicted.
[0011] Secondly, this application also provides a near-infrared spectroscopy-based system for predicting intramuscular fat content in yak meat, characterized by comprising:
[0012] The acquisition unit is used to acquire the raw near-infrared spectral curves and intramuscular fat content values of yak meat samples from multiple pastures, multiple parts of the meat, and multiple age groups.
[0013] The extraction unit is used to perform multi-level feature extraction based on the original near-infrared spectral curve. Each spectral curve is processed by applying a wavelet scattering network. The wavelet scattering network performs stacked convolution and modulus operations through a preset wavelet filter bank to obtain a multi-level scattering feature representation.
[0014] The construction unit is used to construct the implicit manifold structure between samples based on the multi-level scattering feature representation. By treating the scattering feature of each sample as a point in a high-dimensional space and using the preset sample metadata as a geometric prior, a hypergraph structure is constructed with samples as vertices and similarity measured by diffusion distance in feature space and metadata consistency as edges.
[0015] The simulation unit is used to perform collaborative diffusion of feature and label information based on the hypergraph structure. By defining a heat diffusion equation based on its Laplacian operator on the hypergraph, the known intramuscular fat content is used as the initial heat source to simulate the transmission process of its evolution over time on the hypergraph structure. By solving the equation, the steady-state information value of each sample node when it reaches a stable state is obtained, which integrates the topological structure and the initial label information.
[0016] The prediction unit is used to construct a prediction model based on the multi-level scattering feature representation and the steady-state information value, and to predict the near-infrared spectral curve of the sample to be predicted based on the constructed intramuscular fat content prediction model, thereby determining the intramuscular fat content value corresponding to the near-infrared spectral curve of the sample to be predicted.
[0017] The beneficial effects of this invention are as follows:
[0018] This invention automatically extracts deep, stable features from the spectrum using wavelet scattering networks and constructs a hypergraph structure based on metadata such as pasture and location to model complex relationships between samples. Then, it utilizes thermal diffusion to collaboratively diffuse label information onto the hypergraph, ultimately fusing multi-level spectral features with diffused steady-state information values to build a prediction model. This effectively solves the problems of traditional methods struggling to robustly extract fat features from highly variable spectra and failing to utilize high-order relationships between samples. It achieves significant improvements in the accuracy, stability, and cross-scenario generalization ability of predicting intramuscular fat content in yak meat by fully utilizing limited label information without complex manual feature engineering.
[0019] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of the method for predicting intramuscular fat content in yak meat based on near-infrared spectroscopy as described in an embodiment of the present invention;
[0022] Figure 2 This is a schematic diagram of the yak meat intramuscular fat content prediction system based on near-infrared spectroscopy as described in an embodiment of the present invention.
[0023] In the diagram: 701, acquisition unit; 702, extraction unit; 703, construction unit; 704, simulation unit; 705, prediction unit. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0025] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0026] Example 1:
[0027] This embodiment provides a method for predicting intramuscular fat content in yak meat based on near-infrared spectroscopy.
[0028] See Figure 1 The figure shows that the method includes steps S1, S2, S3, S4 and S5.
[0029] Step S1: Obtain the original near-infrared spectral curves and intramuscular fat content values of yak meat samples from multiple pastures, multiple parts of the meat, and multiple age groups.
[0030] It is understandable that step S1 is the foundational data preparation stage for building the predictive model. Its core lies in systematically collecting a calibration sample set that comprehensively reflects the natural variation range of intramuscular fat in yak meat. The "raw near-infrared spectral curve" obtained in this step refers to the unprocessed raw spectral reflectance or absorbance data sequence obtained by directly scanning fresh cut surfaces or specifically treated surfaces within a specific wavelength range using a near-infrared spectrometer on individual yak meat samples from different altitudes, pasture environments, physiological locations, and age groups. This data sequence fully contains the overtone and combination frequency absorption information of organic molecules (such as CH, NH, and OH bonds) in the sample. Simultaneously, the "intramuscular fat content value" refers to the accurate fat content reference value measured by chemical analysis using the Soxhlet extraction method on the adjacent muscle tissue of each of the above spectral scan samples. This value serves as the "standard" for subsequent model training and validation. By collecting samples in a one-to-one correspondence between "spectral and physicochemical values", the dataset is ensured to have sufficient coverage and representativeness in key influencing factors such as pasture, location, and age. The aim is to build a data foundation that can cover the main sources of variation in actual production scenarios from the source, providing a prerequisite for the subsequent establishment of a stable prediction model with high generalization ability.
[0031] Step S2: Perform multi-level feature extraction based on the original near-infrared spectral curves. Process each spectral curve by applying a wavelet scattering network. The wavelet scattering network performs stacked convolution and modulus operations through a preset wavelet filter bank to obtain multi-level scattering feature representations.
[0032] Understandably, this step treats each original spectral curve as a one-dimensional signal. A pre-defined set of wavelet filters covering different scales and frequencies is used to perform multi-level convolution operations, progressively decomposing the signal's local oscillation modes at different resolutions. The subsequent modulus operation then acts as a nonlinear stabilizer, effectively suppressing random disturbances unrelated to chemical composition while preserving its energy distribution characteristics. Through this layered "convolution-modulus operation-smoothing" process, the network can gradually construct a series of scattering coefficients with translation invariance and deformation stability from the original spectrum. These coefficients constitute a "multi-level scattering feature representation." In this step, step S2 includes steps S21, S22, and S23.
[0033] Step S21: Perform time-frequency joint analysis of the near-infrared spectral curve based on the original near-infrared spectral curve. Convolve the spectral curve using a preset Gabor wavelet filter with different scales and center frequencies to decompose the spectral signal in the spectral curve into a time-frequency representation on the wavelength-frequency two-dimensional plane, and obtain the primary time-frequency coefficients containing local spectral oscillation mode information.
[0034] Understandably, this step first performs time-frequency joint analysis on the acquired raw near-infrared spectral curve. The core of this analysis is to convolve the spectral curve using a set of pre-defined Gabor wavelet filters with different scales and center frequencies. The essence of this processing method lies in the fact that Gabor wavelets possess optimal resolution in both the time and frequency domains, enabling the decomposition of a one-dimensional spectral signal (its "time domain" being the wavelength axis) onto a two-dimensional plane composed of wavelength and frequency, forming a so-called "time-frequency representation." This representation clearly reveals the intensity of specific frequencies (which can be understood as the steepness of absorption peaks, shoulder peaks, and other detailed oscillations) contained in the spectral signal at different wavelength positions, thus obtaining "primary time-frequency coefficients." In the specific scenario of detecting intramuscular fat in yak meat, the characteristic absorption peak of carbon-hydrogen bonds (CH bonds) in fat molecules is not an isolated sharp peak in the near-infrared spectrum, but often exhibits a complex envelope with a certain width that may overlap with the absorption peaks of other components (such as water and protein).
[0035] The formulas for convolving the spectral curves with Gabor wavelet filters of different scales and center frequencies are shown below:
[0036] ;
[0037] Among them, W(b) j ,σ k ,f l ) represents the Gabor wavelet transform coefficients in complex form, b j Let σ be the translation parameter. k f is the scale parameter. l Let be the center frequency parameter, s(λ) be the original near-infrared spectral signal, which is a function of wavelength λ, where λ is the wavelength variable, and π is pi. To integrate the product over the entire wavelength domain, i is the summation index, corresponding to each discrete sampling point of the spectrum, and e is the natural constant.
[0038] The formula for calculating the primary time-frequency coefficient is as follows:
[0039] ;
[0040] Where T[j,k,l] represents the primary time-frequency coefficients, and j,k,l are all discrete indices, where j is the index of the translation parameter, k is the index of the scale parameter, and l is the index of the center frequency parameter. To obtain the absolute value, M is the total number of points after discrete sampling of the original spectral signal, s[i] is the sampled value of the original spectral signal at the i-th discrete wavelength point, and λ i Let represent the wavelength value corresponding to the i-th discrete sampling point, where i is the index of the original spectral signal, representing the i-th discrete wavelength point.
[0041] In this context, different scales are specifically represented by the window width of the Gabor wavelet filter on the wavelength axis, for example, set to different values such as 5 nm, 15 nm, and 30 nm. A small-scale filter of 5 nm has a very narrow window, which can finely capture the details of a very sharp absorption peak with a width of only about 8 nm near the wavelength of 1210 nm. On the other hand, a large-scale filter of 30 nm has a wider window, which can smoothly cover a wider band from 1150 nm to 1180 nm, characterizing the slow trend of the overall absorption background in this region, reflecting the broad-band absorption substrate of water or protein in the sample.
[0042] Different center frequencies directly correspond to the signal oscillation frequencies of the primary response of the filter, and their values can be specifically set by the frequency parameters in the filter function. For example, a filter with a high center frequency is extremely sensitive to rapid fluctuations (high-frequency oscillations) of the signal within a short wavelength range (such as 2-3 nanometers). It can extract subtle "shoulder peaks" or asymmetry information on both sides of a sharp absorption peak at 1220 nanometers from the spectral curve. These details may be caused by the slight overlap of absorption peaks from fat and collagen. Conversely, a filter with a low center frequency is more responsive to slow, trending changes (low-frequency oscillations) of the signal within a wavelength range of tens of nanometers. For example, it can effectively depict the "slope" shape of the overall downward shift in spectral reflectance across the entire spectral band from 1100 nanometers to 1300 nanometers.
[0043] Step S22: Based on the primary time-frequency coefficients, feature fusion is performed on the spectral response characteristics of intramuscular fat. By using the pre-set near-infrared absorption bands of known fat chemical bonds as prior information, the coefficients of the time-frequency representation in the corresponding wavelength range are weighted and aggregated and nonlinearly modulated to obtain a robust feature representation that incorporates scene knowledge.
[0044] Understandably, this step utilizes the known characteristic absorption bands of key chemical groups (mainly CH bonds) in intramuscular fat in the near-infrared region (e.g., bands around 1210 nm, 1390 nm, and 1730 nm related to the stretching vibrations of methyl and methylene groups in fat), using these prior wavelength ranges as "attention windows." During processing, the primary time-frequency coefficients are "weighted and aggregated" across all frequency dimensions within these specified wavelength ranges. Essentially, this assigns higher weights to coefficients within these ranges, particularly those reflecting the morphology of fat absorption peaks (such as oscillations at specific frequencies), and performs summation or averaging operations, thereby achieving feature screening and concentration of fat information in the time-frequency domain. The subsequent "nonlinear modulus operation" (i.e., taking the square root and absolute value) further stabilizes the aggregated features, eliminating random fluctuations that may be caused by signal phase and converting energy information into non-negative scalars. This significantly enhances the robustness of the features to interferences such as slight shifts in the spectral signal and minor instrument drift.
[0045] The formula for weighted aggregation is as follows:
[0046] ;
[0047] in, For the aggregation characteristics at band p and scale k, The wavelength weight vector for the p-th fat feature band is generated by a Gaussian function. Let L be the frequency weight vector of the p-th fat feature band, where L represents all frequencies within the frequency dimension, and J is the total number of translation parameters.
[0048] Step S23: Construct deep stable features based on the robust feature representation that integrates scene knowledge. By convolving the robust feature representation with wavelet filter banks and performing modulus operations, and iteratively smoothing and downsampling along the frequency axis, a multi-scale scattering feature representation of small deformations of the spectral signal is constructed, namely the multi-level scattering feature representation.
[0049] Understandably, this step involves convolving the robust feature representation obtained in the previous step with a set of wavelet filter banks (filters with different scales and frequencies than those in the first layer). This is equivalent to further analyzing the finer local pattern changes based on the existing features, thereby capturing the hierarchical structural information within the features. The subsequent modulus operation (i.e., taking the absolute value) is applied again, its core function being to provide a stable nonlinear transformation. This transformation suppresses phase changes caused by minute signal shifts or distortions in the output of the previous convolution layer, retaining only the energy amplitude. This is a key mathematical operation for achieving feature stability. After each convolution-modulus operation, iterative smoothing and downsampling are performed along the frequency axis. The smoothing operation (usually implemented through a low-pass filter) purposefully filters out high-frequency details that may be noise, retaining the main trend information; downsampling reduces the data resolution, achieving information condensation and expanding the receptive field for subsequent operations. Through multiple iterations of this convolution-modulus operation-smoothing-downsampling, the network generates feature maps at different scales layer by layer. Each layer's output represents the stabilized energy distribution information of the original spectral signal at different resolutions (or "scales"). Finally, these stable features from different iterative layers (i.e., different "scales") are concatenated to form the final "multi-level scattering feature representation." In the highly variable scenario of near-infrared spectroscopy applications with yak meat, the surface state of the sample, measurement location, and instrument status all introduce minute deformations and translations of the spectrum. This step, through the aforementioned iterative architecture, systematically constructs a feature representation with "deformation stability" and "translation invariance," ensuring that the extracted multi-level scattering features remain relatively stable regardless of slight shifts in the absorption peak of fat along the wavelength axis or minor changes in the overall baseline.
[0050] The formulas for convolution and modulo operation are shown below:
[0051] ;
[0052] Where V[k] is the intermediate feature after convolution and modulo, N is the length of the input signal, n is the index of the input signal, U[n] is the input signal, and Ψ λ Here, k is the wavelet filter function, and k is the index of the convolution result.
[0053] The smoothing formula is shown below:
[0054] ;
[0055] Where S[n] is the scattering coefficient of the final output, K is the length of the intermediate feature, V[k] is the intermediate feature, 2n-k is the index expression of the downsampling operation, which is equivalent to calculating the convolution once every other point, achieving a 2x downsampling, and Φ is the low-pass smoothing filter. A discrete, finite-length filter sequence, which is a Gaussian filter in this step.
[0056] Step S3: Construct an implicit manifold structure between samples based on the multi-level scattering feature representation. By treating the scattering feature of each sample as a point in a high-dimensional space and using the preset sample metadata as a geometric prior, a hypergraph structure is constructed with samples as vertices and similarity measured by the diffusion distance in the feature space and the consistency of metadata as edges.
[0057] Understandably, this step first utilizes the diffusion distance between feature points to measure their proximity on the intrinsic manifold, capturing the similarity of spectral features. Simultaneously, pre-defined sample metadata (such as pasture number and part category) is used as "geometric prior" knowledge to constrain or enhance the construction of this relationship; for example, forcing samples from the same pasture to connect or adjusting the connection strength based on part priors. Finally, by fusing these two measures (data-driven diffusion distance and knowledge-driven metadata consistency), multiple samples meeting the conditions are defined as a "hyperedge," thus constructing a "hypergraph structure." This structure not only encodes pairwise similarities between samples but, more importantly, characterizes high-order group relationships formed by group attributes (such as belonging to the same pasture) or common features (such as having similar spectral pattern clusters), thereby embedding discrete sample points into a relational network rich in topological constraints, laying the foundation for subsequent information propagation and collaborative learning using graph structures. In this step, step S3 includes steps S31, S32, and S33.
[0058] Step S31: Perform nonlinear measurement calculation of high-dimensional feature space based on the multi-level scattering feature representation. In this step, the Mahalanobis distance between each pair of sample feature vectors is calculated, and a k-nearest neighbor graph reflecting the local neighborhood relationship of the data is constructed based on the calculated Mahalanobis distance to obtain the preliminary similarity relationship between samples based on spectral features.
[0059] Understandably, this step abandons the conventional Euclidean distance and instead calculates the Mahalanobis distance between each pair of sample feature vectors. Unlike Euclidean distance, which assumes that each feature dimension is independent and homoscedastic, Mahalanobis distance effectively considers the potential correlation between different feature dimensions and the different variances of each dimension by introducing a covariance matrix calculated from all sample features. In the specific scenario of yak meat spectral analysis, the features extracted from different levels by the wavelet scattering network are not independent, and their contributions to fat content prediction (importance variance) are also different. Mahalanobis distance can precisely characterize this "ellipsoidal" data distribution, making the calculated distance more reflective of the true "distance" of samples in the essential feature space. Based on the calculated Mahalanobis distance matrix, a "k-nearest neighbor graph" is further constructed. Specifically, for each sample node, only the k other sample nodes with the smallest Mahalanobis distance are connected, thus forming a sparse graph structure. The "preliminary similarity relationship" expressed by this k-nearest neighbor graph is a local neighborhood relationship established purely from a data-driven perspective based on the morphological similarity of spectral features.
[0060] The formula for calculating the Mahalanobis distance is existing technology and will not be elaborated here.
[0061] Step S32: Based on the original near-infrared spectral curves of the samples and the preliminary similarity relationship, perform constraint fusion of cross-pasture geographical relationship. By forcibly connecting sample pairs belonging to the same pasture, and applying a priori adjustment coefficient to the connection strength between different sample pairs according to the differences in fat deposition patterns in different locations, an enhanced adjacency relationship that integrates spectral similarity and scene metadata constraints is obtained.
[0062] Understandably, this step first performs a "forced connection" operation, which adds a connection edge to two samples in the graph as long as they originate from the "same pasture," regardless of their Mahalanobis distance in the feature space. The physical meaning of this operation is that yaks from the same pasture share similar altitudes, climates, feeding methods, and forage sources. These systematic environmental factors lead to an inherent background correlation in their meat spectra that transcends instantaneous chemical composition similarities. Forced connections ensure that this potential similarity caused by shared "geographical" proximity is explicitly and firmly encoded in the graph. Secondly, regarding "location" information, based on the inherent physiological patterns of intramuscular fat deposition in different anatomical locations (e.g., there are objective differences in the fat distribution patterns and content ranges between the tenderloin and the cervical loin), an "a priori adjustment coefficient" is applied to the "connection strength" of all sample pairs in the graph (including edges generated by k-nearest neighbors and forced connections). For example, this can strengthen the connection weight between sample pairs from the same location while weakening the connection weight between sample pairs from different locations. This adjustment is based on a domain understanding: body location is a strong determinant of fat content, and spectral similarity between samples from the same body location is more likely to indicate similar fat content. Through the two operations described above, the original k-nearest neighbor graph, which only reflects spectral morphological similarity, is infused with domain knowledge and transformed into "enhanced adjacency relationships that integrate spectral similarity and scene metadata constraints."
[0063] The formula for constructing enhanced adjacency relationships by fusing spectral similarity and scene metadata constraints is shown below:
[0064] ;
[0065] Among them, W PQ To enhance the adjacency matrix, w knn Let A be the base weight of the k-nearest neighbor connection. PQ Let A be the k-nearest neighbor adjacency matrix, a binary symmetric matrix with elements A. PQ =1 if and only if sample P is one of the k nearest neighbors of sample Q, otherwise 0, w ranch R is the base weight for forced connections to the ranch. PQ Let R be the pasture incidence matrix, a bivariate symmetric matrix with elements R. PQ =1 if and only if sample P and sample Q come from the same pasture, otherwise =0, C PQ Let C be the part consistency matrix, a binary symmetric matrix with elements C. PQ =1 if and only if sample P and sample Q come from the same anatomical site, otherwise 0, α is the connection weight enhancement coefficient for the same site, which is 1.2, and β is the connection weight enhancement coefficient for different sites, which is 0.8.
[0066] Step S33: Encode higher-order sample associations based on the enhanced adjacency relationship. By defining each group of samples that satisfies the enhanced adjacency relationship as a hyperedge, a hypergraph structure is constructed with samples as vertices and the similarity between the diffusion distance in the feature space and the original near-infrared spectral curve as edges.
[0067] It is understandable that this step defines each group of samples satisfying the enhanced adjacency relation as a hyperedge. Satisfying the enhanced adjacency relation here refers not only to sample pairs that are neighbors in the k-nearest neighbor relationship, but more importantly, it explicitly includes the group of all samples associated with the attribute of "same pasture," as well as another group of samples associated with the attribute of "same location." Through this definition, a "hyperedge" can connect multiple sample vertices, for example, connecting all samples from "Pasture A," or all samples belonging to the "thigh" location, or all samples within a small, very close local neighborhood in the feature space. The hypergraph structure constructed in this way has its hyperedges based on the union of the "diffusion distance in the feature space" (reflected in the k-nearest neighbor local group) and the group relations derived from the "similarity of the original near-infrared spectral curves" (reflected in the pasture, location, and other attribute groups).
[0068] Step S4: Perform collaborative diffusion of feature and label information based on the hypergraph structure. Define a heat diffusion equation based on its Laplacian operator on the hypergraph, use the known intramuscular fat content as the initial heat source, simulate its transmission process over time on the hypergraph structure, and obtain the steady-state information value of each sample node when it reaches a stable state by solving the equation.
[0069] Understandably, this step mathematically formalizes how "heat" (i.e., fat content information) flows and diffuses along higher-order paths connected by hyperedges from "heat source" nodes with known labels to nodes with unknown labels by defining a heat diffusion equation based on the hypergraph's Laplace operator. In this simulated transmission process, information does not propagate randomly but strictly follows the degree of connection defined by the hypergraph structure; closely related samples exchange information more frequently and fully. As the simulation progresses, the system eventually reaches a thermal equilibrium "steady state." In this state, the "steady-state information value" carried by each sample node is no longer merely its initial, potentially sparse or noisy, original label, but rather a weighted fusion of its initial value and information from its topological neighbors (especially multi-node groups connected by hyperedges). This process is essentially a label denoising, smoothing, and information supplementation mechanism under the constraints of complex relationship networks. This ensures that the target value of each sample used for modeling incorporates both its own quantitative analysis and its position within the group's relationship network, thus providing higher-quality and more comprehensive supervisory signals for building a more robust prediction model. In this step, step S4 includes steps S41, S42, and S43.
[0070] Step S41: Construct higher-order graph structure operators based on the hypergraph structure. By calculating the Laplacian matrix of the hypergraph, obtain mathematical operators describing the ability of information to diffuse on the hypergraph structure.
[0071] It is understandable that the Laplacian matrix of the hypergraph in this step differs from the Laplacian matrix describing a regular graph (containing only pairwise connections). The construction of the hypergraph's Laplacian matrix requires careful consideration of the higher-order characteristic that the "hyperedge" connects an arbitrary number of vertices (multiple samples). Each element of this matrix not only reflects whether there is a direct connection between any two sample vertices, but more importantly, it accurately characterizes the interaction strength between vertices and the overall topological constraints when information diffuses along a hyperedge connecting multiple samples by introducing the degree of the hyperedge, the degree of the vertices, and the association matrix. In this invention, since the relationships between samples have been modeled through the hypergraph as simultaneously including multiple higher-order associations such as spectral similarity local clusters, same-pasture groups, and same-part categories, the calculated hypergraph Laplacian matrix becomes a "mathematical operator" capable of comprehensively describing these complex relationships. This operator inherently encodes the inherent laws and capabilities of how information (fat content information) diffuses and exchanges within the sample network along higher-order paths such as "belonging to the same pasture" or "being from the same loin part."
[0072] The formula for the Laplace matrix in this step is as follows:
[0073] ;
[0074] Among them, L c For the normalized hypergraph Laplacian matrix, I c It is the identity matrix. Let H be the diagonal matrix of vertex degree raised to the power of -1 / 2, and let H be the hypergraph incidence matrix, a binary matrix whose elements are 1 if and only if sample P belongs to hyperedge e, otherwise 0. e This is a diagonal matrix with hyperedge weights. H is the reciprocal of the hyperdiagonal matrix. ⊤ This is the transpose of the hypergraph incidence matrix.
[0075] Step S42: Construct an initial information field based on the intramuscular fat content value. Normalize the content value of the sample nodes with known intramuscular fat content and use it as the initial temperature value. Set the initial temperature of the sample nodes with unknown content to zero. Combine the intramuscular fat content distribution range within the sample nodes with known intramuscular fat content value to constrain the boundary of the initial temperature field and obtain an initial temperature field that conforms to the physical diffusion law.
[0076] Understandably, this step first "normalizes" the content values of all sample nodes with known content, for example, by linearly scaling them to the [0,1] interval, and then directly assigns the normalized value to the corresponding sample node as its "initial temperature value." For sample nodes with unknown content, their initial temperature value is set to zero. This setting constitutes the initial condition for the diffusion process, that is, the known content samples are the "heat source" carrying heat, and the unknown samples are the "cold zone" to be heated. However, this operation alone may cause the distribution range of the "heat source" temperature (i.e., the normalized content value) in the initial field to not completely match the possible content range of the actual organism, or there may be individual outliers. Therefore, this step introduces a "boundary constraint" operation, which combines the actual distribution range of intramuscular fat content within known content sample nodes (for example, based on historical data, the fat content of this part of yak meat is roughly distributed between a% and b%) to constrain the original content value before normalization or the normalized temperature value. For example, by truncating or scaling, it ensures that all values used as initial temperatures fall within a reasonable range that conforms to prior physiological knowledge and is closer to the real physical world.
[0077] The formulas for linear normalization and boundary truncation are shown below:
[0078] ;
[0079] in, This represents the normalized content value of the Pth known sample after boundary constraint processing. Let be the original intramuscular fat content value of the Pth known sample, a be the theoretical lower limit of intramuscular fat content distribution, b be the theoretical upper limit of intramuscular fat content distribution, a and b are both set by historical data, and they are the minimum and maximum values of historical data, max(·,0) and min(·,1) are both cutoff functions, max(·,0) ensures that the result is not less than 0, and min(·,1) ensures that the result is not greater than 1.
[0080] Step S43: Solve the steady-state problem of the thermal diffusion process based on the mathematical operator and the initial temperature field. Discretize the thermal diffusion equation and solve iteratively until the temperature change of all nodes on the entire hypergraph structure is less than a preset threshold, and obtain the steady-state information value of each sample node.
[0081] Understandably, this step substitutes the hypergraph Laplacian operator and the initial temperature field into the heat diffusion equation describing the heat conduction process. Since this equation cannot be solved analytically directly on the discrete topology of the hypergraph, it needs to be "discretized" to transform it into a linear system that can be iteratively computed by a computer. In each iteration, the current temperature (i.e., information value) of each node is updated based on the current temperature difference between it and all its neighboring nodes connected via hyperedges, as well as the connection strength and direction precisely quantified by the hypergraph Laplacian operator. This process simulates the flow of heat (i.e., fat content information) from high-temperature nodes (known high-fat-content samples) along various higher-order association paths defined by the hypergraph (like pastures, similar locations, spectrally similar clusters) to low-temperature nodes (unknown or low-fat-content samples). The iteration continues until "the temperature change of all nodes in the entire hypergraph structure is less than a preset threshold," marking that the system has reached a "steady state" of thermal equilibrium. At this point, the "steady-state information value" obtained by each sample node is no longer its initial, isolated chemical measurement value, but rather the result of multiple rounds of weighting, mixing, and balancing of its own initial value and the information of all connected neighbor nodes in the entire relational network.
[0082] The thermal diffusion equation is as follows:
[0083] ;
[0084] Among them, P c Let I be the transition matrix, which is the identity matrix. c Subtract the (normalized) Laplace matrix L c .
[0085] The formula for iterative calculation is shown below:
[0086] ;
[0087] Among them, f (t+1)Let Y be the temperature at time step t+1, Y be the sample mask vector, where known samples correspond to elements of 1, and unknown samples correspond to elements of 0. (0) Let f be the initial temperature field vector. The values of the known sample nodes are the fat content normalized by boundary constraints, while the values of the unknown sample nodes are 0. (t) Let be the temperature at time step t. This is an element-wise multiplication operation.
[0088] Step S5: Construct a prediction model based on the multi-level scattering feature representation and the steady-state information value, and predict the near-infrared spectral curve of the sample to be predicted based on the constructed intramuscular fat content prediction model, thereby determining the intramuscular fat content value corresponding to the near-infrared spectral curve of the sample to be predicted.
[0089] It is understandable that the "multi-level scattering feature representation" in this step carries the sample's own detailed spectral fingerprint information, while the "steady-state information value" contains the smoothed and corrected trend of the sample's fat content in the group relationship network. When building the model, the steady-state information value is used as an important supplement or constraint to the supervision signal, guiding the model to simultaneously consider data-driven detailed fitting and the group consistency rules implied by the graph structure when learning the complex mapping from spectral features to fat content. The "intramuscular fat content prediction model" trained in this way inherently integrates the sample's individual chemical characteristics with its topological position information in the group in its decision logic. When applied to a new sample to be predicted, the model first obtains the fused feature representation of the sample through the same feature extraction and graph embedding process (using the existing structure to calculate its relationship), and then directly outputs the predicted value. This process allows the model to not only rely on the spectrum of the new sample itself, but also implicitly reference its "nearest neighbor" information in the historical sample relationship network. This significantly improves the accuracy and stability of predictions when dealing with yak meat samples from new pastures and new batches, effectively alleviating the problem of insufficient model generalization ability caused by limited training data or distribution bias. In this step, step S5 includes steps S51, S52, and S53.
[0090] Step S51: Perform scene-adaptive feature fusion based on the multi-level scattering feature representation and the steady-state information value. By using the steady-state information value of each sample as an attention weight, the multi-level scattering feature representation of the corresponding sample is weighted, and the weight is corrected a priori by combining the sample's part metadata, to obtain a deep fusion feature vector that fuses topological steady-state information and spectral features.
[0091] Understandably, this step transforms the "steady-state information value" of each sample into an "attention weight." The level of the steady-state information value essentially reflects the importance of the sample's position in the hypergraph network and the confidence level of its fat content prediction; a higher steady-state information value may mean that the sample belongs to a well-labeled and closely related group, and its spectral features should be given higher attention. Using this weight to weight the "multi-level scattering feature representation" of the corresponding sample is equivalent to making the model pay more attention to the spectral details of samples that occupy key positions in the network and have reliable label information. Furthermore, to incorporate domain knowledge to correct weight biases that may be caused by network noise, this step "combines the sample's location metadata to perform prior correction on the weights." For example, for locations with low fat deposition potential, such as the buttocks, even if their steady-state information value is high, their weights may be appropriately lowered to conform to physiological laws. Through this combination of dynamic weighting and prior correction, the final "deep fusion feature vector" is no longer the original spectral features, but a new feature modulated by network topology information and calibrated with scene knowledge.
[0092] The formulas for generating and correcting attention weights are as follows:
[0093] ;
[0094] in, Let σ be the final attention weight after correction for the P-th sample, and s be the Sigmoid function. P Let μ be the steady-state information value of the Pth sample. s σ is the mean of the steady-state information values. s c is the standard deviation of the steady-state information value. pb For part p b The prior correction factor.
[0095] The formula for calculating the feature vector of deep fusion is as follows:
[0096] ;
[0097] Among them, z P Let x be the deep fusion feature vector of the Pth sample. P For the multi-level scattering feature representation of the P-th sample, m t This is the feature importance mask vector.
[0098] Step S52: Based on the deep fusion feature vector and the intramuscular fat content value, perform prediction function learning under graph structure constraints. By treating the samples as graph nodes and their deep fusion feature vectors as node attributes, under the constraints of the constructed hypergraph topology, train a preset graph neural network so that while aggregating the features of neighboring nodes, it learns the mapping function from node features to intramuscular fat content, thus obtaining an intramuscular fat content prediction model.
[0099] Understandably, this step first treats each yak meat sample as a node in a hypergraph, and uses the "deeply fused feature vector" as the attribute vector of that node. During model training, the graph neural network does not process each node in isolation, but operates "under the constraints of the constructed hypergraph topology." Its core mechanism is "aggregating neighborhood node features," meaning that each layer of the neural network allows the features of a node to interact and fuse with the features of all its neighboring nodes connected by hyperedges. This information transfer iterates multiple times along higher-order paths defined by the hypergraph (such as "same pasture" group or "same cut" category). In this way, when learning the mapping from individual node features to fat content, the model can explicitly utilize the correlation between samples: the prediction of a sample depends not only on its own features, but also on the group features of its pasture, the common features of its cut, and the smoothing effect of spectral nearest neighbor features. Ultimately, the goal of model optimization during training is to make its final output as close as possible to the true intramuscular fat content value.
[0100] The convolution formula for each layer in this invention is as follows:
[0101] ;
[0102] Among them, E (1) Z is the hyperedge feature matrix, and Z is the node feature matrix. (1) For the updated node feature matrix, W (1) is the weight matrix of the hypergraph convolutional layer.
[0103] The loss function of the prediction model is shown below:
[0104] ;
[0105] Where L is the total loss function value, n train y is the number of training samples. P This represents the true intramuscular fat content of the P-th training sample. Let be the model prediction value for the P-th training sample, and λ be the regularization coefficient. To penalize the size of the first layer convolution weights, To penalize the size of the weights in the second convolutional layer, The size of the output layer weights is used to penalize them.
[0106] Step S53: Based on the intramuscular fat content prediction model, the fat content of the sample to be predicted is estimated. By determining the deep fusion feature vector corresponding to the sample to be predicted and the topological relationship of the sample to be predicted in the hypergraph, and inputting them into the intramuscular fat content prediction model, the predicted value of the intramuscular fat content of the sample to be predicted is obtained.
[0107] Understandably, this step first requires determining the "deep fusion feature vector" of the sample to be predicted. This process recursively depends on the preceding steps: the original near-infrared spectral curve of the sample is processed sequentially through the wavelet scattering network in step S2 to obtain its multi-level scattering feature representation, and then processed through steps S3 and S4 (adding it as a new node to the existing hypergraph structure and calculating its steady-state information value). Finally, through the scene adaptive fusion mechanism in step S51, its unique deep fusion feature vector is generated. Simultaneously, the "topological relationship" of the sample in the hypergraph needs to be determined, i.e., based on its pasture, part metadata, and spectral features, the connection relationship between it and each node in the original training sample set is calculated (such as whether it belongs to the same pasture or part hyperedge, and the k-nearest neighbor relationship based on feature similarity), thereby dynamically embedding it into the constructed hypergraph topology. Subsequently, this fusion feature vector is used as a node attribute, and its connection relationship is used as an extension of the graph structure, both input into the "intramuscular fat content prediction model" trained in step S52. Finally, the value output by the model for this new node is the "predicted intramuscular fat content value of the sample to be predicted".
[0108] The formula for predicting intramuscular fat content is shown below:
[0109] ;
[0110] Among them, s new P is a predicted value for intramuscular fat content. ext Let P be the transition matrix of the hypergraph. ext [new,j] represents the element in the new-th row and j-th column of the transition matrix, P ext [new,new] represents the element in the new-th row and new-th column of the transition matrix. Let N(new) be the steady-state information value of training sample j, and let N(new) be the set of neighboring vertices of the new sample in the extended hypergraph.
[0111] Example 2:
[0112] like Figure 2 As shown, this embodiment provides a near-infrared spectroscopy-based system for predicting intramuscular fat content in yak meat. (See also...) Figure 2The system includes an acquisition unit 701, an extraction unit 702, a construction unit 703, a simulation unit 704, and a prediction unit 705.
[0113] The acquisition unit 701 is used to acquire the original near-infrared spectral curves and intramuscular fat content values of yak meat samples from multiple pastures, multiple parts, and multiple age groups.
[0114] Extraction unit 702 is used to perform multi-level feature extraction based on the original near-infrared spectral curve. Each spectral curve is processed by applying a wavelet scattering network. The wavelet scattering network performs stacked convolution and modulus operations through a preset wavelet filter bank to obtain a multi-level scattering feature representation.
[0115] The construction unit 703 is used to construct an implicit manifold structure between samples based on the multi-level scattering feature representation. By treating the scattering feature of each sample as a point in a high-dimensional space and using the preset sample metadata as a geometric prior, a hypergraph structure is constructed with samples as vertices and similarity measured by diffusion distance in the feature space and metadata consistency as edges.
[0116] The simulation unit 704 is used to perform collaborative diffusion of feature and label information according to the hypergraph structure. By defining a heat diffusion equation based on its Laplacian operator on the hypergraph, the known intramuscular fat content is used as the initial heat source to simulate the transmission process of its evolution over time on the hypergraph structure. By solving the equation, the steady-state information value of each sample node when it reaches a stable state is obtained, which integrates the topological structure and the initial label information.
[0117] The prediction unit 705 is used to construct a prediction model based on the multi-level scattering feature representation and the steady-state information value, and to predict the near-infrared spectral curve of the sample to be predicted based on the constructed intramuscular fat content prediction model, thereby determining the intramuscular fat content value corresponding to the near-infrared spectral curve of the sample to be predicted.
[0118] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0119] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0120] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for predicting intramuscular fat content of yak meat based on near infrared spectroscopy, characterized by, The application relates to a method for predicting intramuscular fat content of yak meat, comprising the following steps: obtaining original near-infrared spectrum curves and intramuscular fat content values of yak meat samples of multiple pastures, multiple parts and multiple age stages; performing multi-level feature extraction according to the original near-infrared spectrum curves, processing each spectrum curve by applying a wavelet scattering network, wherein the wavelet scattering network is subjected to convolution and modulus operation in a layer-by-layer mode through a preset wavelet filter bank to obtain multi-level scattering feature representation; constructing an implicit manifold structure between samples according to the multi-level scattering feature representation, regarding the scattering feature of each sample as a point in a high-dimensional space, and regarding preset sample metadata as geometric prior, to obtain a hypergraph structure with samples as vertices and similarity as edges, wherein the similarity is jointly measured by diffusion distance in a feature space and metadata consistency; performing collaborative diffusion of features and label information according to the hypergraph structure, defining a heat diffusion equation based on a Laplacian operator on the hypergraph, taking the known intramuscular fat content values as an initial heat source, simulating the conduction process of the initial heat source on the hypergraph structure over time, and obtaining a steady-state information value of each sample node when reaching a stable state and fusing the topological structure and the initial label information by solving the equation; constructing a prediction model according to the multi-level scattering feature representation and the steady-state information value, and predicting the near-infrared spectrum curve of a to-be-predicted sample based on the constructed intramuscular fat content prediction model, and then determining the intramuscular fat content value corresponding to the near-infrared spectrum curve of the to-be-predicted sample.
2. The method for predicting intramuscular fat content of yak meat based on near infrared spectroscopy according to claim 1, characterized by, The multi-level feature extraction according to the original near-infrared spectrum curve comprises the following steps: performing time-frequency joint analysis of the near-infrared spectrum curve according to the original near-infrared spectrum curve, performing convolution operation on the spectrum curve by using a preset Gabor wavelet filter with different scales and center frequencies, decomposing the spectrum signal in the spectrum curve into a time-frequency representation on a wavelength-frequency two-dimensional plane, and obtaining primary time-frequency coefficients containing local spectrum oscillation mode information; performing feature fusion for intramuscular fat spectrum response characteristics according to the primary time-frequency coefficients, taking the near-infrared absorption band of a preset known fat chemical bond as prior information, performing weighted aggregation and nonlinear modulus operation on the coefficients of the time-frequency representation in the corresponding wavelength interval, and obtaining a robust feature representation fusing scene knowledge; constructing a deep stable feature according to the robust feature representation fusing scene knowledge, performing convolution and modulus operation on the robust feature representation and a wavelet filter bank, and performing iterative smoothing and down-sampling along the frequency axis, to obtain a multi-scale scattering feature representation of the spectrum signal micro-deformation, that is, the multi-level scattering feature representation.
3. The method of predicting intramuscular fat content of yak meat based on near infrared spectroscopy according to claim 1, characterized by The construction of an implicit manifold structure between samples according to the multi-level scattering feature representation comprises the following steps: performing nonlinear measurement calculation of a high-dimensional feature space according to the multi-level scattering feature representation, wherein the Mahalanobis distance between each pair of sample feature vectors is calculated, and a k-nearest neighbor graph reflecting local neighborhood relationship of data is constructed based on the calculated Mahalanobis distance, to obtain a preliminary similarity relationship between samples based on spectrum features. According to the original near-infrared spectrum curve of the sample and the preliminary similarity relationship, the cross-pasture geographical relationship is constrained and fused, the sample pairs belonging to the same pasture are forced to be connected, and the connection strength between different sample pairs is adjusted according to the priori adjustment coefficient of the fat deposition regularity difference of the parts, so that the enhanced adjacent relationship fused with the spectral similarity and the scene metadata constraint is obtained; According to the enhanced adjacent relationship, the high-order sample correlation is coded, each group of samples satisfying the enhanced adjacent relationship is defined as a super-edge, and the supergraph structure with the sample as the vertex and the similarity of the diffusion distance and the original near-infrared spectrum curve in the feature space as the edge is constructed.
4. The method of predicting intramuscular fat content of yak meat based on near infrared spectroscopy according to claim 1, characterized by According to the supergraph structure, the feature and the label information are cooperatively diffused, including: According to the supergraph structure, a high-order graph structure operator is constructed, the Laplacian matrix of the supergraph is calculated, and a mathematical operator describing the diffusion ability of the information on the supergraph structure is obtained. According to the intramuscular fat content value, an initial information field is constructed, the content value of the sample node with known intramuscular fat content value is normalized and taken as an initial temperature value, the initial temperature of the sample node with unknown content is set to zero, and the initial temperature field is boundary-constrained in combination with the intramuscular fat content distribution interval in the sample node with known intramuscular fat content value, so that the initial temperature field conforming to the physical diffusion law is obtained. According to the mathematical operator and the initial temperature field, the steady-state solution of the heat diffusion process is solved, the heat diffusion equation is discretized and iteratively solved until the temperature change of all nodes on the entire supergraph structure is less than a preset threshold, and the steady-state information value of each sample node is obtained.
5. The method of predicting intramuscular fat content of yak meat based on near infrared spectroscopy according to claim 1, characterized by According to the supergraph structure, the feature and the label information are cooperatively diffused, including: According to the multi-level scattering feature representation and the steady-state information value, a scene-adaptive feature fusion is performed, the steady-state information value of each sample is taken as an attention weight, the multi-level scattering feature representation of the corresponding sample is weighted, and the weight is priorly corrected in combination with the part metadata of the sample, so that a deep fusion feature vector fused with the topological steady-state information and the spectral feature is obtained; According to the deep fusion feature vector and the intramuscular fat content value, a graph structure constrained prediction function learning is performed, the sample is regarded as a graph node and the deep fusion feature vector thereof is taken as a node attribute, a preset graph neural network is trained under the constraint of the constructed supergraph topology structure, so that the mapping function from the node feature to the intramuscular fat content is learned while the neighborhood node features are aggregated, and an intramuscular fat content prediction model is obtained; According to the intramuscular fat content prediction model, the fat content of the to-be-predicted sample is determined, the deep fusion feature vector corresponding to the to-be-predicted sample and the topological relationship of the to-be-predicted sample in the supergraph are determined, and the deep fusion feature vector and the topological relationship are input into the intramuscular fat content prediction model, so that the intramuscular fat content prediction value of the to-be-predicted sample is obtained.
6. A system for predicting intramuscular fat content of yak meat based on near infrared spectroscopy, characterized by, It includes: An acquisition unit is configured to acquire original near-infrared spectrum curves and intramuscular fat content values of yak meat samples of multiple pastures, multiple parts and multiple age groups of individuals. The extraction unit is configured to perform multi-level feature extraction based on the original near-infrared spectrum curve, and each spectrum curve is processed by applying a wavelet scattering network, wherein the wavelet scattering network is subjected to convolution and modulus operation in a layer-by-layer manner by using a preset wavelet filter bank to obtain multi-level scattering feature representation; The construction unit is configured to construct an implicit manifold structure between samples based on the multi-level scattering feature representation, and the scattering feature of each sample is regarded as a point in a high-dimensional space, and a preset sample metadata is regarded as a geometric prior, so that a hypergraph structure is constructed, in which the samples are vertices, and the similarity between the samples is measured by diffusion distance in the feature space and metadata consistency; The simulation unit is configured to perform collaborative diffusion of features and label information based on the hypergraph structure, and a heat diffusion equation based on a Laplacian operator is defined on the hypergraph, and a known intramuscular fat content value is used as an initial heat source to simulate the conduction process of the intramuscular fat content value on the hypergraph structure over time, and a steady-state information value of each sample node is obtained by solving the equation, and the steady-state information value is fused with the topological structure and the initial label information; The prediction unit is configured to construct a prediction model based on the multi-level scattering feature representation and the steady-state information value, and predict the near-infrared spectrum curve of a to-be-predicted sample based on the constructed intramuscular fat content prediction model, and determine the intramuscular fat content value corresponding to the near-infrared spectrum curve of the to-be-predicted sample.
7. The near infrared spectroscopy-based system for predicting intramuscular fat content of yak meat according to claim 6, characterized by, The extraction unit comprises: The first extraction subunit is configured to perform time-frequency joint analysis of the original near-infrared spectrum curve, and the spectrum curve is convolved by using a preset Gabor wavelet filter with different scales and center frequencies, so that the spectrum signal in the spectrum curve is decomposed into a time-frequency representation in a wavelength-frequency two-dimensional plane, and a primary time-frequency coefficient containing local spectrum oscillation mode information is obtained; The second extraction subunit is configured to perform feature fusion for intramuscular fat spectrum response characteristics based on the primary time-frequency coefficient, and the coefficients of the time-frequency representation in the corresponding wavelength interval are weighted, aggregated and nonlinearly modulated by using the near-infrared absorption band of the preset known fat chemical bond as prior information, so as to obtain a robust feature representation fused with scene knowledge; The third extraction subunit is configured to construct deep stable features based on the robust feature representation fused with scene knowledge, and the robust feature representation is convolved and modulated with a wavelet filter bank, and is iteratively smoothed and down-sampled along the frequency axis, so that a multi-scale scattering feature representation of the spectrum signal is obtained, that is, the multi-level scattering feature representation.
8. The near infrared spectroscopy-based system for predicting intramuscular fat content of yak meat according to claim 6, characterized by, The construction unit comprises: The first construction subunit is configured to perform nonlinear distance calculation in a high-dimensional feature space based on the multi-level scattering feature representation, wherein the Mahalanobis distance between each pair of sample feature vectors is calculated, and a k-nearest neighbor graph reflecting the local neighborhood relationship of the data is constructed based on the calculated Mahalanobis distance, so as to obtain a preliminary similarity relationship between the samples based on the spectrum features. The second construction subunit is configured to perform constraint fusion of cross-pasture geographical relations according to the original near-infrared spectrum curve of the sample and the preliminary similarity relation, forcibly connect sample pairs belonging to the same pasture, and apply a prior adjustment coefficient to the connection strength between different sample pairs according to the difference in fat deposition regularity of the parts, to obtain an enhanced adjacency relation that fuses the spectral similarity and scene metadata constraint. The third construction subunit is configured to perform high-order sample correlation coding according to the enhanced adjacency relation, define each group of samples satisfying the enhanced adjacency relation as a hyperedge, and construct a hypergraph structure with the samples as vertices and the similarity of the diffusion distance in the feature space and the original near-infrared spectrum curve as edges.
9. The near infrared spectroscopy-based system for predicting intramuscular fat content of yak meat according to claim 6, characterized by, The simulation unit includes: The first simulation subunit is configured to perform high-order graph structure operator construction according to the hypergraph structure, calculate the Laplacian matrix of the hypergraph, and obtain a mathematical operator describing the diffusion capability of the information on the hypergraph structure. The second simulation subunit is configured to perform initial information field construction according to the intramuscular fat content value, normalize the content value of the sample node with the known intramuscular fat content value, take the normalized content value as an initial temperature value, set the initial temperature of the sample node with the unknown content value to zero, and combine the intramuscular fat content distribution interval in the sample node with the known intramuscular fat content value to perform boundary constraint on the initial temperature field, to obtain an initial temperature field conforming to the physical diffusion law. The third simulation subunit is configured to perform steady-state solution of the heat diffusion process according to the mathematical operator and the initial temperature field, discretize and iteratively solve the heat diffusion equation until the temperature change of all nodes on the entire hypergraph structure is less than a preset threshold, and obtain the steady-state information value of each sample node.
10. The near infrared spectroscopy-based system for predicting intramuscular fat content of yak meat according to claim 6, characterized by, The prediction unit includes: The first prediction subunit is configured to perform scene-adaptive feature fusion according to the multi-level scattering feature representation and the steady-state information value, take the steady-state information value of each sample as an attention weight, weight the multi-level scattering feature representation of the corresponding sample, combine the part metadata of the sample to perform prior correction on the weight, and obtain a deep fusion feature vector that fuses the topological steady-state information and the spectral feature. The second prediction subunit is configured to perform graph structure constraint prediction function learning according to the deep fusion feature vector and the intramuscular fat content value, take the sample as a graph node and the deep fusion feature vector of the sample as a node attribute, train a preset graph neural network under the constraint of the constructed hypergraph topology structure, make the graph neural network aggregate the neighbor node features while learning the mapping function from the node features to the intramuscular fat content, and obtain an intramuscular fat content prediction model. The third prediction subunit is configured to perform fat content estimation of a to-be-predicted sample according to the intramuscular fat content prediction model, determine the deep fusion feature vector corresponding to the to-be-predicted sample and the topological relation of the to-be-predicted sample in the hypergraph, and input the deep fusion feature vector and the topological relation into the intramuscular fat content prediction model, to obtain an intramuscular fat content prediction value of the to-be-predicted sample.
Citation Information
Patent Citations
Solution component concentration determination method and system based on two-dimensional spectrum
CN115184281A
Yak healthy feeding optimization method and system based on machine learning
CN121303612A