Multispectral imaging detection method and system for monitoring and judging probiotic activity based on metabolite
By employing multispectral imaging detection methods, principal component analysis, and spectral clustering techniques, the problems of long time consumption and insufficient specificity in the activity evaluation of probiotic preparations in existing technologies have been solved. This enables rapid and accurate evaluation of the distribution of active ingredients and provides comprehensive quality control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MEIYA SPECIAL MEDICAL (BEIJING) NUTRITION TECHNOLOGY CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies cannot quickly and accurately assess the functional state of live bacteria and the spatial distribution of metabolites in probiotic preparations. Traditional methods are time-consuming and lack spatial resolution, while existing spectroscopic techniques lack specificity and accuracy.
Using multispectral imaging detection methods, principal component analysis is employed for dimensionality reduction, spectral clustering, and feature screening to calculate a comprehensive activity score, thereby directly monitoring metabolites in probiotic liquid preparations and obtaining information on chemical composition and spatial distribution.
It enables rapid and accurate assessment of the distribution uniformity and functional status of active ingredients in probiotic preparations, provides a comprehensive quality control perspective, and improves the automation level of detection and the comparability of results.
Smart Images

Figure CN122016675A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of spectral detection of microbial metabolites, and in particular to a multispectral imaging detection method and system for determining probiotic activity based on metabolite monitoring. Background Technology
[0002] Probiotic preparations are used in many beverage products, including liquid formulations (such as fermented broths and suspensions) and solid formulations (such as freeze-dried powders, tablets, and capsules). Their core quality attributes and nutritional efficacy are based on the number of live probiotics and their metabolic activity. As the market demands increasing consistency in product functionality and quality, traditional testing methods that rely solely on the number of live bacteria are no longer sufficient to accurately and rapidly evaluate the actual functional status of probiotics.
[0003] Currently, the counting of viable bacteria in probiotic preparations mainly relies on the traditional plate culture method. While this method is considered the "gold standard," its drawbacks are particularly prominent: the detection time is as long as 48-72 hours, the operation is cumbersome, and it only reflects the number of culturable viable bacteria, completely failing to characterize the real-time metabolic activity and functional state of the bacteria. Furthermore, whether using chromatography or culture methods, the results are all overall average measurements of the sample, completely losing any information about the spatial distribution of active ingredients (viable bacteria or metabolites) in the preparation. For preparations, the uniformity of the distribution of active ingredients (whether there is aggregation or concentration gradient) is a key factor affecting batch consistency, dosage accuracy, and functional stability, but current technology is completely powerless in addressing this.
[0004] To overcome the limitations of traditional methods, some rapid detection technologies have been attempted in the field of probiotics, such as live bacteria detection technologies based on ATP bioluminescence or specific fluorescent dyes. However, these methods still rely on indirect inference of activity, cannot specifically correlate with direct functional products such as SCFAs, and are susceptible to interference from complex matrices, resulting in insufficient specificity and accuracy. In recent years, spectroscopic analysis techniques (such as near-infrared spectroscopy) have attracted attention due to their speed and non-destructive nature. However, traditional spectroscopic techniques acquire the average spectrum of a point or region on the sample, and the signal is a superposition and mixture of the spectra of all chemical components in the sample (live bacteria, dead bacteria, metabolites, carriers / culture media). This "mixed spectroscopy" method has two fundamental drawbacks: first, poor specificity, making it difficult to extract the weak characteristic signals of specific metabolites (such as SCFAs) with high sensitivity from complex backgrounds; second, a lack of spatial resolution, failing to identify and quantify the microscopic spatial distribution and heterogeneity of metabolites in the sample, which are precisely important dimensions for evaluating the function of the bacterial community and the quality of the formulation process. Summary of the Invention
[0005] The purpose of this invention is to provide a multispectral imaging detection method and system for determining probiotic activity based on metabolite monitoring, which solves the above-mentioned technical problems pointed out in the prior art.
[0006] This invention provides a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, comprising the following steps:
[0007] A cuboid of multispectral image data was collected from a probiotic liquid preparation sample. The cuboid of multispectral image data contains the pixel coordinates and their corresponding spectral vectors under multiple preset characteristic wavelengths. The preset characteristic wavelengths are the characteristic spectral response wavelengths of short-chain fatty acids.
[0008] Principal component analysis is used to reduce the dimensionality of the multispectral image data cuboid to obtain a feature spectral data matrix; the principal components in the feature spectral data matrix are used for clustering to obtain initial clusters; the feature spectral vectors in the feature spectral data matrix are used to obtain target spectral similarity clusters in the initial clusters; the intensity values of pixels at characteristic wavelengths are obtained for the target spectral similarity clusters, and a comprehensive activity score is calculated.
[0009] A preset activity threshold is set; the detection results of the probiotic liquid preparation sample are obtained by judging the comprehensive activity score and the activity threshold.
[0010] Accordingly, this invention also proposes a multispectral imaging detection system for determining probiotic activity based on metabolite monitoring, comprising: an acquisition module; an analysis module; and an identification module;
[0011] The acquisition module is used to acquire a cuboid of multispectral image data from a probiotic liquid preparation sample. The cuboid of multispectral image data contains the coordinates of pixels at multiple preset feature wavelengths and their corresponding spectral vectors.
[0012] The analysis module is used to reduce the dimensionality of the multispectral image data cuboid using principal component analysis to obtain a feature spectral data matrix; to perform clustering using the principal components in the feature spectral data matrix to obtain initial clusters; to obtain target spectral similarity clusters in the initial clusters using the feature spectral vectors in the feature spectral data matrix; and to obtain the intensity values of pixels at characteristic wavelengths for the target spectral similarity clusters and calculate a comprehensive activity score.
[0013] The identification module is used to preset an activity threshold; the detection result of the probiotic liquid preparation sample is obtained by judging the comprehensive activity score and the activity threshold.
[0014] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages:
[0015] Analysis of the multispectral imaging detection method and system for determining probiotic activity based on metabolite monitoring provided by the present invention shows that, in specific applications, multispectral imaging acquires a data cuboid containing spatial coordinates and full-spectrum information, which can simultaneously obtain chemical composition information and precise spatial distribution information of probiotic metabolites in the sample. This allows the detection results to not only reflect the amount of active ingredients, but also reveal whether the distribution of active ingredients is uniform, providing an unprecedented comprehensive perspective for quality control.
[0016] Furthermore, addressing the challenges of processing high-dimensional multispectral data, this method innovatively employs a four-step analysis strategy: PCA dimensionality reduction, spectral clustering, feature selection, and quantitative scoring. This strategy first effectively removes spectral redundancy and noise through PCA, preserving core chemical characteristics. Next, unsupervised clustering automatically identifies different material regions. Then, spectral matching precisely pinpoints feature regions related to high activity. Finally, mathematical integration generates a single comprehensive activity score. The entire process is highly automated, transforming complex image and spectral data into intuitive and comparable quantitative indicators, achieving a rapid closed loop from data to decision-making. In addition, the method directly monitors function-related metabolites, providing a more accurate reflection of the actual functional state of probiotic products compared to simple viable bacteria counting. Attached Figure Description
[0017] Figure 1 This is a flowchart of a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, as described in Example 1.
[0018] Figure 2 This is a schematic diagram of a probiotic liquid preparation sample from an example of a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, as described in Example 1.
[0019] Figure 3 This is a schematic diagram of the spectrum of a probiotic liquid preparation sample in Example 1, which is a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring.
[0020] Figure 4 This is a flowchart of a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, as described in Example 1, showing the target spectral similarity clusters.
[0021] Figure 5 This is a schematic diagram illustrating the screening of k initial cluster centers in a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, as described in Example 1.
[0022] Figure 6 This is a flowchart of the extended growth region of a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, as described in Example 1.
[0023] Figure 7 This is the iterative operation flow of initial clustering in a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, as described in Example 1.
[0024] Figure 8 This is a flowchart of a multispectral imaging detection system for determining probiotic activity based on metabolite monitoring, as described in Example 2.
[0025] Labels: Acquisition module 10; Analysis module 20; Identification module 30. Detailed Implementation
[0026] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings.
[0028] Example 1
[0029] like Figure 1 , 2 As shown in Figure 3, this embodiment of the invention provides a multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, comprising the following steps:
[0030] S10: Collect multispectral image data cuboids from probiotic liquid preparation samples. The multispectral image data cuboid (i.e., treating the colony as a cuboid) contains the pixel coordinates (i.e., the pixel coordinates on the spectral image plane with the x and y axes) and their corresponding spectral vectors (i.e., the spectral vectors are spectral bands with wavelengths) at multiple preset characteristic wavelengths. The preset characteristic bands are the characteristic spectral response bands of short-chain fatty acids (SCFA).
[0031] It should be noted that the probiotic liquid preparation sample is pretreated by uniformly coating it onto a glass slide to form the test film (i.e., the liquid is spread evenly, reducing spectral distortion caused by particle size and surface scattering; the uniform coating itself implicitly solves the x and y coordinate positions, making the subsequent search for enrichment areas more focused on the real differences in biological activity); the test film is scanned using a multispectral imaging system to acquire a cuboid of multispectral image data containing coordinate and spectral information;
[0032] A three-dimensional data structure was acquired from probiotic liquid preparation samples using a multispectral camera with set wavelengths. This acquired three-dimensional data structure includes the x-axis and y-axis coordinate systems and the wavelengths (z, λ), which together form a multispectral image data cuboid. The x-axis and y-axis coordinate systems represent the planar position of the image, similar to the pixel coordinates of a regular camera. The wavelengths represent the spectral dimension (spectral vector), indicating different wavelength channels, with each wavelength corresponding to a specific spectral band.
[0033] As mentioned above, short-chain fatty acids (SCFAs) are composed of carbon (C), hydrogen (H), and oxygen (O) atoms. The characteristic wavelengths related to the vibrations of the CH and C=O bonds in SCFA molecules are, for example, ~1720 cm⁻¹ (C=O stretching), ~2950 cm⁻¹, and ~2870 cm⁻¹ (CH stretching) in the mid-infrared region; and 1700-1800 nm and 2300-2500 nm (combination and combination absorption of CH and C=O) in the near-infrared region. However, research has found that these short-chain fatty acids (SCFAs), as "metabolites," exhibit typical characteristics in their spectral images, possessing clear spectral features. This makes the multispectral imaging method theoretically sound and reliable, allowing for identification and implementation.
[0034] S20: Dimensionality reduction of the multispectral image data cuboid is performed using principal component analysis to obtain a feature spectral data matrix; clustering is performed using the principal components in the feature spectral data matrix to obtain initial clusters; target spectral similarity clusters in the initial clusters are obtained using the feature spectral vectors in the feature spectral data matrix; intensity values of pixels at characteristic wavelengths are obtained for the target spectral similarity clusters, and a comprehensive activity score is calculated.
[0035] It should be noted that the spectral information of hundreds or thousands of related wavelength channels is compressed into a few (e.g., 3-5) independent principal components that can represent the majority of the data variance. Each principal component is a "feature image" representing a major variation pattern in the original spectrum. A clustering algorithm is used to process the feature spectral data matrix after PCA dimensionality reduction. All pixels in the image are automatically divided into several clusters based on the similarity of their spectral features. Each cluster represents a region of substances with similar chemical composition. Target spectral similarity clusters are obtained and their intensity values are calculated. Based on the initial clusters, the feature spectral vectors are further analyzed to select the cluster whose spectral shape is most similar to the "highly active" reference spectrum ("target spectral similarity cluster"). Then, the pixel intensity values of this cluster at all feature wavelengths are extracted. Precise positioning... Identify the image regions most likely representing highly active probiotics and their metabolites, and acquire their spectral intensity data. Initial clustering may include substances with similar spectra but different origins. This step ensures that the analyzed region is directly related to the target activity by matching the spectra with those of known active metabolites, thus improving specificity. The intensity value directly reflects the relative concentration or abundance of the characteristic metabolite. Mathematically integrate the characteristic wavelength intensity values of the selected "target spectral similarity clusters" (e.g., weighted summation) to generate a single, quantitative value (score) to characterize the probiotic activity level of the entire sample. Multidimensional spatial and spectral information (multiple regions, multiple wavelengths) is condensed into an intuitive and comparable indicator. The score considers not only the concentration (intensity) of the metabolite but also its spatial distribution range (cluster size).
[0036] S30: Preset activity threshold; the detection results of the probiotic liquid preparation sample are obtained by judging the comprehensive activity score and the activity threshold;
[0037] It should be noted that an activity threshold is set in advance, and the calculated comprehensive activity score is compared with this threshold. If the score is higher than the threshold, the activity is judged to be qualified; if it is lower than the threshold, the activity is judged to be insufficient or unqualified.
[0038] Specifically, such as Figure 4 As shown, in step S20, principal component analysis is used to reduce the dimensionality of the multispectral image data cuboid to obtain a feature spectral data matrix; the principal components in the feature spectral data matrix are used for clustering to obtain initial clusters; the feature spectral vectors in the feature spectral data matrix are used to obtain target spectral similarity clusters in the initial clusters; the intensity values of pixels at the feature wavelengths are obtained for the target spectral similarity clusters, and a comprehensive activity score is calculated. The specific operation steps are as follows:
[0039] S21: The multispectral image data cuboid is converted into a two-dimensional matrix using principal component analysis, and the two-dimensional matrix is standardized to obtain a standardized matrix.
[0040] Calculate the covariance for each row and column of the standardized matrix, and then transform the standardized matrix into a covariance matrix;
[0041] Principal component analysis is used to decompose the covariance matrix to obtain multiple eigenvalues (i.e., eigenvalues). The eigenvalue represents the variance of the standardized spectral data (i.e., the set of spectral vectors of all pixels) in each principal component direction; the first eigenvalue represents the projection variance of the data in the first principal component axis direction. This variance is the largest, indicating that the first principal component captures the concentration changes of the most abundant short-chain fatty acids in the most important samples in the data, and the subsequent eigenvalues decrease in turn. The largest (meaning the first principal component direction retains the most change information from the original data) and its corresponding eigenvector (i.e., eigenvector). It is an L-dimensional vector that defines the direction of the i-th principal component axis in the original L spectral band space. It can be regarded as a "spectral weight template" that indicates which original bands contribute significantly to the principal component.
[0042] Sort all eigenvalues in descending order, and rearrange the eigenvectors corresponding to the eigenvalues in the same order; calculate the sum of the first k eigenvalues in descending order, and further calculate the sum of all eigenvalues; divide the sum of the first k eigenvalues by the sum of all eigenvalues to calculate the cumulative contribution rate (i.e., the cumulative contribution rate is a value between 0 and 1). = (λ1+ λ2+ ... +λ k ) / (λ1+λ2+ ... +λ I ), where λ is the eigenvalue; for example, = 0.95 indicates that the first k principal components collectively retain 95% of the information (variance) of the original data; the eigenvalues obtained by PCA in this scheme are unordered; the descending order is to organize the principal components in descending order of information content (variance), and the operation always starts with the most important component; the sum of the first k eigenvalues represents how much of the original total variance (i.e., how much original information) the newly constructed projection matrix can retain if we only use the first k principal components (i.e., discard the last Lk components); the final descending order is to sort by importance, and the summation and calculation of the cumulative contribution rate are to provide an objective and quantitative stopping criterion for dimensionality reduction, indicating "how many principal components are enough", balancing the simplicity of the data and the integrity of the information;
[0043] Preset cumulative contribution rate threshold; determine whether each cumulative contribution rate is greater than or equal to the cumulative contribution rate threshold;
[0044] If so, then all the eigenvalues corresponding to the cumulative contribution rate are taken as principal components, and the eigenvectors corresponding to the principal components are extracted to construct the projection matrix (that is, the eigenvectors corresponding to the selected eigenvalues are taken as principal components, and these principal components (eigenvectors) are used to construct the projection matrix; the eigenvalue is a scalar representing the importance (variance) of the direction of the corresponding principal component, and it is not the principal component itself; the eigenvector is an L-dimensional vector that defines the direction of the principal component in the original spectral space (i.e., the most original spectral vector in S10), and it is the principal component).
[0045] The projection matrix is reduced in dimensionality using the normalized matrix to obtain the feature spectral data matrix;
[0046] It should be noted that the multispectral image data cuboid (which is constructed as a three-dimensional matrix containing pixel height, pixel width, and spectral dimension, assuming a spatial size of H (height) * W (width) pixels and a spectral dimension of L bands) is converted into a two-dimensional matrix x with shape (HW, L). Each row of this two-dimensional matrix represents a pixel, and each column represents a specific spectral vector (i.e., a spectral band). Each column (i.e., each spectral band) of the two-dimensional matrix x is independently normalized. For example, for all pixel gray values in the first column (band 1)... ( ), calculate the mean of this column. and standard deviation ; Use the mean of this column and standard deviation Calculation standardization ,in For each pixel, the output is a standardized matrix xs, but the mean of each column is 0 and the standard deviation is 1.
[0047] If the covariance matrix It is square matrix The first in Line number Column element unit Indicates the first The band and the first The covariance of each band is calculated using the following formula: ;in, It is the transpose (shape) of the normalized matrix Xs ), Represents matrix multiplication; because The mean of each column is 0; this calculation is equivalent to calculating the covariance between the columns.
[0048] The projection matrix P is reduced in dimensionality using the normalized matrix Xs to obtain the feature spectral data matrix; matrix multiplication Y = Xs * P is then performed; where Xs represents the normalized data matrix with shape (M, L); P represents the projection matrix with shape (L, N); and Y represents the dimensionality-reduced feature spectral data matrix with shape (M, N). Each row of Y corresponds to a pixel in the original data cuboid, representing the coefficients of that pixel on the N principal components. The N-dimensional vector is the feature spectral vector, which is simpler than the original L-dimensional vector (i.e., the original spectral vector, the spectral vector in the multispectral image data cuboid), but contains most of the effective information. In the above scheme, each row of the feature spectral data matrix Y corresponds to the N-dimensional feature spectral vector of a pixel; this feature spectral vector is obtained by projecting the original spectral vector through principal component analysis. Specifically, each column of the projection matrix P is a principal component direction (eigenvector), and the elements in Y... This represents the projected coordinates of the spectral data of the i-th pixel along the j-th principal component direction;
[0049] S22: Extract the first principal component of the first column and the last principal component of the last column from the feature spectral data matrix;
[0050] An initial cluster center candidate set is constructed by analyzing all feature spectral vectors of the corresponding rows of the first and last principal components.
[0051] For the initial cluster center candidate set, determine the feature spectral vector of the center point, calculate the Euclidean distance from all feature spectral vectors in the initial cluster center candidate set to the feature spectral vector of the center point, and select the feature spectral vector with the farthest Euclidean distance from the center point as the first initial cluster center.
[0052] The Euclidean distance between the first initial cluster center and all feature spectral vectors in the candidate set of initial cluster centers is calculated again. The feature spectral vector with the furthest Euclidean distance from the first initial cluster center is selected as the next initial cluster center, and so on, until k feature spectral vectors are selected as initial cluster centers. (That is, the first initial cluster center is calculated by determining the feature spectral vector of the center point, and then the feature spectral vector with the furthest Euclidean distance from the first initial cluster center is selected as the next initial cluster center, ensuring that the initial cluster centers are representative (i.e., each initial center represents a potential extreme type of data) and dispersed (i.e., since they come from different corners of the data space (the extremes of the principal components), the distance between them is likely to be large, which meets the condition of a high-quality initial center (wide distribution)). Figure 5 (as shown)
[0053] Calculate the distance from the feature spectral vector of each pixel in the feature spectral data matrix to the k initial cluster centers, select the initial cluster center with the minimum distance, and cluster the feature spectral vectors of all pixels within the minimum distance range to the initial cluster center to obtain the initial cluster.
[0054] It should be noted that in the above steps, the first principal component (PC1) represents the direction with the largest variance in the data, i.e., the main spectral feature that best distinguishes different categories; the last principal component (PCN) represents the direction with the smallest variance in the data, usually containing noise or very subtle spectral differences. This step selects points (feature spectral vectors) from the "two ends" of the data distribution (the extreme points with the largest and smallest information content), ensuring that the candidate points can cover the main boundaries of the data space and the unique categories that may exist. When step S22 is executed, an "initial cluster center candidate set" is constructed from the feature spectral vectors corresponding to the "first principal component" and the "last principal component". The distance from all points in the candidate set to a certain center point is calculated. The point farthest from the center point is selected as the first initial cluster center. Based on the selected centers, the point farthest from the selected center is repeatedly selected as the next center until K centers are selected, at which point the operation ends.
[0055] S23: Calculate the average vector of the feature spectral vectors in the initial cluster; find the new centroid in the corresponding initial cluster based on the average vector (i.e., the new centroid is the feature spectral vector with the same average feature spectral vector, which serves as the new cluster center and becomes the more representative center point of the cluster).
[0056] Calculate the vector distance between the new centroid (i.e., the new initial cluster center) and the initial cluster center with the minimum distance (i.e., the initially selected initial cluster center);
[0057] Set a convergence threshold; determine whether the vector distance is greater than the convergence threshold;
[0058] If so, then the initial cluster is determined as the final target spectral similarity cluster;
[0059] If not, then repeat step S22 based on the new centroid (i.e., repeat the steps starting from the above "calculate the distance from the feature spectral vector of each pixel in the feature spectral data matrix to the k initial cluster centers (i.e., the new centroids), filter the initial cluster centers (i.e., the new centroids) with the smallest distance, and cluster the feature spectral vectors of all pixels within the minimum distance range to the initial cluster centers to obtain the initial clusters;") until the final target spectral similarity cluster is obtained.
[0060] It should be noted that the above scheme improves the accuracy and stability of clustering results through iterative optimization; and by matching similarity with reference spectra, it accurately selects the target spectral similarity clusters from all clusters, i.e., the regions most likely to represent highly active probiotic metabolites. The above steps calculate the distance between the new centroid (average vector within the cluster) and the old centroid (i.e., the initial cluster center) and iterate by setting a convergence threshold, so that the cluster center continuously moves towards the true center of the points within the cluster until it stabilizes, ensuring the highest internal consistency of each cluster and more accurate segmentation results.
[0061] The above method calculates the matching degree (e.g., spectral angle) between the average spectral vector of each target spectral similarity cluster and the reference spectrum, using the following formula: ;in, It is the average spectral vector of the target cluster. It is the reference spectral vector, and "•" represents the dot product. By using the Euclidean norm (modulus) of the vector, clusters that are spectrally similar but not the target (such as certain excipients) are excluded, and the target metabolite enrichment region is specifically screened, which greatly improves the specificity and purposefulness of the method.
[0062] S24: Determine potential metabolite enrichment regions by analyzing the target spectral similarity clusters using feature spectral vectors; obtain active closed contours for the potential metabolite enrichment regions; determine seed point growth for the pixels within the active closed contours to obtain extended growth regions; obtain the intensity values of pixels at characteristic wavelengths for the extended growth regions, calculate the sample activity level index, further obtain the sample activity level index of the same batch of probiotic liquid preparation samples, and calculate the comprehensive activity score.
[0063] It should be noted that the identified target spectral similarity clusters are expanded from discrete sets of pixels into continuous active regions, and precisely quantified. Finally, multiple indicators are fused to generate a comprehensive score representing the overall sample activity. Because target spectral similarity clusters may be broken or discontinuous due to noise or threshold segmentation, region growth is performed using pixels within the cluster as seeds. This is used to merge neighboring similar pixels based on spatial adjacency and spectral similarity criteria to form complete and connected closed active contours (i.e., closed contours of metabolite-enriched regions), which better reflects the morphology of actual colonies or product-enriched regions. The method not only extracts the average intensity at characteristic wavelengths within the expanded region (reflecting the relative concentration of target metabolites), but also calculates the area / pixel count of the region (reflecting the enrichment range), the uniformity of intensity distribution, etc., thus quantifying the activity state more comprehensively from both concentration and distribution dimensions.
[0064] The activity level index of each formulation sample in the same batch is calculated separately, and then these indexes are integrated (e.g., by averaging) to obtain the comprehensive activity score of the batch. This eliminates the random errors that may exist in individual samples and evaluates the overall batch level, making the conclusion more reliable. This score is the direct basis for comparing with the preset threshold and making the final pass / fail judgment.
[0065] Specifically, such as Figure 6 As shown, in step S24, potential metabolite enrichment regions are determined by the target spectral similarity clusters using feature spectral vectors; active closed contours are obtained for the potential metabolite enrichment regions; seed points are determined for the pixels inside the active closed contours to grow, resulting in extended growth regions; the intensity values of pixels under characteristic wavelengths are obtained for the extended growth regions, and the sample activity level index is calculated. Furthermore, the sample activity level indexes of the same batch of probiotic liquid preparation samples are obtained, and a comprehensive activity score is calculated. The specific operation steps are as follows:
[0066] S241: Calculate the difference vector of the mean vector for each feature spectral vector in the target spectral similarity cluster, and further calculate the square norm of the difference vector;
[0067] The squared norms of all difference vectors are calculated and summed. Based on this sum, a scalar variance value is determined (i.e., a scalar variance value corresponding to each target spectral similarity cluster; the smaller this value, the more tightly the pixels within that cluster are clustered in the dimensionality-reduced feature space, and the higher the spectral consistency). For example, for a cluster containing m feature spectral vectors... For a cluster whose mean vector is μ, then the scalar variance value is... ;
[0068] Set a tightness threshold; determine whether the scalar variance value is less than the tightness threshold;
[0069] If so, the target spectral similarity cluster is determined to be a potential metabolite enrichment region (i.e., if the low variance cluster has highly consistent spectra, it is likely to correspond to a region with uniform and continuous material composition in the image, such as a region with uniform deposition of metabolites, a pure background matrix, etc.; if the high variance cluster has large spectral differences, it may correspond to a mixed region of different substances, an edge transition zone, or a noisy region containing multiple spectral types. These regions are not suitable as analytical targets representing a single metabolic activity).
[0070] It should be noted that the scalar variance value within each target spectral similarity cluster is calculated and compared with a preset tightness threshold. Clusters with highly consistent internal spectra and low variance are selected and identified as potential metabolite enrichment regions. The target spectral similarity clusters identified in the previous step undergo a first quality control process, excluding regions that, although their average spectra match the reference spectra, have chaotic internal spectra and high heterogeneity. Research has shown that an ideal enrichment region representing a specific metabolite should have highly consistent internal spectra (i.e., uniform metabolite concentration and composition). High scalar variance indicates that the region may be a mixture of different substances, a blurred transition zone, or heavily influenced by noise. Heavy interference, using these regions as analytical targets, will introduce huge errors and instabilities. Therefore, only clean regions with uniform internal spectra (i.e., low scalar variance values, which are less than the tightness threshold) have a clear physicochemical significance in their spatial morphology (outline). Only then can the results of subsequent spatial operations such as outline extraction and region growing be reliable. Therefore, this application needs to calculate the scalar variance value inside each "target spectral similarity cluster" and compare it with the preset tightness threshold to screen out clusters with highly consistent internal spectra and low variance, which are identified as potential metabolite enrichment regions. Regions that, although their average spectra match the reference spectra, have chaotic internal spectra and high heterogeneity are excluded.
[0071] S242: Obtain the set of pixel coordinates of the suspected enriched region (that is, determine the coordinates of the pixels in the suspected enriched region based on the coordinates of the pixels in the target spectral similarity cluster, and form a coordinate set of all coordinates, which is the coordinate set of the suspected enriched region).
[0072] All suspected enriched regions are clustered to form spectrally uniform regions (i.e., these suspected enriched regions have highly similar spectra; a spectrally uniform region is a collection of discrete pixels in the image. These pixels have high consistency in spectral features and may be spatially disconnected (because they belong to different clusters, which may be spatially separated). Therefore, it is region clustering, not pixel clustering).
[0073] Morphological operations and edge detection algorithms are performed on the spectral uniform region to obtain the active closed contour (i.e., since the pixels in the spectral uniform region are discrete, morphological dilation and erosion operations are performed to smooth the boundary of the spectral uniform region. After edge detection, the pixels at the boundary of the spectral uniform region form a complete closed contour, which is the boundary of the active associated region. The specific detection is common knowledge and will not be described in detail).
[0074] The aforementioned morphological operations (dilation / erosion) are used to connect regions, smooth boundaries, and output a continuous closed contour. Adjacent suspected enriched pixels are grouped into the same spatially connected region, each representing an independent candidate enriched region. Mathematical morphological processing (such as moderate dilation and erosion using circular structuring elements) is applied to each candidate region. The dilation operation bridges minor breaks within the region caused by noise or thresholding, connecting very close subregions; the erosion operation shrinks boundaries that have been overextended by dilation, bringing them closer to the actual edges of the region while smoothing irregular contours.
[0075] It should be noted that all discrete pixels from the selected suspected enriched regions are clustered into a set of spectrally uniform regions. Morphological operations (dilation / erosion) and edge detection are then performed on this set to generate smooth, closed, active, and unidirectional contours. Scattered, potentially disconnected, hyperspectrally consistent pixel groups are transformed into one or more regions with clear, complete, and continuous spatial boundaries. Based on the clustering results of spectral similarity, these pixels may not be spatially connected (e.g., the same metabolite may be distributed across multiple isolated patches in the sample). This step spatially re-aggregates these spectrally similar pixels, treating them as the same spectrally uniform region. True metabolite-rich regions should have a continuous morphology in the image. The aforementioned morphological operations (dilation connecting neighboring points, then erosion restoring approximate size) are used to fill small holes, connect neighboring patches, and smooth irregular boundaries, making the region shape more natural. The closed circular contours formed after edge detection provide a precise and quantifiable spatial definition for subsequent area calculations, pixel statistics within the region, and region growth.
[0076] S243: Randomly select pixels inside the active closed contour (i.e., internal pixels in the spectrally uniform region excluding contour points at the edges) as seed points.
[0077] Obtain the corresponding feature spectral vector in the feature spectral data matrix for the seed point; calculate the spectral similarity between the seed point and the feature spectral vectors corresponding to the neighboring pixels;
[0078] A preset similarity threshold is set; it is then determined whether the spectral similarity is greater than the similarity threshold.
[0079] If so, then grow the adjacent pixel, use the adjacent pixel as a new seed point and repeat the growth step until growth stops, thus obtaining an expanded growth region.
[0080] It should be noted that seed points are randomly selected within the active closed contour, and the region is grown into its neighborhood based on spectral similarity, incorporating pixels with similar spectral features and spatial proximity to form an extended growth region. Based on the initial contour, all connected pixels with the same spectral characteristics as the core region are precisely captured, potentially slightly expanding or precisely correcting the initial contour. Morphological operations, based on geometric rules, may incorrectly include spatially adjacent pixels with different spectra or omit pixels with the same spectrum but blocked by a threshold. Region growth uses spectral similarity as the fundamental criterion, achieving more accurate semantic segmentation. The true boundary of the extended growth region should be determined by the abrupt change points of spectral features, rather than a fixed geometric distance. The seed growth algorithm can automatically detect positions where the spectral similarity is below the threshold and stop, thus finding the most natural region boundary that best conforms to the definition of spectral consistency.
[0081] The above steps reduce the dependence on the position of a single seed point by randomly selecting multiple seed points from inside the contour and taking the union of the results, thus making the final region more stable.
[0082] S244: Perform a binarization mask on the extended growth region to obtain the extended growth region mask, and obtain the intensity value of the pixel at the characteristic wavelength to calculate the sample activity level index; obtain the sample activity level index of the same batch of the probiotic liquid preparation sample, and calculate the batch uniformity index; determine the allowable fluctuation range of the batch uniformity index, and screen the final metabolite enrichment region; obtain the intensity value of the pixel at the characteristic wavelength of the final metabolite enrichment region to calculate the comprehensive activity score.
[0083] It should be noted that the activity level indicators (such as average intensity, area, etc.) of the extended growth region of a single sample are calculated; the indicators of multiple samples in the same batch are statistically analyzed to calculate the batch uniformity index, and the final metabolite enrichment regions that meet the consistency requirements are screened out accordingly; based on these final metabolite enrichment regions, the comprehensive activity score of the batch is calculated, which elevates the analysis from a single sample to the batch as a whole. Intra-batch consistency verification is introduced before the final scoring to ensure that the data on which the scoring is based is reliable and representative; the activity of the formulation in the same batch should be relatively uniform. The above steps, by calculating the batch uniformity index, are used to identify and exclude outlier regions caused by individual sample preparation, imaging, or local contamination, ensuring that the regions used for the final scoring can truly reflect the general level of the batch;
[0084] The comprehensive activity score in the above steps is an integrated endpoint indicator, which is calculated by weighting factors such as the average spectral intensity (concentration), total area (enrichment degree) of the final region of each sample after screening, and batch uniformity. Such a score not only quantifies how strong the activity is, but also covers how wide the activity distribution is and how stable the batch quality is, providing an extremely comprehensive and robust basis for production quality control and release decisions.
[0085] Specifically, such as Figure 7 As shown, in step S244, a binarized mask is applied to the extended growth region to obtain the extended growth region mask, and the intensity values of pixels under the characteristic wavelength are obtained to calculate the sample activity level index; the sample activity level index of the same batch of probiotic liquid preparation samples is obtained, and the batch uniformity index is calculated; the allowable fluctuation range of the batch uniformity index is determined, and the final metabolite enrichment region is screened; the intensity values of pixels under the characteristic wavelength are obtained for the final metabolite enrichment region to calculate the comprehensive activity score. The specific operation steps are as follows:
[0086] S2441: Obtain a single-band grayscale image of the reflectance of the multispectral image data cuboid at the characteristic wavelength of the characteristic band (that is, a grayscale image of the multispectral image data cuboid at the maximum absorption wavelength; at this specific wavelength, the grayscale value (e.g., 0-255) of each pixel in the image reflects the intensity of light absorbed at that point).
[0087] Reference image of the preset characteristic band reference wavelength (i.e., single-band grayscale image of the non-absorbed wavelength, at which the colorimetric product hardly absorbs light, and the change in light intensity mainly reflects the background characteristics of the sample, such as the smoothness, roughness, and unevenness of the illumination of the sample surface).
[0088] The difference index image is calculated using the single-band grayscale image and the reference image, and is used as the reaction intensity image. The formula is as follows:
[0089]
[0090] in, Represented as a reaction intensity image, the gray value of each pixel in the reaction intensity image directly corresponds to the concentration of short-chain fatty acids at that spatial location (i.e., the three-dimensional spatial location along the x, y, and z axes; the reaction intensity image is essentially extracted from the cuboid of multispectral image data, which is a three-dimensional space along the x, y, and z axes), i.e., the original concentration of SCFA. Represented as a reference image; Represented as a single-band grayscale image;
[0091] The extended growth region is binarized and masked to obtain the extended growth region mask;
[0092] The average gray value of all pixels within the mask of the extended growth region in the reaction intensity image (i.e., the average concentration of short-chain fatty acids) is calculated as an indicator of the relative concentration of short-chain fatty acids.
[0093] The intensity values of all pixels in the extended growth region mask are obtained according to the preset characteristic wavelength (that is, the intensity values of the pixels can be directly obtained through the preset characteristic wavelength in S10, reflecting the absorption / reflection intensity of the metabolites under the characteristic absorption wavelength).
[0094] The average spectral response intensity of pixels in the extended growth region mask is calculated using the intensity values of pixels at characteristic wavelengths.
[0095] The relative concentration of short-chain fatty acids (SCFAs) and the average spectral response intensity are weighted and calculated to obtain the sample activity level index (i.e., the activity index of the probiotic liquid preparation sample, representing the average spectral response intensity of the suspected enrichment region of metabolites in the sample at a characteristic wavelength. Generally, the higher the concentration of short-chain fatty acids in the metabolites, the stronger (or weaker, depending on the specific substance) the signal at the characteristic absorption wavelength; therefore, the average spectral response intensity can be used as a proxy index of metabolic activity; that is, short-chain fatty acids serve as a marker of activity. This is because the core function of many important probiotics (such as Clostridium butyricum, Ferencella ferruginea, and certain Lactobacillus strains) is to ferment dietary fiber to produce SCFAs, especially butyric acid and propionic acid. If these bacteria are inactivated, this SCFA-producing metabolic pathway stops; therefore, the presence and concentration of SCFAs directly indicate the presence and metabolic intensity of these specific probiotics).
[0096] It should be noted that by using the extended growth region mask, the average spectral response intensity of pixels within that region at a preset characteristic wavelength is calculated as an indicator of the sample activity level of a single sample. This condenses the complex spectral information of a spatial region (composed of hundreds or thousands of pixels) into a single quantitative value directly related to the concentration of metabolites. The average intensity value mentioned above is a standard and comparable metric, enabling subsequent batch statistics. Since the mask region is obtained through layers of screening (spectral matching, variance screening, and spatial growth), this average intensity value highly specifically represents the signal level of the target metabolite, eliminating the influence of background and interfering substances.
[0097] S2442: Obtain the sample activity level index of each sample in the same batch of the probiotic liquid preparation sample by performing steps S21-S2441.
[0098] For each formulation sample in the same batch, the batch average (representing the average metabolic activity level of the batch samples, used as an overall evaluation of the batch activity) and batch standard deviation (measuring the dispersion of activity values of individual samples within a batch; a larger batch standard deviation indicates greater differences between samples and poorer batch consistency; a smaller batch standard deviation indicates smaller differences between samples and better batch consistency) are calculated.
[0099] The batch uniformity index (i.e., ...) is calculated using the batch average and batch standard deviation. = / ;in, Expressed as the average value of the batch; Expressed as batch standard deviation; This is expressed as a batch uniformity index; when the batches are very uniform, the batch standard deviation is... Very small Batch average Large; batch standard deviation when batches are uneven. Very large Batch average Very small; therefore, the batch average A higher value indicates better batch uniformity.
[0100] It should be noted that for multiple samples in the same batch, the average value of their "sample activity level index" is calculated, and the batch uniformity index is defined by dividing the average intensity by the standard deviation, generating the batch average activity level and batch uniformity index. The batch average value directly reflects the average activity level of the product in that batch and is the core of quality rating. A high-quality batch requires not only high average activity but also small activity differences between individual samples (i.e., high uniformity). A high mean, low standard deviation, and large uniformity index value indicate high and consistent batch activity, which is the ideal state. A low mean, high standard deviation, and small uniformity index value indicate low and unstable batch activity and poor quality. This index penalizes both low activity and high volatility, and can reflect batch quality more comprehensively than simply looking at the average value or standard deviation.
[0101] S2443: Collect the qualified average uniformity value (that is, the standard average uniformity value) and allowable fluctuation range (that is, the allowable upper and lower fluctuation threshold) of historical reference batches.
[0102] Determine whether the batch uniformity index is outside the allowable fluctuation range;
[0103] If so (i.e., exceeding the allowable fluctuation range), it is determined that there are differences in the sample activity level indicators among each formulation sample in the same batch (i.e., the batch uniformity index is less than the allowable fluctuation range, the batch uniformity is too low, and the difference between samples is too large. The possible reasons are poor clustering effect, which may have mistakenly excluded some active regions or included too many heterogeneous regions, resulting in inaccurate calculation of the batch uniformity index of each sample and increasing the intra-batch difference; or the batch uniformity index is greater than the allowable fluctuation range, the batch uniformity is too high, exceeding the normal fluctuation, which may be due to overly conservative clustering, the segmented region is too small or the real difference is missed, which may have masked the real intra-batch variation, or the region segmentation is incomplete; that is, when the batch uniformity index exceeds the allowable range, the system automatically determines that the current regional results obtained based on the original clustering are unreliable and the clustering process needs to be re-optimized).
[0104] Returning to step S22 above, the K-means algorithm is used to randomly select a feature spectral vector from the feature spectral data matrix as the initial cluster center. At this point, step S22 is re-entered for a reselection operation. Principal component analysis is no longer used, but the return to step S22, "re-selecting a feature spectral vector as the initial cluster center," corresponds to the "first initial cluster center" in step S22.
[0105] Calculate the distance between the initial cluster center and the remaining feature spectral vectors in the feature spectral data matrix, and further calculate the squared distance between the feature spectral vectors;
[0106] Calculate the sum of squared distances of all feature spectral vectors based on the squared distances of the feature spectral vectors; calculate the probability value of the next feature spectral vector as the initial cluster center based on the squared distances and the sum of squared distances of the feature spectral vectors (i.e., generate a random number, select the next center point according to the cumulative probability distribution, the farther the point is from the selected center, the greater the probability of being selected).
[0107] K initial cluster centers are selected based on the probability values of the initial cluster centers (that is, the probability values are calculated again based on the next feature spectral vector as the new initial cluster center until K initial cluster centers are selected).
[0108] For example, returning to step S22, the principal component analysis algorithm is no longer used (and the farthest distance selection method is abandoned). Instead, a new feature spectral vector is randomly selected from the feature spectral data matrix as the first initial cluster center c1. For each feature spectral vector x in the matrix that has not been selected as a center, its shortest distance to the first initial cluster center c1 (i.e., the Euclidean distance to the nearest center) is calculated and denoted as D(x). The probability that each feature spectral vector x will be selected as the next center is calculated: , where the denominator is the sum of D(z)² of all characteristic spectral vectors; based on the calculated probability distribution;
[0109] Continue by calculating the distance from the feature spectral vector of each pixel in the feature spectral data matrix to the k initial cluster centers, selecting the initial cluster center with the minimum distance, and clustering the feature spectral vectors of all pixels within the minimum distance range to the initial cluster center to obtain the initial cluster (i.e., the continued execution of the steps is part of S22); and repeat the above steps S23-S24 to obtain a new batch uniformity index, determine that the new batch uniformity index is within the allowable fluctuation range, and obtain the extended growth region mask as the final metabolite enrichment region (i.e., a binary image, where pixels with a value of 1 represent the finally determined metabolite enrichment region; the new batch uniformity index can be directly determined to be within the allowable fluctuation range because the new batch uniformity index is double-verified by the principal component selection in the original step S22 and the new optimized K-means algorithm, so the re-iteration of the K-means algorithm maximizes the accuracy of clustering, so that the final batch uniformity index reaches the standard range).
[0110] It should be noted that the calculated batch evenness index is compared with the allowable fluctuation range of historical qualified batches. If it exceeds the range, a feedback optimization loop is triggered to return to S22, and the cluster centers are reinitialized using the K-means algorithm. All subsequent analyses are then re-executed until the newly obtained batch evenness index falls within the qualified range. Finally, a mask of metabolite-enriched regions is obtained, ensuring that the analysis method can automatically adjust to obtain stable and reliable region segmentation results when facing different batches and different samples.
[0111] Simultaneously, the final output (batch uniformity) is linked to the core analysis step (cluster initialization). When the final result (batch uniformity) is abnormal, the system does not simply issue an alarm, but automatically traces the possible cause (poor initial clustering) and attempts to optimize it. The above steps, by introducing the K-means algorithm (selecting initial centers based on distance probability, making the centers far apart), significantly improve the stability and quality of the clustering results, thereby reducing abnormal fluctuations in the batch sample analysis results caused by the randomness of the algorithm. After this step, the metabolite enrichment region used for the final calculation has been verified for batch consistency, ensuring that the analysis results of different times and different batches are based on the same optimized analysis standard, making the comprehensive activity score highly reliable.
[0112] S2444: Calculate the area of the region where the metabolites are enriched (i.e., use the number of pixels as the area).
[0113] The intensity values of pixels in the region where metabolites are enriched are obtained, and the average spectral intensity is calculated (i.e., this step is the same as obtaining the pixel intensity values in S2441 above, and will not be repeated here).
[0114] The comprehensive activity score is obtained by weighting the area of the region and the average spectral intensity.
[0115] It should be noted that for the finally confirmed metabolite enrichment region, two dimensions of indicators are extracted: region area (A) and average spectral intensity (I). The above are used to calculate a comprehensive activity score through weighted calculation (e.g., Score = w1*A + w2*I), where w1 and w2 are weights, which comprehensively characterize the activity status of the probiotic preparation. The weights w1 and w2 are preset values determined by experience. The region area (A) reflects the total range of the active metabolite enrichment region. Even if the average intensity is high, if the area is small, the overall contribution is limited. The average spectral intensity (I) reflects the average concentration of metabolites in the region.
[0116] Example 2
[0117] like Figure 8 As shown, this embodiment of the invention also provides a multispectral imaging detection system for determining probiotic activity based on metabolite monitoring, comprising: a data acquisition module 10; an analysis module 20; and an identification module 30.
[0118] The acquisition module 10 is used to acquire a cuboid of multispectral image data from a probiotic liquid preparation sample. The cuboid of multispectral image data contains the coordinates of pixels under multiple preset feature wavelengths and their corresponding spectral vectors.
[0119] The analysis module 20 is used to reduce the dimensionality of the multispectral image data cuboid using principal component analysis to obtain a feature spectral data matrix; to perform clustering using the principal components in the feature spectral data matrix to obtain an initial cluster; to obtain a target spectral similarity cluster in the initial cluster using the feature spectral vectors in the feature spectral data matrix; and to obtain the intensity value of the pixel at the feature wavelength for the target spectral similarity cluster and calculate a comprehensive activity score.
[0120] The identification module 30 is used to preset an activity threshold; the detection result of the probiotic liquid preparation sample is obtained by judging the comprehensive activity score and the activity threshold.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; those skilled in the art can modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multispectral imaging detection method for determining probiotic activity based on metabolite monitoring, characterized in that, The following steps are included: A cuboid of multispectral image data was collected from a probiotic liquid preparation sample. The cuboid of multispectral image data contains the pixel coordinates and their corresponding spectral vectors under multiple preset characteristic wavelengths. The preset characteristic wavelengths are the characteristic spectral response wavelengths of short-chain fatty acids. Principal component analysis is used to reduce the dimensionality of the multispectral image data cuboid to obtain a feature spectral data matrix; the principal components in the feature spectral data matrix are used for clustering to obtain initial clusters; the feature spectral vectors in the feature spectral data matrix are used to obtain target spectral similarity clusters in the initial clusters; the intensity values of pixels at characteristic wavelengths are obtained for the target spectral similarity clusters, and a comprehensive activity score is calculated. Preset activity threshold; The detection results of the probiotic liquid preparation sample are obtained by judging the comprehensive activity score and the activity threshold.
2. The multispectral imaging detection method for determining probiotic activity based on metabolite monitoring according to claim 1, characterized in that, Principal component analysis was used to reduce the dimensionality of the multispectral image data cuboid to obtain the feature spectral data matrix. The specific steps are as follows: The multispectral image data cuboid is converted into a two-dimensional matrix using principal component analysis, and the two-dimensional matrix is then standardized to obtain a standardized matrix. Calculate the covariance for each row and column of the standardized matrix, and then transform the standardized matrix into a covariance matrix; The covariance matrix is decomposed using principal component analysis to obtain multiple eigenvalues; Sort all eigenvalues in descending order, and rearrange the corresponding eigenvectors in the same order; calculate the sum of the first k eigenvalues in descending order, and then calculate the sum of all eigenvalues. The cumulative contribution rate is calculated by combining the sum of the first k eigenvalues with the sum of all eigenvalues. Preset cumulative contribution rate threshold; determine whether each cumulative contribution rate is greater than or equal to the cumulative contribution rate threshold; If so, then all the feature values corresponding to the cumulative contribution rate are taken as principal components, and the feature vectors corresponding to the principal components are extracted to form a projection matrix. The projection matrix is reduced in dimensionality using the normalized matrix to obtain the feature spectral data matrix.
3. The multispectral imaging detection method for determining probiotic activity based on metabolite monitoring according to claim 2, characterized in that, Clustering is performed using the principal components in the feature spectral data matrix to obtain initial clusters; target spectral similarity clusters within the initial clusters are then obtained using the feature spectral vectors in the feature spectral data matrix. The specific steps are as follows: Extract the first principal component of the first column and the last principal component of the last column from the feature spectral data matrix; construct an initial cluster center candidate set from all feature spectral vectors in the rows corresponding to the first and last principal components; For the initial cluster center candidate set, determine the feature spectral vector of the center point, calculate the Euclidean distance from all feature spectral vectors in the initial cluster center candidate set to the feature spectral vector of the center point, and select the feature spectral vector with the farthest Euclidean distance from the center point as the first initial cluster center. The Euclidean distance between the first initial cluster center and all feature spectral vectors in the initial cluster center candidate set is calculated again. The feature spectral vector with the farthest Euclidean distance from the first initial cluster center is selected as the next initial cluster center, until k feature spectral vectors are selected as initial cluster centers. Calculate the distance from the feature spectral vector of each pixel in the feature spectral data matrix to the k initial cluster centers, select the initial cluster center with the minimum distance, and cluster the feature spectral vectors of all pixels within the minimum distance range to the initial cluster center to obtain the initial cluster. Calculate the average vector of the feature spectral vectors in the initial cluster; find the new centroid in the corresponding initial cluster based on the average vector; Calculate the vector distance between the new centroid and the initial cluster center with the minimum distance; preset a convergence threshold; determine whether the vector distance is greater than the convergence threshold; If so, then the initial cluster is determined as the final target spectral similarity cluster; If not, repeat the above steps based on the new centroid until the final target spectral similarity cluster is obtained.
4. The multispectral imaging detection method for determining probiotic activity based on metabolite monitoring according to claim 3, characterized in that, The intensity values of pixels at characteristic wavelengths are obtained for the target spectral similarity clusters, and a comprehensive activity score is calculated. The specific operation steps are as follows: Potential metabolite enrichment regions are identified by analyzing the target spectral similarity clusters using feature spectral vectors. Active closed contours are obtained for these regions. Seed points are determined for the pixels within the active closed contours to grow extended growth regions. Intensity values of pixels at characteristic wavelengths are obtained for these extended growth regions, and sample activity level indicators are calculated. Furthermore, the sample activity level indicators of the same batch of probiotic liquid preparation samples are obtained, and a comprehensive activity score is calculated.
5. The multispectral imaging detection method for determining probiotic activity based on metabolite monitoring according to claim 4, characterized in that, Potential metabolite enrichment regions are identified by analyzing the target spectral similarity clusters using feature spectral vectors; active closed contours are then obtained for these potential metabolite enrichment regions. The specific steps are as follows: For each feature spectral vector in the target spectral similarity cluster, calculate the difference vector of the mean vector, and further calculate the square norm of the difference vector; calculate and sum the square norms of all difference vectors, and determine the scalar variance value in one step based on the sum; Set a tightness threshold; determine whether the scalar variance value is less than the tightness threshold; If so, the target spectral similarity cluster is determined to be a potential region of suspected metabolite enrichment; The two-dimensional coordinates of the pixels corresponding to the feature spectral vectors in the target spectral similarity cluster are determined; the set of pixel coordinates of suspected enriched regions is determined based on the two-dimensional coordinates of the corresponding pixels; all suspected enriched regions are clustered to form spectrally uniform regions. Morphological operations are performed on the spectrally uniform region to obtain an active closed profile.
6. The multispectral imaging detection method for determining probiotic activity based on metabolite monitoring according to claim 5, characterized in that, Seed points are determined for the pixels within the active closed contour to grow an extended growth region; the intensity values of the pixels under a characteristic wavelength are obtained from the extended growth region, and the sample activity level index is calculated. Furthermore, the sample activity level indexes of the same batch of probiotic liquid preparation samples are obtained, and a comprehensive activity score is calculated. The specific operation steps are as follows: Randomly select pixels inside the active closed contour as seed points; Obtain the corresponding feature spectral vector in the feature spectral data matrix for the seed point; calculate the spectral similarity between the seed point and the feature spectral vectors corresponding to the neighboring pixels; A preset similarity threshold is set; it is then determined whether the spectral similarity is greater than the similarity threshold. If so, then grow the adjacent pixel, use the adjacent pixel as a new seed point and repeat the growth step until growth stops, thus obtaining an expanded growth region. The extended growth region is binarized and masked to obtain the extended growth region mask, and the intensity value of the pixel at the characteristic wavelength is obtained to calculate the sample activity level index; the sample activity level index of the same batch of the probiotic liquid preparation sample is obtained to calculate the batch uniformity index. The allowable fluctuation range of the batch uniformity index is determined, and the final metabolite enrichment region is screened; the intensity value of the pixel at the characteristic wavelength of the final metabolite enrichment region is obtained to calculate the comprehensive activity score.
7. The multispectral imaging detection method for determining probiotic activity based on metabolite monitoring according to claim 6, characterized in that, The extended growth region is binarized and masked to obtain the extended growth region mask. The intensity values of the pixels at the characteristic wavelength are obtained, and the sample activity level index is calculated. The specific operation steps are as follows: A single-band grayscale image of reflectance at the characteristic wavelength of the characteristic band is obtained by processing the cuboid of the multispectral image data; a reference image of the reference wavelength of the preset characteristic band is also obtained. The difference index image is calculated using the single-band grayscale image and the reference image, and is used as the reaction intensity image. The extended growth region is binarized and masked to obtain the extended growth region mask; Calculate the average gray value of all pixels within the mask of the extended growth region in the reaction intensity image, and use it as an indicator of the relative concentration of short-chain fatty acids; The intensity values of all pixels in the extended growth region mask at the preset characteristic wavelength are obtained; the average spectral response intensity of the pixels in the extended growth region mask is calculated using the intensity values; and the sample activity level index is obtained by weighting the relative concentration index of the short-chain fatty acids with the average spectral response intensity.
8. The multispectral imaging detection method for determining probiotic activity based on metabolite monitoring according to claim 7, characterized in that, The activity level index of the same batch of probiotic liquid formulation samples was obtained, and the batch uniformity index was calculated. The allowable fluctuation range of the batch uniformity index was determined, and the final metabolite enrichment region was screened. The specific operation steps are as follows: To obtain the activity level index of each sample in the same batch of the probiotic liquid preparation, the above steps are performed on the sample samples in the same batch; the batch average and batch standard deviation of the activity level index of each sample in the same batch are calculated. The batch uniformity index is obtained by calculating the batch average and batch standard deviation. Collect the average uniformity value and allowable fluctuation range of qualified samples from historical reference batches; Determine whether the batch uniformity index is outside the allowable fluctuation range; If so, then it is determined that there are differences in the sample activity level indicators among each formulation sample in the same batch; Returning to the steps above, the K-means algorithm is used to reselect a feature spectral vector from the feature spectral data matrix as the initial cluster center; Calculate the distance between the initial cluster center and the remaining feature spectral vectors in the feature spectral data matrix, and further calculate the squared distance of the feature spectral vectors; Calculate the sum of squared distances of all feature spectral vectors based on the squared distances of the feature spectral vectors; calculate the probability value of the next feature spectral vector as the initial cluster center based on the squared distances and the sum of squared distances of the feature spectral vectors. K initial cluster centers are selected based on the probability values of the initial cluster centers; Continue by calculating the distance from the feature spectral vector of each pixel in the feature spectral data matrix to the k initial cluster centers, filter out the initial cluster centers with the smallest distance, and cluster the feature spectral vectors of all pixels within the minimum distance range to the initial cluster centers to obtain the initial clusters; and repeat the above steps to obtain a new batch uniformity index, determine that the new batch uniformity index is within the allowable fluctuation range, and obtain the extended growth region mask as the final metabolite enrichment region.
9. A multispectral imaging detection system for determining probiotic activity based on metabolite monitoring, characterized in that, include: Data acquisition module; Analysis module; Recognition module; The acquisition module is used to acquire a cuboid of multispectral image data from a probiotic liquid preparation sample. The cuboid of multispectral image data contains the coordinates of pixels at multiple preset feature wavelengths and their corresponding spectral vectors. The analysis module is used to reduce the dimensionality of the multispectral image data cuboid using principal component analysis to obtain a feature spectral data matrix; to perform clustering using the principal components in the feature spectral data matrix to obtain initial clusters; to obtain target spectral similarity clusters in the initial clusters using the feature spectral vectors in the feature spectral data matrix; and to obtain the intensity values of pixels at characteristic wavelengths for the target spectral similarity clusters and calculate a comprehensive activity score. The identification module is used to preset the activity threshold; The detection results of the probiotic liquid preparation sample are obtained by judging the comprehensive activity score and the activity threshold.