Immune marker analysis system based on machine learning
By processing 3D data of immune markers through machine learning, entropy density gradient maps and boundary topology maps of the markers are generated. Combined with tensor modeling and support vector machine classifiers, the problems of inconsistent interpretation standards and insufficient capture of nonlinear correlation features in traditional systems are solved, thereby improving the accuracy of disease diagnosis models.
Patent Information
- Application Number
- CN202511468519.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Traditional immune marker analysis systems are susceptible to subjective experience, have inconsistent interpretation standards, and are difficult to adapt to the dynamic expression changes of cell population surface markers. Threshold determination models cannot effectively capture nonlinear correlation features, resulting in insufficient accuracy in the construction of disease diagnostic models.
An immune marker analysis system based on machine learning is adopted. The system processes three-dimensional data through an entropy density rheology module, generates an entropy density gradient map of markers by combining the moving average method, detects gradient abrupt change points by a boundary dynamic annotation module, constructs an immune marker boundary topology map, and achieves multi-dimensional data fusion of immune subtype, time dimension and marker expression through a collaborative tensor modeling module. Finally, a support vector machine classifier is used for discrimination.
It enables accurate identification of the morphological boundary features of immune markers, improves the completeness of biomarker association analysis and the accuracy of high-specificity marker identification, and enhances the accuracy of disease diagnostic models.
Smart Images

Figure CN120954685A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data analysis technology, and in particular to an immune marker analysis system based on machine learning. Background Technology
[0002] The field of medical data analysis technology involves a technical system for extracting effective information from clinical test data to support disease diagnosis decisions. Its core lies in the integration and processing of multi-source, heterogeneous medical test data, feature pattern recognition, and visualization output. This field encompasses key technical aspects such as standardized processing of laboratory test data, biomarker correlation analysis, and diagnostic model construction. It requires the comprehensive application of data cleaning algorithms, statistical analysis methods, and visualization technologies to transform medical test data into clinical diagnostic information. Traditional immunomarker analysis systems refer to analytical devices that calculate antibody concentration based on enzyme-linked immunosorbent assay (ELISA) data using optical density values. These systems employ manual interpretation combined with semi-automatic image processing software to achieve quantitative analysis of test results. Alternatively, they utilize flow cytometry data and gate strategies to identify the expression levels of surface markers on specific cell populations, employing threshold determination models to quantitatively analyze fluorescence signal intensity.
[0003] Traditional immunomarker analysis systems rely on manual interpretation and semi-automatic image processing software. In complex and ever-changing clinical testing scenarios, they are easily influenced by subjective experience, leading to inconsistent interpretation standards. The fixed gating strategy used in flow cytometry is difficult to adapt to the dynamic expression changes of cell surface markers. The threshold determination model's linear analysis of fluorescence signal intensity cannot effectively capture non-linear correlation features. The lack of a multi-dimensional data collaborative analysis mechanism in the standardization process of laboratory test data leads to the loss of potential correlation information between different time slices and immunotypes. Semi-quantitative analysis methods are not sensitive enough to low-density data and are prone to boundary identification bias in regions of antibody concentration gradient mutations. Ultimately, this affects the accuracy of disease diagnostic model construction and the effectiveness of clinical decision support. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing an immune marker analysis system based on machine learning.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: an immune marker analysis system based on machine learning, comprising:
[0006] The entropy density rheology module is used to process the three-dimensional data of immune markers through the kernel density estimation algorithm, calculate the spatial neighborhood information entropy, generate the marker entropy density gradient map using the moving average method, and transmit the marker entropy density gradient map to the boundary dynamic annotation module.
[0007] The boundary dynamic annotation module is used to calculate the directional derivative based on the entropy density gradient map of the marker, detect 8-neighbor gradient abrupt change points, and generate an immune marker boundary topology map by using the morphological closing operation of dynamic structuring elements. The immune marker boundary topology map is then passed to the collaborative tensor modeling module.
[0008] The collaborative tensor modeling module is used to prune low-density data based on the immune marker boundary topology map and construct immune subtypes. Time slice The third-order tensor of the marker is used to extract the core factor matrix using the dimension-adaptive Tucker decomposition algorithm. The Manhattan distance is calculated to generate the co-expression strength value. The co-expression strength value and the decomposition residual matrix are then passed to the marker discrimination module.
[0009] The marker discrimination module is used to perform a Hadamard product operation on the co-expression intensity value and the decomposed residual matrix, input the result into a support vector machine classifier for discrimination, and generate a set of highly specific immune markers.
[0010] As a further aspect of the present invention, the entropy density gradient map specifically includes spatial entropy distribution characteristics, gradient vector field, and density contour lines; the boundary topology map includes boundary curvature parameters, topological connected domains, and morphological closed contours; the collaborative expression intensity value covers core factor loadings, temporal evolution patterns, and Manhattan distance matrix; the decomposed residual matrix includes orthogonal residual components, tensor reconstruction error, and low-rank approximation bias; and the high-specificity immune marker set specifically refers to the classification decision surface, feature weight vector, and sample confidence score.
[0011] The radius of the dynamic structural element With the standard deviation of the gradient vector field magnitude satisfy , It is obtained by calculating the standard deviation of all magnitude values of the gradient vector field;
[0012] in, The radius represents the structuring element of the morphological closing operation. This represents the standard deviation of all magnitude values in the gradient vector field. Represents the logarithmic function with base 2. This indicates the rounding up operation;
[0013] The spacing of the density contour lines The moving average is obtained by calculating the Euclidean distance between the centroids of adjacent contour lines. The size of the moving average window is the square root of the number of voxels in the three-dimensional data.
[0014] As a further aspect of the present invention, the directional derivative is calculated using a modified Sobel-Feldman operator, whose convolution kernel size is dynamically adapted to the average spacing of the density contour lines. :
[0015] ;
[0016] in The average Euclidean distance of the density contour lines in the gradient vector field is obtained by calculating the average L2 norm of the difference in centroid coordinates of adjacent contour lines.
[0017] The struct element radius of the morphological closing operation satisfy ;
[0018] in, The radius represents the structuring element of the morphological closing operation. This represents the set of all magnitude values in the gradient vector field. This represents the arithmetic mean of the magnitudes of the gradient vector field. This represents the maximum value in the set of gradient vector field magnitudes. This indicates the rounding up operation.
[0019] As a further aspect of the present invention, the core tensor dimension setting in the Tucker decomposition algorithm follows the following: First dimension ;
[0020] Second dimension ;
[0021] Third dimension ;
[0022] The Manhattan distance matrix is calculated using a sliding window mechanism, with a window width of... With the rank of the core factor matrix satisfy , The optimal selection is made within the preset interval {[}3,8} using cross-validation.
[0023] in, This represents the size of the first dimension of the core tensor. The total number of categories representing immune subtype classification. Represents the logarithmic function with base 2. This indicates the rounding up operation. This represents the size of the second dimension of the core tensor. This represents the number of time slices. This indicates the floor function. This represents the size of the third dimension of the core tensor. This represents the total number of initial markers.
[0024] As a further aspect of the present invention, the kernel function parameters of the support vector machine classifier Dynamically adapt the Frobenius norm of the decomposed residual matrix :
[0025] ;
[0026] The update step size of the feature weight vector With respect to the low-rank approximation deviation Negative correlation, satisfying , For the original third-order tensor, and All values are standardized dimensionless values;
[0027] in, The kernel function parameters represent the kernel parameters of the support vector machine classifier. Represents the total number of initial markers. This represents the variance of all elements in the decomposed residual matrix. The Frobenius norm represents the residual matrix of the decomposition. The function is used to ensure that the denominator is not zero. The update step size represents the feature weight vector. The numerical value representing the low-rank approximation deviation. Frobenius norm represents the original third-order tensor.
[0028] As a further aspect of the present invention, the extraction process of the orthogonal residual components includes:
[0029] Gram-Schmidt orthogonalization is performed on the core factor matrix to generate a set of orthogonal basis vectors;
[0030] Calculate the projection components of the decomposed residual matrix onto each orthogonal basis vector;
[0031] Retaining projection coefficients exceeding a threshold The components, of which The i-th singular value of the core factor matrix is obtained through singular value decomposition.
[0032] threshold The calculation result should be rounded to three decimal places.
[0033] in, The screening threshold representing the projection coefficient. This represents the size of the first dimension of the core tensor. This represents the i-th singular value obtained after singular value decomposition of the core factor matrix. From 1 to Integer index.
[0034] As a further aspect of the present invention, the determination of the topological connected components adopts an improved 8-neighborhood connectivity algorithm, and the merging threshold is set as follows:
[0035] ;
[0036] The area selection criteria for the closed contour are as follows: , , ;
[0037] in, The threshold representing the merging of connected components. This represents the arithmetic mean of all values of the boundary curvature parameter. This represents the standard deviation of all values of the boundary curvature parameter. Represents the total number of initially connected components. This represents the natural exponential function. The area representing the closed contour of the shape. Represents pi (π) The radius of the smallest structuring element used in the morphological closing operation is 1. The maximum structuring element radius used in the morphological closing operation is determined by... Divide by 2 and round up to obtain the result. This represents the average Euclidean distance between the centroids of adjacent density contour lines.
[0038] As a further aspect of the present invention, the extraction of the time evolution mode employs a constrained dynamic time warping algorithm, wherein the slope of the curved path is limited as follows:
[0039] ;
[0040] The calculation of the sample confidence score incorporates the co-expression strength value. and the decomposed residual matrix :
[0041] ;
[0042] in, Represents the index increment in the i-direction of the time series. The index increment in the j-direction of the time series is represented. This represents the number of time slices. Represents the logarithmic function with base 2. A three-dimensional matrix representing the strength values of collaborative expression, with dimension 1. , Represents the number of immune subtypes. Represents the initial number of markers. A three-dimensional matrix representing the decomposed residual matrix with dimensions equal to 1. same, The confidence score represents the sample. The L1 norm of a matrix is the sum of the absolute values of all its elements. The function is used to take the larger of the two values. This represents the Hadamard product operation.
[0043] As a further aspect of the present invention, the calculation of the spatial neighborhood information entropy employs adaptive bandwidth kernel density estimation: bandwidth Local variance of the three-dimensional data of the immune markers satisfy ;
[0044] The window length of the moving average method Dynamically adapt the number of data voxels ,satisfy ;
[0045] in, The bandwidth parameter representing kernel density estimation, This represents the variance of the data within a 5×5×5 cube neighborhood centered at the current voxel. This represents the total sample size of the three-dimensional data of the immune markers. represent -1 / 5, This represents the length of the moving average window. Represents the total number of data voxels.
[0046] As a further aspect of the present invention, the kernel density estimation algorithm employs an improved Epanechnikov kernel function:
[0047] ;
[0048] Streamline tracking step size of gradient vector field Resolution of each axis of the three-dimensional data satisfy , All units were converted to millimeters before being used in the calculation.
[0049] in, The kernel function representing kernel density estimation, Represents the normalized distance parameter and , For the current voxel coordinates, Let the coordinates be those of the neighboring voxels. The step size represents the streamline tracing step. Represents the resolution of the 3D data along the x-axis. Represents the resolution of the 3D data along the y-axis. Represents the resolution of the 3D data along the z-axis. The function is used to find the minimum value among the three resolutions.
[0050] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0051] In this invention, three-dimensional immune data are processed by a kernel density estimation algorithm, and spatial information entropy is calculated. This effectively integrates the spatial distribution characteristics of multi-source heterogeneous medical data. The gradient map generated by the moving average method can dynamically capture the continuous trend of biomarker concentration changes. Based on the boundary topology model constructed by directional derivative detection and morphological closing operation, the morphological boundary features of immune biomarkers are accurately identified. Third-order tensor modeling is used to achieve multi-dimensional data fusion of immune subtype, time dimension, and biomarker expression. The core factor matrix is extracted by the dimension-adaptive tensor decomposition algorithm, which significantly improves the completeness of biomarker correlation analysis. The co-expression intensity value quantified by Manhattan distance is combined with the Hadamard product operation of the residual matrix to enhance the recognition accuracy of the support vector machine classifier for highly specific biomarkers. Finally, the technology of quantitative analysis of immune biomarkers is realized from single threshold judgment to multi-dimensional co-discrimination. Attached Figure Description
[0052] Figure 1 This is a flowchart of the machine learning-based immune marker analysis system of the present invention.
[0053] Figure 2 This is a flowchart of the entropy density rheology module of the present invention;
[0054] Figure 3 This is a flowchart of the boundary dynamic annotation module of the present invention;
[0055] Figure 4 This is a flowchart of the collaborative tensor modeling module of the present invention;
[0056] Figure 5 This is a flowchart of the mark discrimination module of the present invention. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of this invention clearer, the software-based technical solution is described in detail below with reference to system architecture diagrams and embodiments. It should be understood that the specific embodiments described herein are only for explaining the technical solutions of this invention and do not constitute a limitation on the scope of protection.
[0058] In the description of this invention, the system architecture relationships or data processing flows indicated by terms such as "layer," "module," "interface," "data flow," "client," and "server" are all defined based on the architecture diagram or flowchart corresponding to the embodiments. This way of describing is only used to clearly illustrate the logical relationships between the elements in the technical solution, and not to limit the physical deployment form. The term "multiple" includes two or more technical units, including but not limited to multiple data nodes, processing threads, service instances, or functional components and other scalable elements. The specific number is determined according to the actual business scenario and needs to be specifically specified.
[0059] Please see Figure 1 and Figure 2 This invention provides a technical solution: an immune marker analysis system based on machine learning, comprising:
[0060] The entropy density rheology module is used to process the three-dimensional data of immune markers through the kernel density estimation algorithm, calculate the spatial neighborhood information entropy, generate the marker entropy density gradient map using the moving average method, and pass the marker entropy density gradient map to the boundary dynamic annotation module.
[0061] The entropy density gradient map specifically includes spatial entropy distribution characteristics, gradient vector field, and density contour lines;
[0062] The spatial neighborhood information entropy is calculated using adaptive bandwidth kernel density estimation: bandwidth Local variance of three-dimensional data of immune markers satisfy ;
[0063] in, The bandwidth parameter representing kernel density estimation, This represents the variance of the data within a 5×5×5 cube neighborhood centered at the current voxel. The total sample size representing the three-dimensional data of immune markers. represent -1 / 5;
[0064] The kernel density estimation algorithm employs a modified Epanechnikov kernel function: ;
[0065] in, The kernel function representing kernel density estimation, Represents the normalized distance parameter and , For the current voxel coordinates, The coordinates are those of the neighboring voxels.
[0066] Window length of moving average method Dynamically adapt the number of data voxels ,satisfy ;
[0067] in, This represents the length of the moving average window. Represents the total number of data voxels;
[0068] Spacing of density contour lines It is obtained by calculating the moving average of the Euclidean distance between the centroids of adjacent contour lines. The size of the moving average window is the integer value of the square root of the number of voxels in the three-dimensional data.
[0069] Streamline tracking step size of gradient vector field Resolution of each axis of the three-dimensional data satisfy , All units were converted to millimeters before being used in the calculation.
[0070] in, The step size represents the streamline tracing step. Represents the resolution of the 3D data along the x-axis. Represents the resolution of the 3D data along the y-axis. Represents the resolution of the 3D data along the z-axis. The function is used to find the minimum value among the three resolutions.
[0071] First, the three-dimensional data of the immunomarkers to be analyzed were acquired. This data was obtained from biopsy tissue samples from lymphoma patients, using multicolor immunofluorescence (mIF) imaging technology. The imaging device was a Leica STELLARIS 8 DIVE with an objective magnification of 40x. This three-dimensional data is stored digitally as a series of consecutive stacked two-dimensional image slices, with each pixel recording the fluorescence intensity value of a specific immunomarker (e.g., CD8, PD-1, FoxP3). In this embodiment, a tissue region measuring 500 μm × 500 μm × 100 μm was extracted for analysis, and this region was digitized into a 1000 × 1000 × 200 voxel three-dimensional data array. Therefore, the total sample size of this immunomarker three-dimensional data is... 2×10 8 10 data points. The resolution of each axis of the 3D data, after being converted to millimeters, is as follows: =0.0005mm, =0.0005mm, =0.0005mm.
[0072] Next, the 3D data is processed using a kernel density estimation algorithm to calculate the spatial neighborhood information entropy. This process is performed at each voxel location. Taking the voxel with coordinates (150, 250, 80) as an example, its spatial neighborhood is first determined. Here, the neighborhood is a 5×5×5 cube region centered on the voxel, containing a total of 125 voxels (5×5×5=125). Then, the local variance of the 125 data points (i.e., fluorescence intensity values) within this neighborhood is calculated. In this example, the local variance was obtained by statistically calculating these 125 fluorescence intensity values. The value was 215.3. Subsequently, based on this local variance and the total sample size... Calculate adaptive bandwidth .
[0073] formula The adaptive bandwidth used to calculate kernel density estimation. This represents the bandwidth parameter, whose value is dynamically adjusted based on the local data density, decreasing in densely populated areas and increasing in sparsely populated areas. (Parameter) The proportionality coefficient is an empirically selected value determined through bandwidth optimization experiments on over 500 immunofluorescence image datasets from different tissue types. The experimental procedure was as follows: for each dataset, a series of bandwidth values (from 0.5 to 2.0, with a step size of 0.1) were preset. Kernel density was estimated using each bandwidth value, and the KL divergence between the resulting density distribution and the actual cell distribution (manually labeled by pathologists) was calculated. The proportionality coefficient that minimizes the average KL divergence across all samples was selected. Experimental results show that 1.06 achieves optimal or near-optimal estimation results in various scenarios. Represents local variance, reflecting the degree of dispersion of data within a neighborhood. For the total sample size, the index term It is the optimal bandwidth scaling factor derived based on the criterion of minimizing mean square integral error (MISE). The innovation of this formula lies in reflecting the overall volume of the data. Compared with reflecting local features This combination enables adaptive bandwidth adjustment, avoiding the over-smoothing or under-smoothing issues that arise when processing data with uneven density using a globally uniform bandwidth. Substituting the above values into the formula for calculation: ,but . Therefore, the bandwidth used for kernel density estimation of voxels (150, 250, 80) is... It is 4.98.
[0074] Obtain bandwidth Then, the improved Epanechnikov kernel function was used to estimate the information entropy of the voxel. A voxel with coordinates (152, 251, 81) in its neighborhood was used as an example. For example, first calculate its relationship with the central voxel. Normalized distance parameter . .because Therefore, the first part of the kernel function is used for calculation. Formula This is an improved Epanechnikov kernel function. Wherein, The kernel function value determines the contribution weight of neighboring voxels to the central voxel density estimation. It is the normalized distance parameter, derived from the central voxel. With neighboring voxels Euclidean distance Divide by bandwidth The innovation of this kernel function lies in the fact that it is a piecewise function, which, compared to the standard Epanechnikov kernel function, expands the range of influence from... Expanded to This creates a "soft boundary". When When, it maintains the optimality of the Epanechnikov kernel in the sense of mean square integral error; when At that time, the weights smoothly decrease to zero, instead of... The abrupt change to zero occurs at a certain point. This design, while preserving computational efficiency, can handle voxels at neighborhood boundaries more smoothly, reducing boundary effects in the estimation results and improving the robustness of density estimation. Substitute into the formula: This calculation is repeated for all 124 neighboring voxels within a 5×5×5 neighborhood, and all kernel function values are summed with weights to obtain the entropy density value at the central voxel (150, 250, 80). After traversing all voxels, a complete spatial entropy distribution feature map is generated.
[0075] Subsequently, the spatial entropy distribution feature map is smoothed using the moving average method to generate an entropy density gradient map of the markers. First, the window length of the moving average is determined. Total number of data voxels 1000×1000×200=2×10 8 . formula Used to dynamically adapt the length of the moving average window. Among them, For window length, For floor operations, The total number of data voxels. This formula determines the window size using the square root of the data voxel count, allowing the smoothness to adapt to different data sizes. A larger window is used for large datasets to effectively suppress noise, while a smaller window is used for small datasets to retain more detail. Adding 1 ensures the window length is odd, giving the window a unique center point. Substitute into the formula: This window length applies to one-dimensional entropy density data. In three-dimensional space, a [length] is typically used. The cubic window, in which Here, we assume a cubic window with a side length of 585 voxels. For each point in the entropy density map, we calculate the average of all entropy density values within its 585×585×585 neighborhood, using this as the smoothed value for that point. On the smoothed entropy density map, we calculate the gradient, forming a gradient vector field. Then, we draw density contour lines based on the gradient vector field. The spacing between the density contour lines... The distance between the centroids of adjacent contour lines is obtained by calculating the moving average of the distance between them. The size of the moving average window is set to the square root of the number of voxels in the 3D data, which is 14142. For example, if the centroid coordinates of three adjacent contour lines are calculated as C1(10,20,30), C2(12,23,31), and C3(15,25,33), then the distance between them is... , Assuming these two distance values belong to a portion of the moving average window, the final average distance is obtained by averaging all 14,142 such distance values within the window. The calculation obtained in this embodiment is as follows: Finally, the streamline tracking step size of the gradient vector field is calculated. . formula Used to calculate the streamline tracking step size. Step size, It is the resolution (in millimeters) of each axis of the three-dimensional data. The function takes the minimum value among the three. The coefficient 0.5 is to comply with the Nyquist sampling theorem, ensuring that the tracking step size is less than half of the minimum resolution, thereby avoiding missing key gradient change features during streamline tracing. Substituting the resolution value into the formula: mm. This step size This data is used for subsequent streamline visualization and analysis. Finally, the marker entropy density gradient map, which includes spatial entropy distribution characteristics, gradient vector field, and density contour lines, is passed to the boundary dynamic annotation module.
[0076] Please see Figure 1 and Figure 3 The boundary dynamic annotation module is used to calculate the directional derivative based on the entropy density gradient map of the marker, detect the gradient abrupt change point in the 8-neighborhood, generate the immune marker boundary topology map by using the morphological closing operation of the dynamic structuring element, and pass the immune marker boundary topology map to the co-tensor modeling module.
[0077] The boundary topology graph includes boundary curvature parameters, topological connected components, and morphological closed contours.
[0078] The directional derivative is calculated using a modified Sobel-Feldman operator, whose kernel size dynamically adapts to the average spacing of density contour lines. : ;
[0079] in The average Euclidean distance between density contour lines in the gradient vector field is obtained by calculating the average L2 norm of the difference in centroid coordinates between adjacent contour lines.
[0080] radius of dynamic structuring element With the standard deviation of the gradient vector field magnitude satisfy , It is obtained by calculating the standard deviation of all magnitude values of the gradient vector field;
[0081] in, The radius represents the structuring element of the morphological closing operation. This represents the standard deviation of all magnitude values in the gradient vector field. Represents the logarithmic function with base 2. This indicates the rounding up operation;
[0082] The radius of the structuring element in morphological closing operations satisfy ;
[0083] in, The radius represents the structuring element of the morphological closing operation. Represents the set of all magnitude values in the gradient vector field. The arithmetic mean of the magnitudes of the gradient vector field is represented by... This represents the maximum value in the set of gradient vector field magnitudes. This indicates the rounding up operation;
[0084] The determination of topological connectivity uses an improved 8-neighborhood connectivity algorithm, with the merging threshold set as follows: ;
[0085] in, The threshold representing the merging of connected components. This represents the arithmetic mean of all values of the boundary curvature parameter. This represents the standard deviation of all values of the boundary curvature parameter. Represents the total number of initially connected components. Represents the natural exponential function;
[0086] The area selection criteria for closed contours are as follows: , , ;
[0087] in, The area representing the closed contour of the shape. Represents pi (π) The radius of the smallest structuring element used in the morphological closing operation is 1. The maximum structuring element radius used in the morphological closing operation is determined by... Divide by 2 and round up to obtain the result. This represents the average Euclidean distance between the centroids of adjacent density contour lines.
[0088] This module first calculates the directional derivative based on the biomarker entropy density gradient map to identify the boundary between immune cell aggregation regions and the tissue matrix. This process employs a modified Sobel-Feldman operator. The kernel size of this operator is based on the average spacing of the density contour lines calculated in the previous module. Dynamic adaptation is performed. In the aforementioned embodiments, calculations have been performed. .
[0089] formula Used to generate convolution kernels in the x-direction. Wherein, The convolution kernel matrix, For floor operations, Let be the average Euclidean distance between the density contour lines. The innovation of this formula lies in incorporating the weights (1 and 2) from the standard Sobel-Feldman operator with the contour line spacing. Connect them. When A larger value indicates a gradual change in entropy density and blurred boundaries. In this case, increasing the weights of the convolution kernel can enhance the response to subtle edges; conversely, when... A smaller weight indicates clearer boundaries, thus a smaller weight is used to avoid over-amplifying noise. This adaptive mechanism ensures that the sensitivity and accuracy of edge detection remain stable across image regions of varying sharpness. Substitute into the formula to calculate the weighting factor: Therefore, the convolution kernel in the x-direction convolution kernel in the y direction They are respectively: , These two convolutional kernels are applied to each 3×3 pixel neighborhood of the entropy density gradient map to calculate the gradient components Gx and Gy in the x and y directions, thus obtaining the gradient magnitude at that point. and direction By traversing the entire image, points where gradient values change abruptly within an 8-neighborhood are detected, and these points are initially marked as boundary points.
[0090] Next, morphological closing operations on dynamic structuring elements are used to process the initially labeled boundary points, connecting broken boundaries and filling small holes to generate a complete immune marker boundary topology map. The radius of the structuring element... Based on the standard deviation of the gradient vector field magnitude Dynamic calculation. First, calculate all magnitude values in the entire gradient vector field. Standard deviation In this embodiment, the standard deviation is obtained by statistically analyzing the gradient magnitudes of all points in the gradient graph. .
[0091] formula Used to calculate the radius of the structuring element in morphological closing operations. Among them, For radius, For floor operations, It is a logarithm with base 2. It is the standard deviation of the magnitude of the gradient vector field. The innovation of this formula lies in relating the size of the structuring element to the overall discreteness of the gradient change (i.e., (This is related to) When the value is large, it indicates that the intensity and sharpness of the boundaries in the image vary greatly, and there may be wide or blurry boundary regions. In this case, a larger structuring element is needed to bridge these uncertain regions; when When the scale is small, the boundaries are more consistent and clear, allowing for the use of smaller structuring elements and avoiding over-smoothing that could lead to loss of detail. This adaptability to logarithmic scales makes the response of structuring element size changes more smooth and robust to variations in gradient discretization. Substitute into the formula: Therefore, a circular (or spherical, in 3D) structuring element with a radius of 3 is used to perform a morphological closing operation. The closing operation involves first dilation and then erosion, effectively connecting adjacent boundary points and forming a preliminary morphological closed contour.
[0092] Subsequently, topological connectivity is determined for the generated multiple closed contours to merge regions belonging to the same immune cell cluster. This process employs an improved 8-neighborhood connectivity algorithm with a merging threshold. It is dynamically set. First, the boundary curvature parameters of all initial connected components are calculated to obtain a set of curvature values. In this embodiment, a total of [number missing] were identified. Given an initial connected region, calculate its mean curvature. Standard deviation of curvature .
[0093] formula The threshold used to calculate connected component merging. The merging threshold, and Let be the arithmetic mean and standard deviation of the boundary curvature of all initial connected domains, respectively. This represents the total number of initially connected components. It is a natural exponential function. The innovation of this formula lies in its adaptability: the threshold is based on the mean curvature. And based on the standard deviation of curvature and the number of connected components The relevant sigmoid function terms are adjusted. This is based on the initial number of connected components. When the number of terms is small, the exponent in the denominator is close to 1, the adjustment term is large, and there is a tendency to combine terms; when the number of terms is small... When Nc is very large, the exponent approaches 0, the denominator approaches 1, and the adjustment range is also large. This seems counterintuitive, but it can be explained by the fact that when regions are over-segmented (Nc is large), a more lenient threshold is needed to merge them. The S-shaped curve transition in the middle ensures the smoothness of the threshold change. This design allows the merging decision to consider not only the average curvature of the region boundaries, but also the dispersion of the curvature and the degree of fragmentation, thus achieving smarter region merging. Substitute into the formula: The merging threshold is set to 0.33. If the average curvature of the boundary between two adjacent connected components is less than 0.33, they are merged into one connected component.
[0094] Finally, the merged morphological closed contours are filtered by area to remove excessively small regions that may be noise or artifacts. The filtering criteria are related to the morphological operation parameters and contour spacing. Minimum structuring element radius. Set to 1, maximum structuring element radius according to Calculation. Formula Used to determine the radius of the maximum structuring element. Substitute: The area screening threshold is determined by the formula. Determined. Where A is the area of the outline. This is pi (π). The innovation of this formula lies in linking the lower limit of the area selection to the minimum and maximum feature sizes that morphological operations can handle. Logically, any contour smaller than the "average feature area" defined by the minimum and maximum structuring elements is likely unreliable and should be discarded. Area threshold calculation: All closed contours with an area smaller than 3.927 voxels will be removed. After all the above steps, the final generated immune marker boundary topology map contains accurate boundary curvature parameters, topological connectivity, and filtered morphological closed contours, and is passed to the co-tensor modeling module.
[0095] Please see Figure 1 and Figure 4The collaborative tensor modeling module is used to trim low-density data based on the immune marker boundary topology map, construct a third-order tensor of immune subtype × time slice × marker, extract the core factor matrix using the dimension-adaptive Tucker decomposition algorithm, calculate the Manhattan distance to generate collaborative expression intensity values, and pass the collaborative expression intensity values and decomposition residual matrix to the marker discrimination module.
[0096] The co-expression strength values encompass core factor loadings, temporal evolution patterns, and the Manhattan distance matrix;
[0097] The decomposed residual matrix includes orthogonal residual components, tensor reconstruction error, and low-rank approximation bias.
[0098] The core tensor dimension setting in the Tucker decomposition algorithm follows the principle of: First dimension ;
[0099] Second dimension ;
[0100] Third dimension ;
[0101] in, This represents the size of the first dimension of the core tensor. The total number of categories representing immune subtype classification. Represents the logarithmic function with base 2. This indicates the rounding up operation. This represents the size of the second dimension of the core tensor. This represents the number of time slices. This indicates the floor function. This represents the size of the third dimension of the core tensor. Represents the total number of initial markers;
[0102] The Manhattan distance matrix is calculated using a sliding window mechanism, with a window width of... With the rank of the core factor matrix satisfy , The optimal selection is made within the preset interval [3,8] using cross-validation.
[0103] The extraction process of orthogonal residual components includes: performing Gram-Schmidt orthogonalization on the core factor matrix to generate a set of standard orthogonal basis vectors;
[0104] Calculate the projection components of the decomposed residual matrix onto each orthogonal basis vector;
[0105] Retaining projection coefficients exceeding a threshold The components, of which The i-th singular value of the core factor matrix is obtained through singular value decomposition.
[0106] threshold The calculation result should be rounded to three decimal places.
[0107] in, The screening threshold representing the projection coefficient. This represents the size of the first dimension of the core tensor. The i-th singular value obtained after singular value decomposition of the core factor matrix represents the i-th singular value. From 1 to Integer index;
[0108] The extraction of the time evolution pattern employs a constrained dynamic time warping algorithm, with the slope of the curved path limited as follows: ;
[0109] in, Represents the index increment in the i-direction of the time series. The index increment in the j-direction of the time series is represented. This represents the number of time slices. This represents the logarithmic function with base 2.
[0110] First, the module crops the original 3D immunofluorescence data based on the immunomarker boundary topology map, setting the low-density region data outside the boundary contour (e.g., fluorescence intensity below a preset value, obtained by subtracting twice the standard deviation from the mean of all pixel intensities within the boundary) to zero. This operation focuses on the core region of immune cell aggregation. Next, a third-order tensor is constructed to integrate multi-dimensional information. The data organization of this tensor is: immunotype × time slice × marker. In this embodiment, the research object is three immunotypes (… (e.g., Treg, Th1, Th2 cells), 49 time points were collected ( For example, tissue samples were collected every 2 hours for 4 days after drug intervention, and 20 different immune markers were tested on each sample. This constructs a 3×49×20 third-order tensor. Each element in the tensor Representative at the At the 1st time point, the 1st Of the 10 immune subtypes, the first The average fluorescence intensity values of each marker. To illustrate the data structure, a portion of the data (data from time point 1) is listed in tabular form below:
[0111] Table 1. Mean fluorescence intensity of immunomarkers at time point 1.
[0112] Immune subtypes Marker 1 (CD3) Marker 2 (CD4) ... Marker 20 (IL-10) Subtype 1 (Treg) 158.2 120.5 ... 210.7 Subtype 2 (Th1) 165.3 115.1 ... 45.2 Subtype 3 (Th2) 149.8 125.7 ... 98.6
[0113] As shown in Table 1, the table clearly displays the expression levels of various biomarkers in different immune subtypes at the initial time point. The entire third-order tensor contains such data for all 49 time points.
[0114] Subsequently, the core factor matrix is extracted using the dimension-adaptive Tucker decomposition algorithm. The core tensor dimension of this algorithm is dynamically set according to the scale of the input data. The formula for setting the core tensor dimension is: First dimension... The second dimension The third dimension The innovation of these formulas lies in the fact that they do not employ a fixed or empirically chosen rank, but rather combine the dimension of the core tensor (i.e., the complexity of the model) with the size of each mode in the original data. To perform mathematical associations. Logarithmic relationships are used because the number of immune subtypes is usually small, and logarithmic growth can effectively compress information without excessively losing class specificity. Using the square root relationship is suitable for capturing the main trends in time series rather than high-frequency noise. The use of a power of 0.4 is an empirical index that balances model fit and generalization ability, derived from empirical analysis of a large amount of proteomics and genomics data. This index was obtained by searching for the optimal average value in the range [0.2, 0.6] through cross-validation experiments on more than 100 public datasets. Substituting the data scale of this example into the calculation: . . Therefore, the core tensor dimension is set to 2×7×4. After performing Tucker decomposition, a 2×7×4 core tensor and three factor matrices are obtained, corresponding to the immune subtype (3×2), time evolution (49×7), and biomarker (20×4) patterns, respectively.
[0115] Next, the Manhattan distance is calculated to generate collaborative expression strength values. This calculation employs a sliding window mechanism on the biomarker factor matrix. The choice of window width is related to the rank of the core factor matrix. Related. The optimal rank for the factor matrix is selected within the preset interval [3,8] using cross-validation. In this example, cross-validation (e.g., dividing the factor matrix data into 10 folds, training the model with 9 folds each time to predict the remaining 1 fold, and selecting the rank that minimizes the prediction error) indicates the optimal rank. The width of the sliding window is determined by the formula. Confirmed. This formula originates from optimal filter length design in signal processing; its exponential relationship ensures that the window width can cover the range of rank... The main interaction modes at the complexity level represented. Substitute into the formula: Due to the number of markers If the window width is less than 63, the sliding window here will cover all markers, that is, calculate the Manhattan distance of each pair of markers in the 4-dimensional factor space, and generate a 20×20 Manhattan distance matrix. The element values of this matrix reflect the cooperative expression strength between markers.
[0116] During the decomposition process, orthogonal residual components also need to be extracted. This process first performs Gram-Schmidt orthogonalization on the core factor matrix (taking the subtype factor matrix as an example, with a size of 3×2) to generate an orthonormal basis. Then, the projection components of the decomposition residual matrix onto these orthonormal bases are calculated. Projection coefficients exceeding a threshold are retained. The component. This threshold is given by the formula. Confirmed. The innovation of this formula lies in correlating the threshold with the singular values of the factor matrix. These singular values represent the energy or importance of the data in that direction. The Frobenius norm of the factor matrix reflects the total amount of information captured by the factor matrix. The threshold is proportional to this information content, allowing the screening criteria to adapt to the importance of different factor matrices. The coefficient 0.1 was determined through ROC analysis on simulated data, aiming to maximize the removal of injected random noise while retaining 95% of the true signal. First, singular value decomposition (SVD) is performed on the 3×2 subtype factor matrix to obtain two singular values. . . The calculation results are rounded to three decimal places, with a threshold of 1.560. Any residual component with an absolute value of projection coefficient less than 1.560 will be discarded.
[0117] Finally, the constrained dynamic time warping (DTW) algorithm is used to extract the time evolution pattern. The slope of the curved path is constrained to prevent unreasonable matching. The slope constraint is given by the formula... Given. Among them... To match the local slope of the path, The number of time slices. The innovation of this constraint is that it is not a fixed value, but rather varies with the length of the time series. Logarithmic growth. For longer time series, it allows for greater flexibility and non-linearity in matching paths, while for shorter series, it imposes stronger linear constraints, consistent with the intuition that long-term trends in time series analysis can be more complex and variable. Substitute into the formula: In the DTW calculation, any match that causes the local slope of the matching path to exceed 1.5615 will be banned. Finally, the co-expression strength value and the decomposed residual matrix are passed together to the flag discrimination module.
[0118] Please see Figure 1 and Figure 5 The marker discrimination module is used to perform Hadamard product operation on the co-expression intensity value and the decomposition residual matrix, input it into the support vector machine classifier for discrimination, and generate a set of highly specific immune markers;
[0119] High-specificity immune marker set specifically refers to classification decision surface, feature weight vector, and sample confidence score;
[0120] Kernel function parameters of support vector machine classifier Frobenius norm of dynamically adapting decomposition residual matrix : ;
[0121] in, The kernel function parameters represent the kernel parameters of the support vector machine classifier. Represents the total number of initial markers. This represents the variance of all elements in the decomposed residual matrix. The Frobenius norm represents the residual matrix of the decomposition. The function is used to ensure that the denominator is not zero;
[0122] Update step size of feature weight vector Deviation from low-rank approximation Negative correlation, satisfying , For the original third-order tensor, and All values are standardized dimensionless values;
[0123] in, The update step size represents the feature weight vector. The numerical value representing the low-rank approximation deviation. The Frobenius norm representing the original third-order tensor;
[0124] The calculation of sample confidence scores incorporates co-expression strength values. and decomposition residual matrix : ;
[0125] in, A three-dimensional matrix representing the strength values of collaborative expression, with dimension 1. , Represents the number of immune subtypes. Represents the initial number of markers. A three-dimensional matrix representing the decomposed residual matrix with dimensions equal to 1. same, The confidence score represents the sample. The L1 norm of a matrix is the sum of the absolute values of all its elements. The function is used to take the larger of the two values. This represents the Hadamard product operation.
[0126] First, the module will collaboratively express the intensity value matrix. (A reflection of the core correlation between markers) That is, a 3×49×20 third-order matrix and the decomposition residual matrix (A 3x3 matrix of the same dimension containing secondary information ignored by the core model) performs the Hadamard product operation. This operation is an element-wise multiplication, denoted as . The purpose of this operation is to fuse the main pattern and residual information. Features that have large values in both the core pattern and the residual (which may represent nonlinear or unique interactions) will be amplified, thus providing richer features for subsequent classifiers.
[0127] Next, the result of the Hadama product... As input, the data is fed into a Support Vector Machine (SVM) classifier for discrimination, aiming to screen for highly specific immune markers from 20 initial markers. The classifier's goal is to differentiate, for example, effectively predicting the immune subtype status of drug response. The performance of the SVM classifier largely depends on the kernel function and its parameters. This embodiment uses a Gaussian kernel (RBF kernel), whose key parameters... Dynamically adapt based on the characteristics of the decomposed residual matrix.
[0128] Kernel function parameters are given by the formula Dynamic adaptation. The innovation of this formula lies in its scaling of the kernel function width (by...). The reciprocal of the denominator (determined by the input data) is directly related to the statistical properties of the input data. It is a scaling factor commonly used in Hebb's learning rules and random matrix theory to standardize the data scale based on the number of features and variance. (Molecular part) It is a robust normalization term, when the Frobenius norm of the residual matrix... When the value is very small (close to zero, which could lead to numerical instability), this term is approximately 1, thus avoiding [the problem]. The value grows explosively; otherwise, the value is 1. This design makes... It can adaptively adjust based on the variance, norm, and feature dimensions of the data itself, avoiding time-consuming manual parameter tuning or cross-validation processes. (Constant) This is a small positive number set to prevent division by zero. Its value was selected through experiments with thousands of values to ensure sufficient safety in double-precision floating-point operations and to not affect the vast majority of calculation results. In this embodiment, By calculation, the residual matrix is decomposed. Variance of all elements Its Frobenius norm Substitute these values into the formula: .this The value is used to configure the SVM classifier.
[0129] During the training of SVM, if optimization methods such as gradient descent are used, the update step size of the feature weight vector... It is also dynamically adjusted. The update step size is determined by the formula. Confirmed. Among them, To update the step size, This represents the low-rank approximation bias (i.e., the reconstruction error of the Tucker decomposition). Let Frobenius be the Frobenius norm of the original third-order tensor. The one used here... and All are standardized dimensionless values, for example, by dividing the original value by... The innovation of this formula lies in relating the learning rate to the model's goodness of fit (derived from the relative reconstruction error). Reflects the correlation. When the model fits well ( A small step size means the current weights are close to optimal, and fine-tuning should be done using a smaller step size; when the model fits poorly ( If the initial learning rate is large, a larger step size can be used to accelerate convergence. The coefficient 0.01 is an empirically set initial learning rate benchmark. Convergence tests on multiple standard machine learning datasets showed that a value between 0.005 and 0.02 performed well, and 0.01 is a robust choice. In this embodiment, the low-rank approximation bias is calculated. The Frobenius norm of the original tensor . This update step size It is used for weight adjustment during classifier training.
[0130] After training, the SVM model generates a classification decision surface and feature weight vectors for each marker. Furthermore, to evaluate the reliability of the classification result for each sample, the system also calculates a sample confidence score. The sample confidence score is calculated using the formula... Calculation. Among them, For the collaborative expression intensity value matrix, To decompose the residual matrix, For Hadama accumulation, It is the L1 norm (the sum of the absolute values of all elements of the matrix). The function takes the larger of the two values. The innovation of this formula lies in its use of a normalized metric to quantify the consistency or resonance strength between "core information" and "residual information." (Molecule) This amplifies signals that are in the same direction (both positive or both negative) and have large absolute values. The denominator uses... Normalization, compared to using the sum or product of the two, is more efficient for... and It is more robust to cases with large norm differences. A high The values indicate that the information contained in the residuals is highly correlated with the core patterns extracted by tensor decomposition, which enhances the confidence in classification decisions. Taking a specific sample as an example, the L1 norm of its co-expression strength matrix was calculated. L1 norm of the decomposed residual matrix The L1 norm of the Hadamard product of the two . The result of 0.536 indicates that the classification confidence of this sample is moderately high. A threshold (e.g., 0.75) can be set for the confidence score to filter out classification results with high confidence. Finally, based on the feature weight vector output by the SVM, the 20 biomarkers are ranked, and the top 5 biomarkers with the highest absolute weight values are identified as a set of high-specificity immune biomarkers, such as {PD-L1, GZMB, Ki-67, CD8, CTLA-4}. This set is the final output of this system.
[0131] The above embodiments illustrate preferred embodiments of the present invention. Any equivalent adjustments to the technical solution based on software engineering methods are within the scope of protection, including but not limited to: implementing algorithm logic using different programming languages, refactoring functional modules into services, adjusting data interaction protocols, and optimizing resource scheduling strategies. Any implementation scheme derived from reasonable modifications to the data processing flow, service call chain, or system architecture layer without departing from the core technology of the present invention should be considered within the scope of protection defined by the claims of the present invention.
Claims
1. An immune marker analysis system based on machine learning, characterized in that, The system includes: The entropy density rheology module is used to process the three-dimensional data of immune markers through the kernel density estimation algorithm, calculate the spatial neighborhood information entropy, generate the marker entropy density gradient map using the moving average method, and transmit the marker entropy density gradient map to the boundary dynamic annotation module. The boundary dynamic annotation module is used to calculate the directional derivative based on the entropy density gradient map of the marker, detect 8-neighbor gradient abrupt change points, and generate an immune marker boundary topology map by using the morphological closing operation of dynamic structuring elements. The immune marker boundary topology map is then passed to the collaborative tensor modeling module. The collaborative tensor modeling module is used to trim low-density data according to the immune marker boundary topology map, construct a third-order tensor of immune subtype × time slice × marker, extract the core factor matrix using the dimension-adaptive Tucker decomposition algorithm, calculate the Manhattan distance to generate collaborative expression intensity values, and pass the collaborative expression intensity values and decomposition residual matrix to the marker discrimination module. The marker discrimination module is used to perform a Hadamard product operation on the co-expression intensity value and the decomposed residual matrix, input the result into a support vector machine classifier for discrimination, and generate a set of highly specific immune markers.
2. The machine learning-based immune marker analysis system according to claim 1, characterized in that, The entropy density gradient map specifically includes spatial entropy distribution features, gradient vector field, and density contour lines; the boundary topology map includes boundary curvature parameters, topological connected domains, and morphological closed contours; the collaborative expression intensity value covers core factor loadings, temporal evolution patterns, and Manhattan distance matrix; the decomposed residual matrix includes orthogonal residual components, tensor reconstruction error, and low-rank approximation bias; and the high-specificity immune marker set specifically refers to the classification decision surface, feature weight vector, and sample confidence score. The radius r of the dynamic structuring element and the standard deviation of the gradient vector field magnitude satisfy , It is obtained by calculating the standard deviation of all magnitude values of the gradient vector field; in, The radius represents the structuring element of the morphological closing operation. This represents the standard deviation of all magnitude values in the gradient vector field. Represents the logarithmic function with base 2. This indicates the rounding up operation; The spacing Δd between the density contour lines is obtained by calculating the moving average of the Euclidean distance between the centroids of adjacent contour lines. The size of the moving average window is the square root of the number of voxels in the three-dimensional data.
3. The machine learning-based immune marker analysis system according to claim 2, characterized in that, The directional derivative is calculated using a modified Sobel-Feldman operator, whose kernel size is dynamically adapted to the average spacing of the density contour lines. : ; in The average Euclidean distance of the density contour lines in the gradient vector field is obtained by calculating the average L2 norm of the difference in centroid coordinates of adjacent contour lines. The struct element radius of the morphological closing operation satisfy ; in, The radius represents the structuring element of the morphological closing operation. This represents the set of all magnitude values in the gradient vector field. This represents the arithmetic mean of the magnitudes of the gradient vector field. This represents the maximum value in the set of gradient vector field magnitudes. This indicates the rounding up operation.
4. The machine learning-based immune marker analysis system according to claim 3, characterized in that, The core tensor dimension setting in the Tucker decomposition algorithm follows the following principle: First dimension ; Second dimension ; Third dimension ; The Manhattan distance matrix is calculated using a sliding window mechanism, with a window width of... With the rank of the core factor matrix satisfy , The optimal selection is made within the preset interval [3,8] using cross-validation. in, This represents the size of the first dimension of the core tensor. The total number of categories representing immune subtype classification. Represents the logarithmic function with base 2. This indicates the rounding up operation. This represents the size of the second dimension of the core tensor. This represents the number of time slices. This indicates the floor function. This represents the size of the third dimension of the core tensor. This represents the total number of initial markers.
5. The machine learning-based immune marker analysis system according to claim 4, characterized in that, The kernel function parameters of the support vector machine classifier Dynamically adapt the Frobenius norm of the decomposed residual matrix : ; The update step size of the feature weight vector With respect to the low-rank approximation deviation Negative correlation, satisfying , For the original third-order tensor, and All values are standardized dimensionless values; in, The kernel function parameters represent the kernel parameters of the support vector machine classifier. Represents the total number of initial markers. This represents the variance of all elements in the decomposed residual matrix. The Frobenius norm represents the residual matrix of the decomposition. The function is used to ensure that the denominator is not zero. The update step size represents the feature weight vector. The numerical value representing the low-rank approximation deviation. Frobenius norm represents the original third-order tensor.
6. The machine learning-based immune marker analysis system according to claim 5, characterized in that, The extraction process of the orthogonal residual components includes: Gram-Schmidt orthogonalization is performed on the core factor matrix to generate a set of orthogonal basis vectors; Calculate the projection components of the decomposed residual matrix onto each orthogonal basis vector; Retaining projection coefficients exceeding a threshold The components, of which The i-th singular value of the core factor matrix is obtained through singular value decomposition. threshold The calculation result should be rounded to three decimal places. in, The screening threshold representing the projection coefficient. This represents the size of the first dimension of the core tensor. This represents the i-th singular value obtained after singular value decomposition of the core factor matrix. From 1 to Integer index.
7. The machine learning-based immune marker analysis system according to claim 6, characterized in that, The determination of the topological connectivity domain adopts an improved 8-neighborhood connectivity algorithm, and the merging threshold is set as follows: ; The area selection criteria for the closed contour are as follows: , , ; in, The threshold representing the merging of connected components. This represents the arithmetic mean of all values of the boundary curvature parameter. This represents the standard deviation of all values of the boundary curvature parameter. Represents the total number of initially connected components. This represents the natural exponential function. The area representing the closed contour of the shape. Represents pi (π) The radius of the smallest structuring element used in the morphological closing operation is 1. The maximum structuring element radius used in the morphological closing operation is determined by... Divide by 2 and round up to obtain the result. This represents the average Euclidean distance between the centroids of adjacent density contour lines.
8. The machine learning-based immune marker analysis system according to claim 7, characterized in that, The extraction of the time evolution pattern employs a constrained dynamic time warping algorithm, with the slope of the curved path limited as follows: ; The calculation of the sample confidence score incorporates the co-expression strength value. and the decomposed residual matrix : ; in, Represents the index increment in the i-direction of the time series. The index increment in the j-direction of the time series is represented. This represents the number of time slices. Represents the logarithmic function with base 2. A three-dimensional matrix representing the strength values of collaborative expression, with dimension 1. , Represents the number of immune subtypes. Represents the initial number of markers. A three-dimensional matrix representing the decomposed residual matrix with dimensions equal to 1. same, The confidence score represents the sample. The L1 norm of a matrix is the sum of the absolute values of all its elements. The function is used to take the larger of the two values. This represents the Hadamard product operation.
9. The machine learning-based immune marker analysis system according to claim 8, characterized in that, The spatial neighborhood information entropy is calculated using adaptive bandwidth kernel density estimation: bandwidth Local variance of the three-dimensional data of the immune markers satisfy ; The window length of the moving average method Dynamically adapt the number of data voxels ,satisfy ; in, The bandwidth parameter representing kernel density estimation, This represents the variance of the data within a 5×5×5 cube neighborhood centered at the current voxel. This represents the total sample size of the three-dimensional data of the immune markers. represent -1 / 5, This represents the length of the moving average window. Represents the total number of data voxels.
10. The machine learning-based immune marker analysis system according to claim 9, characterized in that, The kernel density estimation algorithm employs a modified Epanechnikov kernel function: ; Streamline tracking step size of gradient vector field Resolution of each axis of the three-dimensional data satisfy , All units were converted to millimeters before being used in the calculation. in, The kernel function representing kernel density estimation, Represents the normalized distance parameter and , For the current voxel coordinates, Let the coordinates be those of the neighboring voxels. The step size represents the streamline tracing step. Represents the resolution of the 3D data along the x-axis. Represents the resolution of the 3D data along the y-axis. Represents the resolution of the 3D data along the z-axis. The function is used to find the minimum value among the three resolutions.
Citation Information
Patent Citations
Quantitative analysis model of protein based on immunofluorescence image and establishment method
CN113724195A
Method of training machine learning model for analyzing immunohistochemical staining images and computing system executing same
CN118985027A
Intelligent management system for energy consumption optimization and fault self-diagnosis of cleaning equipment
CN120258775A
Mathematical algorithm calculation system and optimization method based on big data
CN120371256A