Machine learning based immune signature analysis system
By using a machine learning-based immune marker analysis system to process three-dimensional immune data and construct a boundary topology model, the system addresses the shortcomings of inconsistent interpretation standards and threshold judgment models in traditional systems, thereby achieving multi-dimensional data fusion of highly specific immune markers and accurate disease diagnosis.
Patent Information
- Application Number
- CN202511468519.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Traditional immune marker analysis systems are susceptible to subjective experience, have inconsistent interpretation standards, and are difficult to adapt to the dynamic expression changes of cell population surface markers. Threshold determination models cannot effectively capture nonlinear correlation features, resulting in insufficient accuracy in the construction of disease diagnostic models.
An immune marker analysis system based on machine learning is adopted. The system processes three-dimensional data through an entropy density rheology module to generate an entropy density gradient map of the markers. Combined with a boundary dynamic annotation module and a co-tensor modeling module, a boundary topology model is constructed using directional derivatives and morphological closing operations. The core factor matrix is then extracted to perform highly specific immune marker set discrimination.
This approach enables multi-dimensional data fusion of immune markers, significantly improving the completeness of biomarker correlation analysis and the accuracy of quantitative analysis of immune markers, thereby enhancing the accuracy of disease diagnostic models.
Smart Images

Figure CN120954685B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical data analysis, and in particular to an immune marker analysis system based on machine learning. BACKGROUND
[0002] The technical field of medical data analysis involves a technical system for extracting effective information from clinical test data to support disease diagnosis decision-making, the core of which is the integrated processing, feature pattern recognition and visualized result output of multi-source heterogeneous medical test data. This field covers key technical links such as laboratory test data standardization processing, biomarker correlation analysis, diagnosis model construction, and requires the comprehensive use of data cleaning algorithms, statistical analysis methods and visualization presentation techniques to realize the transformation of medical test data into clinical diagnosis information. The traditional immune marker analysis system refers to an analysis device that calculates antibody concentration based on enzyme-linked immunosorbent assay test data through optical density value, which uses manual interpretation combined with semi-automatic image processing software to realize quantitative analysis of test results, or relies on flow cytometry test data to identify the expression level of specific cell surface markers through gating strategy, and uses threshold determination model to quantitatively analyze fluorescence signal intensity.
[0003] The traditional immune marker analysis system relies on manual interpretation and semi-automatic image processing software, which is easily influenced by subjective experience in complex and variable clinical testing scenarios, leading to inconsistent interpretation standards. The fixed gating strategy used in flow cytometry is difficult to adapt to the dynamic expression changes of cell surface markers, and the linear analysis method of threshold determination model for fluorescence signal intensity cannot effectively capture nonlinear correlation characteristics. In the standardization processing of laboratory test data, there is a lack of multi-dimensional data collaborative analysis mechanism, resulting in the loss of potential correlation information between different time slices and immune subtypes. The sensitivity of semi-quantitative analysis methods to low-density region data is insufficient, and boundary recognition bias is prone to occur in antibody concentration gradient mutation regions, ultimately affecting the accuracy of disease diagnosis model construction and the effectiveness of clinical decision support. SUMMARY
[0004] The purpose of the present application is to solve the problems existing in the prior art, and to provide an immune marker analysis system based on machine learning.
[0005] In order to achieve the above purpose, the present application adopts the following technical scheme: an immune marker analysis system based on machine learning, comprising:
[0006] An entropy density rheological module is used to process immune marker three-dimensional data through kernel density estimation algorithm, calculate spatial neighborhood information entropy, generate marker entropy density gradient map using sliding average method, and pass the marker entropy density gradient map to a boundary dynamic labeling module.
[0007] The boundary dynamic labeling module is configured for calculating a directional derivative based on the marker entropy density gradient map, detecting 8-neighbor gradient mutation points, and generating an immune marker boundary topology map by using a morphological closing operation of a dynamic structure element, and transferring the immune marker boundary topology map to the collaborative tensor modeling module.
[0008] The collaborative tensor modeling module is configured for cropping low-density data according to the immune marker boundary topology map, constructing an immune subtype Time slice The marker third-order tensor is configured for extracting a core factor matrix by using a dimension-adaptive Tucker decomposition algorithm, calculating a Manhattan distance to generate a collaborative expression intensity value, and transferring the collaborative expression intensity value and a decomposition residual matrix to the marker discrimination module.
[0009] The marker discrimination module is configured for performing a Hadamard product operation on the collaborative expression intensity value and the decomposition residual matrix, inputting a support vector machine classifier for discrimination, and generating a high-specificity immune marker set.
[0010] As a further scheme of the application, the entropy density gradient map specifically includes a spatial entropy distribution feature, a gradient vector field, and a density contour line, the boundary topology map contains a boundary curvature parameter, a topological connected domain, and a morphological closed contour, the collaborative expression intensity value covers a core factor load, a time evolution mode, and a Manhattan distance matrix, the decomposition residual matrix includes an orthogonal residual component, a tensor reconstruction error, and a low-rank approximation deviation, and the high-specificity immune marker set specifically refers to a classification decision surface, a feature weight vector, and a sample confidence score.
[0011] The radius of the dynamic structure element The standard deviation of the length of the gradient vector field Satisfies , obtained by calculating the standard deviation of all length values of the gradient vector field.
[0012] wherein, represents the radius of the morphological closing operation structure element, represents the standard deviation of all length values of the gradient vector field, represents a logarithmic function with a base of 2, represents a rounding-up operation.
[0013] The interval of the density contour line obtained by calculating a moving average value of the Euclidean distance between the centers of adjacent contour lines, and the moving average window size is an integer value obtained by taking the square root of the number of three-dimensional data voxels.
[0014] As a further scheme of the present application, the calculation of the directional derivative employs a modified Sobel-Feldman operator whose kernel size is dynamically adapted to the average spacing of the density contour :
[0015] ;
[0016] wherein is the average Euclidean distance of the density contour in the gradient vector field, obtained by computing the L2 norm average of the coordinate difference between adjacent contour centroids;
[0017] the radius of the structuring element of the morphological closing operation satisfies ;
[0018] wherein, represents the radius of the structuring element of the morphological closing operation, represents the set of all modulus values in the gradient vector field, represents the arithmetic average of the modulus values in the gradient vector field, represents the maximum value in the set of modulus values in the gradient vector field, denotes the ceiling operation.
[0019] As a further scheme of the present application, the dimension of the core tensor in the Tucker decomposition algorithm is set to follow: the first dimension ;
[0020] the second dimension ;
[0021] the third dimension ;
[0022] the calculation of the Manhattan distance matrix employs a sliding window mechanism, the window width satisfies , , is selected by cross-validation method in the preset interval {[}3,8{]};
[0023] wherein, represents the first dimension size of the core tensor, represents the total number of immune subtype categories, denotes the logarithm function with base 2, denotes the ceiling operation, represents the second dimension size of the core tensor, represents the number of time slices, denotes the floor operation, represents the third dimension size of the core tensor, This represents the total number of initial markers.
[0024] As a further aspect of the present invention, the kernel function parameters of the support vector machine classifier Dynamically adapt the Frobenius norm of the decomposed residual matrix :
[0025] ;
[0026] The update step size of the feature weight vector With respect to the low-rank approximation deviation Negative correlation, satisfying , For the original third-order tensor, and All values are standardized dimensionless values;
[0027] in, The kernel function parameters represent the kernel parameters of the support vector machine classifier. Represents the total number of initial markers. This represents the variance of all elements in the decomposed residual matrix. The Frobenius norm represents the residual matrix of the decomposition. The function is used to ensure that the denominator is not zero. The update step size represents the feature weight vector. The numerical value representing the low-rank approximation deviation. Frobenius norm represents the original third-order tensor.
[0028] As a further aspect of the present invention, the extraction process of the orthogonal residual components includes:
[0029] Gram-Schmidt orthogonalization is performed on the core factor matrix to generate a set of orthogonal basis vectors;
[0030] Calculate the projection components of the decomposed residual matrix onto each orthogonal basis vector;
[0031] Retaining projection coefficients exceeding the threshold The components, of which The i-th singular value of the core factor matrix is obtained through singular value decomposition.
[0032] threshold The calculation result should be rounded to three decimal places.
[0033] in, The screening threshold representing the projection coefficient. This represents the size of the first dimension of the core tensor. represents the i-th singular value of the core factor matrix after singular value decomposition, is an integer index from 1 to .
[0034] As a further scheme of the present application, the determination of the topological connected domain employs an improved 8-neighbor connected algorithm, and the merging threshold is set as:
[0035] ;
[0036] The area screening condition of the morphologically closed contour is , , ;
[0037] wherein, represents the merging threshold of the connected domain, represents the arithmetic mean of all values of the boundary curvature parameter, represents the standard deviation of all values of the boundary curvature parameter, represents the total number of initial connected domains, denotes the natural exponential function, represents the area of the morphologically closed contour, represents the constant pi, represents the minimum structural element radius adopted by the morphological closing operation and takes the value of 1, represents the maximum structural element radius adopted by the morphological closing operation, which is obtained by rounding up the value of divided by 2, represents the average Euclidean distance between adjacent density contour centroids.
[0038] As a further scheme of the present application, the extraction of the time evolution pattern employs a constrained dynamic time warping algorithm, and the slope limit of the curved path is:
[0039] ;
[0040] The calculation of the sample confidence score fuses the synergistic expression intensity value and the decomposition residual matrix :
[0041] ;
[0042] wherein, represents the index increment in the time series i direction, represents the index increment in the time series j direction, represents the number of time slices, denotes the logarithmic function with base 2, a three-dimensional matrix representing the co-expression intensity values and having dimensions , a three-dimensional matrix representing the number of immune subtype classes, a three-dimensional matrix representing the number of initial markers, a three-dimensional matrix representing the decomposition residual matrix and having the same dimensions as , a three-dimensional matrix representing the sample confidence score, denotes the L1 norm of a matrix, i.e. the sum of the absolute values of all elements, the function is used to take the larger value of the two, denotes the Hadamard product operation.
[0043] As a further aspect of the present application, the calculation of the spatial neighborhood information entropy employs an adaptive bandwidth kernel density estimation: the bandwidth is proportional to the local variance of the immunological marker three-dimensional data satisfies ;
[0044] the window length of the moving average method is dynamically adapted to the number of data voxels satisfies ;
[0045] wherein denotes the bandwidth parameter of the kernel density estimation, denotes the variance of the data within a 5x5x5 cubic neighborhood centered at the current voxel, denotes the total sample size of the immunological marker three-dimensional data, denotes the -1 / 5th power of , denotes the length of the moving average window, denotes the total number of data voxels.
[0046] As a further aspect of the present application, the kernel density estimation algorithm employs an improved Epanechnikov kernel function:
[0047] ;
[0048] the step size of the streamline tracking of the gradient vector field is proportional to the axial resolution of the three-dimensional data along each axis satisfies , all converted to millimeter units before participating in the calculation;
[0049] wherein denotes the kernel function of the kernel density estimation, denotes the normalized distance parameter and , is the current voxel coordinate, is a coordinate of a neighborhood voxel, represents a step size of streamline tracking, represents a resolution of the three-dimensional data in the x-axis direction, represents a resolution of the three-dimensional data in the y-axis direction, represents a resolution of the three-dimensional data in the z-axis direction, The function is used to take the minimum value in the three resolutions.
[0050] Compared with the prior art, the application has the advantages and positive effects that:
[0051] In the application, the three-dimensional immune data is processed by a kernel density estimation algorithm and the spatial information entropy is calculated, the spatial distribution characteristics of the multi-source heterogeneous medical data are effectively integrated, the gradient graph generated by the sliding average method can dynamically capture the continuous trend of the concentration change of the marker, the boundary topology model constructed based on the directional derivative detection and the morphological closing operation can accurately identify the morphological boundary characteristics of the immune marker, the three-order tensor modeling is used to realize the multi-dimensional data fusion of the immune subtype, the time dimension and the marker expression, the core factor matrix is extracted through the dimension-adaptive tensor decomposition algorithm, the integrity of the biomarker correlation analysis is significantly improved, the Manhattan distance quantified collaborative expression intensity value is combined with the Hadamard product operation of the residual matrix, the recognition accuracy of the high-specificity marker by the support vector machine classifier is enhanced, and finally the technical leap from the single threshold determination to the multi-dimensional collaborative determination of the quantitative analysis of the immune marker is realized. BRIEF DESCRIPTION OF DRAWINGS
[0052] Figure 1 is a general flowchart of the immune marker analysis system based on machine learning of the application;
[0053] Figure 2 is a flowchart of the entropy density rheological module of the application;
[0054] Figure 3 is a flowchart of the boundary dynamic labeling module of the application;
[0055] Figure 4 is a flowchart of the collaborative tensor modeling module of the application;
[0056] Figure 5 is a flowchart of the marker discrimination module of the application. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical scheme and advantages of the application clearer, the technical scheme realized by software will be described in detail below in combination with the system architecture diagram and the embodiments. It should be understood that the specific embodiments described herein are only used to explain the technical scheme of the application and do not constitute a limitation on the protection scope.
[0058] In the description of the present application, the system architecture relationship or data processing flow indicated by the terms "hierarchy", "module", "interface", "data flow", "client", "server" and the like are defined based on the corresponding architecture diagram or flowchart of the embodiment. This way of expression is only used to clearly explain the logical relationship of each element in the technical solution, and is not limited to the physical deployment form. The "multiple" contains two or more technical units, including but not limited to multiple data nodes, processing threads, service instances or functional components, and other scalable elements. The specific number is determined according to the actual business scenario.
[0059] Please refer to Figure 1 and Figure 2 , the present application provides a technical solution: an immune marker analysis system based on machine learning includes:
[0060] An entropy density rheological module is used to process the immune marker three-dimensional data by a kernel density estimation algorithm, calculate the spatial neighborhood information entropy, generate a marker entropy density gradient map using a sliding average method, and pass the marker entropy density gradient map to a boundary dynamic labeling module.
[0061] The entropy density gradient map specifically includes spatial entropy distribution characteristics, gradient vector field, and density contour lines.
[0062] The calculation of the spatial neighborhood information entropy uses an adaptive bandwidth kernel density estimation: bandwidth and the local variance of the immune marker three-dimensional data satisfy ;
[0063] wherein, represents the bandwidth parameter of the kernel density estimation, represents the variance of the data in the 5x5x5 cubic neighborhood centered on the current voxel, represents the total sample size of the immune marker three-dimensional data, represents -1 / 5 power;
[0064] The kernel density estimation algorithm uses an improved Epanechnikov kernel function: ;
[0065] wherein, represents the kernel function of the kernel density estimation, represents the normalized distance parameter and , is the current voxel coordinate, is the neighborhood voxel coordinate;
[0066] The window length of the sliding average method is dynamically adapted to the number of data voxels , satisfying ;
[0067] wherein, represents the length of the moving average window, represents the total number of data voxels;
[0068] the interval of the density contour obtained by calculating the moving average of the Euclidean distance between adjacent contour centroids, with a moving average window size of the square root of the number of voxels in three dimensions, rounded to the nearest integer;
[0069] the step size for streamline tracking of the gradient vector field the axial resolution of the three-dimensional data in each axis satisfies , all converted to millimeter units before calculation;
[0070] wherein, represents the step size for streamline tracking, represents the resolution of the three-dimensional data in the x-axis direction, represents the resolution of the three-dimensional data in the y-axis direction, represents the resolution of the three-dimensional data in the z-axis direction, the function is used to take the minimum value of the three resolutions.
[0071] First, the immune marker three-dimensional data to be analyzed is obtained. This data comes from a biopsy tissue sample of a lymphoma patient, obtained using multi-color immunofluorescence (mIF) imaging technology, with the imaging equipment being a Leica STELLARIS 8DIVE and the objective magnification being 40 times. This three-dimensional data is stored in digital form, specifically as a series of consecutive two-dimensional image slices stacked together, with each pixel point recording the fluorescence intensity value of a specific immune marker (such as CD8, PD-1, FoxP3). In this embodiment, a tissue region with a size of 500 microns x 500 microns x 100 microns is intercepted for analysis, which is digitized into a 1000 x 1000 x 200 voxel three-dimensional data array. Thus, the total sample size of the immune marker three-dimensional data is 2 x 10 8 data points. The axial resolution of the three-dimensional data in each axis, after conversion to millimeter units, is = 0.0005 mm, = 0.0005 mm, = 0.0005 mm.
[0072] Next, the three-dimensional data is processed by kernel density estimation algorithm to calculate the spatial neighborhood information entropy. This process is performed at each voxel position. Taking the voxel with coordinates (150, 250, 80) as an example, first determine its spatial neighborhood. The neighborhood here is a 5x5x5 cubic region centered on the voxel, containing a total of 125 (5x5x5=125) voxels. Then, calculate the local variance of the 125 data points (i.e. fluorescence intensity values) in the neighborhood In this example, by statistical calculation of the 125 fluorescence intensity values, the local variance is obtained 215.3. Subsequently, based on the local variance and the total sample size the adaptive bandwidth is calculated.
[0073] The formula is used to calculate the adaptive bandwidth of kernel density estimation. Wherein, represents the bandwidth parameter, whose value is dynamically adjusted according to the local density of the data, decreasing in dense areas and increasing in sparse areas. The parameter is an empirical selection-based scaling factor, determined by bandwidth optimization experiments on more than 500 different tissue type immunofluorescence image data. The experimental process is as follows: for each data, a series of bandwidth values (from 0.5 to 2.0, step 0.1) are preset, kernel density estimation is performed using each bandwidth value, and the KL divergence of the obtained density distribution and the true cell distribution (annotated manually by a pathologist) is calculated; select the scaling factor that minimizes the average KL divergence of all samples, and the experimental results show that 1.06 can achieve optimal or suboptimal estimation effect in various scenarios. represents the local variance, reflecting the dispersion degree of the data in the neighborhood. is the total sample size, and the exponential term is the optimal bandwidth scaling factor derived based on the minimum integrated squared error (MISE) criterion. The innovation of this formula is to combine the reflecting the global volume of data with the reflecting the local features, achieving adaptive adjustment of the bandwidth and avoiding the problem of over-smoothing or under-smoothing when using a global uniform bandwidth to process data with uneven density. The above numerical values are substituted into the formula to calculate: , then . . Therefore, the bandwidth used for kernel density estimation of voxel (150, 250, 80) is 4.98.
[0074] After obtaining the bandwidth , the information entropy of the voxel with coordinates (152, 251, 81) in the neighborhood is estimated using the improved Epanechnikov kernel function. For example, first calculate its relationship with the central voxel. Normalized distance parameter . .because Therefore, the first part of the kernel function is used for calculation. Formula This is an improved Epanechnikov kernel function. Wherein, The kernel function value determines the contribution weight of neighboring voxels to the central voxel density estimation. It is the normalized distance parameter, derived from the central voxel. With neighboring voxels Euclidean distance Divide by bandwidth The innovation of this kernel function lies in the fact that it is a piecewise function, which, compared to the standard Epanechnikov kernel function, expands the range of influence from... Expanded to This creates a "soft boundary". When When, it maintains the optimality of the Epanechnikov kernel in the sense of mean square integral error; when At that time, the weights smoothly decrease to zero, instead of... The abrupt change at a certain point is zero. This design, while maintaining computational efficiency, can handle voxels at neighborhood boundaries more smoothly, reducing boundary effects in the estimation results and improving the robustness of density estimation. Substitute into the formula: This calculation is repeated for all 124 neighboring voxels within a 5×5×5 neighborhood, and all kernel function values are summed with weights to obtain the entropy density value at the central voxel (150, 250, 80). After traversing all voxels, a complete spatial entropy distribution feature map is generated.
[0075] Subsequently, the spatial entropy distribution feature map is smoothed using the moving average method to generate an entropy density gradient map of the markers. First, the window length of the moving average is determined. Total number of data voxels 1000×1000×200=2×10 8 . formula Used to dynamically adapt the length of the moving average window. Among them, For window length, For floor operations, The total number of data voxels. This formula determines the window size using the square root of the data voxel count, allowing the smoothness to adapt to different data sizes. A larger window is used for large datasets to effectively suppress noise, while a smaller window is used for small datasets to retain more detail. Adding 1 ensures the window length is odd, giving the window a unique center point. Substitute into the formula: This window length applies to one-dimensional entropy density data. In three-dimensional space, a [length] is typically used. The cubic window, in which Here, we assume a cubic window with a side length of 585 voxels. For each point in the entropy density map, we calculate the average of all entropy density values within its 585×585×585 neighborhood, using this as the smoothed value for that point. On the smoothed entropy density map, we calculate the gradient, forming a gradient vector field. Then, we draw density contour lines based on the gradient vector field. The spacing between the density contour lines... The distance between the centroids of adjacent contour lines is obtained by calculating the moving average of the distance between them. The size of the moving average window is set to the square root of the number of voxels in the 3D data, which is 14142. For example, if the centroid coordinates of three adjacent contour lines are calculated as C1(10,20,30), C2(12,23,31), and C3(15,25,33), then the distance between them is... , Assuming these two distance values belong to a portion of the moving average window, the final average distance is obtained by averaging all 14,142 such distance values within the window. The calculation obtained in this embodiment is as follows: Finally, the streamline tracking step size of the gradient vector field is calculated. . formula Used to calculate the streamline tracking step size. Step size, It is the resolution (in millimeters) of each axis of the three-dimensional data. The function takes the minimum value among the three. The coefficient 0.5 is to comply with the Nyquist sampling theorem, ensuring that the tracking step size is less than half of the minimum resolution, thereby avoiding missing key gradient change features during streamline tracing. Substituting the resolution value into the formula: mm. This step size This data is used for subsequent streamline visualization and analysis. Finally, the marker entropy density gradient map, which includes spatial entropy distribution characteristics, gradient vector field, and density contour lines, is passed to the boundary dynamic annotation module.
[0076] Please see Figure 1 and Figure 3 The boundary dynamic annotation module is used to calculate the directional derivative based on the entropy density gradient map of the marker, detect the gradient abrupt change point in the 8-neighborhood, generate the immune marker boundary topology map by using the morphological closing operation of the dynamic structuring element, and pass the immune marker boundary topology map to the co-tensor modeling module.
[0077] The boundary topology graph includes boundary curvature parameters, topological connected components, and morphological closed contours.
[0078] The calculation of directional derivative adopts an improved Sobel-Feldman operator, whose convolution kernel size dynamically adapts to the average interval of density contours : ;
[0079] wherein is the average Euclidean distance of density contours in the gradient vector field, which is obtained by calculating the average value of L2 norm of coordinate difference between adjacent contour centroids;
[0080] the radius of dynamic structuring element the standard deviation of the length of gradient vector field satisfies , which is obtained by calculating the standard deviation of all length values of the gradient vector field;
[0081] wherein, represents the radius of morphological closing structuring element, represents the standard deviation of all length values of the gradient vector field, represents the logarithm function with base 2, represents the upward rounding operation;
[0082] the radius of morphological closing structuring element satisfies ;
[0083] wherein, represents the radius of morphological closing structuring element, represents the set of all length values of the gradient vector field, represents the arithmetic mean of length values of the gradient vector field, represents the maximum value in the set of length values of the gradient vector field, represents the upward rounding operation;
[0084] The determination of topologically connected domain adopts an improved 8-neighbor connected algorithm, and the merging threshold is set as: ;
[0085] wherein, represents the merging threshold of connected domain, represents the arithmetic mean of all values of boundary curvature parameter, represents the standard deviation of all values of boundary curvature parameter, represents the total number of initial connected domains, represents the natural exponential function;
[0086] The area screening condition of morphologically closed contour is , , ;
[0087] wherein, represents the area of the morphologically closed contour, represents the value of pi, represents the minimum structuring element radius used in the morphological closing operation and takes the value 1, represents the maximum structuring element radius used in the morphological closing operation and is obtained by rounding up to the nearest integer after dividing by 2, represents the average Euclidean distance between the centroids of adjacent density contour lines.
[0088] This module first calculates directional derivatives based on the marker entropy density gradient map to identify the boundaries of immune cell aggregates and tissue stroma. This process uses an improved Sobel-Feldman operator. The size of the convolution kernel of this operator is dynamically adapted according to the average distance between density contour lines calculated in the previous module . In the aforementioned embodiment, the average distance between density contour lines has been calculated.
[0089] The formula is used to generate the convolution kernel in the x direction. Wherein, is the convolution kernel matrix, is the floor operation, is the average Euclidean distance of the density contour lines. The innovation of this formula is to associate the weights (1 and 2) in the standard Sobel-Feldman operator with the contour line spacing . When is large, it means that the entropy density changes smoothly and the boundary is blurred, at this time, by increasing the weight of the convolution kernel, the response to weak edges can be enhanced; on the contrary, when is small, it means that the boundary is clear, then use a smaller weight to avoid over-enhancing noise. This adaptive mechanism makes the sensitivity and accuracy of edge detection stable in image regions of different clarity. Bring into the formula to calculate the weight factor: . Therefore, the convolution kernel in the x direction and the convolution kernel in the y direction are respectively: , . Apply these two convolution kernels to each 3x3 pixel neighborhood of the entropy density gradient map to calculate the gradient components Gx and Gy in the x and y directions, respectively, to obtain the gradient magnitude and direction of this point. By traversing the entire image, points where the gradient value suddenly changes in the 8-neighborhood are preliminarily marked as boundary points.
[0090] Next, the morphological closing operation with dynamic structuring element is applied to the preliminary marked boundary points to connect the broken boundaries and fill the small holes, and to generate the complete immunolabel boundary topology map. The radius of the structuring element is calculated as follows: According to the standard deviation of the gradient vector field module Dynamic calculation. First, the standard deviation of all the module values in the entire gradient vector field is calculated In this embodiment, the standard deviation of the gradient amplitude of all points in the gradient map is calculated
[0091] The formula is used to calculate the radius of the morphological closing operation structuring element. Wherein, is the radius, is the rounding up operation, is the logarithm with base 2, is the standard deviation of the gradient vector field module. The innovation of this formula is to link the size of the structuring element with the overall dispersion of the gradient change (i.e. ). When is large, it means that the intensity and clarity of the boundary in the image differ greatly, and there may be wide or blurred boundary regions, so a larger structuring element is needed to bridge these uncertain regions; when is small, it means that the boundary is consistent and clear, so a smaller structuring element can be used, thereby avoiding the loss of details caused by excessive smoothing. The adaptability of this logarithmic scale makes the response of the change of the size of the structuring element to the dispersion of the gradient more smooth and robust. Bring into the formula: Therefore, a circular (or spherical, in three dimensions) structuring element with a radius of 3 is used to perform the morphological closing operation. The closing operation process is to dilate first and then erode, effectively connecting adjacent boundary points to form a preliminary morphological closed contour.
[0092] Subsequently, the generated multiple closed contours are subjected to topological connected component determination to merge regions belonging to the same immune cell cluster. This process uses an improved 8-neighbor connected component algorithm, and the merging threshold is dynamically set. First, the boundary curvature parameters of all initial connected components are calculated to obtain a set of curvature values. In this embodiment, a total of initial connected components are identified, and the average curvature and the curvature standard deviation are calculated.
[0093] The formula is used to calculate the threshold for merging connected components. Wherein, is the merging threshold, and respectively the arithmetic mean and the standard deviation of the boundary curvature of all initial connected components, is the total number of initial connected components, is the natural exponential function. The innovation of this formula lies in its adaptivity: the basis of the threshold is the average curvature which is adjusted according to the standard deviation of the curvature and a sigmoid function term related to the number of connected components . When the number of initial connected components is small, the exponential term in the denominator approaches 1, the adjustment term is large, and the tendency is to merge. When is very large, the exponential term tends to 0, the denominator approaches 1, and the adjustment is also large, which seems a bit counterintuitive, but can be explained as follows: when the regions are over-segmented (Nc is large), a more lenient threshold is needed to merge them. The sigmoid curve transition in the middle ensures the smoothness of the threshold variation. This design makes the merging decision not only consider the average bending degree of the region boundary, but also consider the dispersion of the bending degree and the fragmentation degree, so as to realize more intelligent region merging. Bring into the formula: . Set the merging threshold to 0.33. If the average boundary curvature between two adjacent connected components is less than 0.33, they will be merged into one connected component.
[0094] Finally, the area of the closed contour after merging is screened to remove small regions that may be noise or artifacts. The screening condition is related to the morphological operation parameters and the contour interval. The minimum structural element radius is set to 1, and the maximum structural element radius is calculated according to . The formula is used to determine the maximum structural element radius. Bring into: . The area screening threshold is determined by the formula . Where A is the contour area, is the circumference. The innovation of this formula is to link the lower limit of area screening with the minimum and maximum feature sizes that can be handled by morphological operation. Logically, any contour smaller than the "average feature area" defined by the minimum and maximum structural elements may be unreliable and should be removed. Calculate the area threshold: . All closed contours with an area less than 3.927 voxel units will be removed. After all the above steps, the final immune marker boundary topology map contains accurate boundary curvature parameters, topological connected components, and screened morphological closed contours, and is passed to the collaborative tensor modeling module.
[0095] Please refer to Figure 1 and Figure 4, the synergistic tensor modeling module is configured to crop low-density data according to the immune marker boundary topology graph, construct a third-order tensor of immune subtypes x time slices x markers, extract core factor matrices by using a dimension-adaptive Tucker decomposition algorithm, calculate Manhattan distance to generate synergistic expression intensity values, and pass the synergistic expression intensity values and a decomposition residual matrix to the marker discrimination module;
[0096] The synergistic expression intensity values include core factor loadings, time evolution patterns, and Manhattan distance matrices.
[0097] The decomposition residual matrix includes orthogonal residual components, tensor reconstruction errors, and low-rank approximation deviations.
[0098] The core tensor dimension in the Tucker decomposition algorithm is set according to the following rules: the first dimension ;
[0099] The second dimension ;
[0100] The third dimension ;
[0101] wherein represents the size of the first dimension of the core tensor, represents the total number of immune subtype classifications, represents a logarithm function with a base of 2, represents a rounding-up operation, represents the size of the second dimension of the core tensor, represents the number of time slices, represents a rounding-down operation, represents the size of the third dimension of the core tensor, and represents the total number of initial markers.
[0102] The Manhattan distance matrix is calculated by using a sliding window mechanism, and the window width and the rank of the core factor matrix satisfy , and are selected by cross-validation in a preset interval [3, 8].
[0103] The extraction process of the orthogonal residual components includes: performing Gram-Schmidt orthogonalization on the core factor matrix to generate a set of standard orthogonal basis vectors.
[0104] The projection components of the decomposition residual matrix on each orthogonal basis vector are calculated.
[0105] Components with projection coefficients exceeding a threshold are retained, wherein is the i-th singular value of the core factor matrix, which is obtained by singular value decomposition.
[0106] threshold The calculation result should be rounded to three decimal places.
[0107] in, The screening threshold representing the projection coefficient. This represents the size of the first dimension of the core tensor. This represents the i-th singular value obtained after singular value decomposition of the core factor matrix. From 1 to Integer index;
[0108] The extraction of the time evolution pattern employs a constrained dynamic time warping algorithm, with the slope of the curved path limited as follows: ;
[0109] in, Represents the index increment in the i-direction of the time series. The index increment in the j-direction of the time series is represented. This represents the number of time slices. This represents the logarithmic function with base 2.
[0110] First, the module crops the original 3D immunofluorescence data based on the immunomarker boundary topology map, setting the low-density region data outside the boundary contour (e.g., fluorescence intensity below a preset value, obtained by subtracting twice the standard deviation from the mean of all pixel intensities within the boundary) to zero. This operation focuses on the core region of immune cell aggregation. Next, a third-order tensor is constructed to integrate multi-dimensional information. The data organization of this tensor is: immunotype × time slice × marker. In this embodiment, the research object is three immunotypes (… (e.g., Treg, Th1, Th2 cells), 49 time points were collected ( For example, tissue samples were collected every 2 hours for 4 days after drug intervention, and 20 different immune markers were tested on each sample. This constructs a 3×49×20 third-order tensor. Each element in the tensor Representative at the At the 1st time point, the 1st Among the various immune subtypes, the first The average fluorescence intensity values of each marker. To illustrate the data structure, a portion of the data (data from time point 1) is listed in tabular form below:
[0111] Table 1. Mean fluorescence intensity of immunomarkers at time point 1.
[0112] Immune Subtype Marker 1 (CD3) Marker 2 (CD4) ... Marker 20 (IL-10) Subtype 1 (Treg) 158.2 120.5 ... 210.7 Subtype 2 (Th1) 165.3 115.1 ... 45.2 Subtype 3 (Th2) 149.8 125.7 ... 98.6
[0113] As shown in Table 1, the table clearly displays the expression levels of various biomarkers in different immune subtypes at the initial time point. The entire third-order tensor contains such data for all 49 time points.
[0114] Subsequently, the core factor matrix is extracted using the dimension-adaptive Tucker decomposition algorithm. The core tensor dimension of this algorithm is dynamically set according to the scale of the input data. The formula for setting the core tensor dimension is: First dimension... The second dimension The third dimension The innovation of these formulas lies in the fact that they do not employ a fixed or empirically chosen rank, but rather combine the dimension of the core tensor (i.e., the complexity of the model) with the size of each mode in the original data. To perform mathematical association. Logarithmic relationships are used because the number of immune subtypes is usually small, and logarithmic growth can effectively compress information without excessively losing class specificity. Using the square root relationship is suitable for capturing the main trends in time series rather than high-frequency noise. The use of a power of 0.4 is an empirical index that balances model fit and generalization ability, derived from empirical analysis of a large amount of proteomics and genomics data. This index was obtained by searching for the optimal average value in the range [0.2, 0.6] through cross-validation experiments on more than 100 public datasets. Substituting the data scale of this example into the calculation: . . Therefore, the core tensor dimension is set to 2×7×4. After performing Tucker decomposition, a 2×7×4 core tensor and three factor matrices are obtained, corresponding to the immune subtype (3×2), time evolution (49×7), and biomarker (20×4) patterns, respectively.
[0115] Next, the Manhattan distance is calculated to generate co-expression strength values. This calculation employs a sliding window mechanism on the biomarker factor matrix. The choice of window width is related to the rank of the core factor matrix. Related. The optimal rank for the factor matrix is selected within the preset interval [3,8] using cross-validation. In this example, cross-validation (e.g., dividing the factor matrix data into 10 folds, training the model with 9 folds each time to predict the remaining 1 fold, and selecting the rank that minimizes the prediction error) indicates the optimal rank. The width of the sliding window is determined by the formula. Confirmed. This formula originates from optimal filter length design in signal processing; its exponential relationship ensures that the window width can cover the range of rank... The main interaction modes at the complexity level represented. Bringing in the formula: . Since the number of markers is less than the window width 63, the sliding window here will cover all markers, i.e. the Manhattan distance of each pair of markers in the 4-dimensional factor space is calculated, generating a 20x20 Manhattan distance matrix, whose element values reflect the co-expression strength between markers.
[0116] In the decomposition process, the orthogonal residual components also need to be extracted. This process first performs Gram-Schmidt orthogonalization on the core factor matrix (take the subtype factor matrix as an example, the size is 3x2) to generate a standard orthogonal basis. Then, the projection components of the decomposition residual matrix on these orthogonal bases are calculated. Components with projection coefficients exceeding the threshold are retained. The threshold is determined by the formula The innovation of this formula is to associate the threshold with the singular value of the factor matrix, which represents the energy or importance of the data in this direction, so (the Frobenius norm of the factor matrix) reflects the total amount of information captured by the factor matrix. The threshold is proportional to this amount of information, so that the screening criteria can be adapted to the importance of different factor matrices. The coefficient 0.1 is determined by ROC analysis on simulation data, aiming to maximize the rejection of injected random noise while retaining 95% of the true signals. First, singular value decomposition (SVD) is performed on the 3x2 subtype factor matrix to obtain two singular values . . . The calculation result retains three decimal places, and the threshold is 1.560. Any residual component with an absolute value of the projection coefficient below 1.560 will be discarded.
[0117] Finally, the extraction of the time evolution pattern uses the constrained dynamic time warping (DTW) algorithm. The slope of the curved path is limited to prevent unreasonable matching. The slope constraint is given by the formula . Where is the local slope of the matching path, is the number of time slices. The innovation of this constraint is that it is not a fixed value, but grows logarithmically with the length of the time series. For longer time series, it allows the matching path to have more flexibility and nonlinearity, while for short sequences, it imposes stronger linear constraints, which conforms to the intuition that long-term trends may be more complex and variable in time series analysis. Bring into the formula: . In the DTW calculation, any matching that results in a local slope of the matching path exceeding 1.5615 will be prohibited. Finally, the co-expression strength value and the decomposition residual matrix are passed to the marker discrimination module together.
[0118] Please refer to Figure 1 and Figure 5 , a marker discrimination module, for performing Hadamard product operation on the synergistic expression intensity value and the decomposition residual matrix, inputting a support vector machine classifier for discrimination, and generating a high-specificity immune marker set;
[0119] The high-specificity immune marker set specifically refers to a classification decision surface, a feature weight vector, and a sample confidence score.
[0120] Kernel function parameters of the support vector machine classifier Dynamic adaptation of the Frobenius norm of the decomposition residual matrix : ;
[0121] wherein, represents the kernel function parameters of the support vector machine classifier, represents the total number of initial markers, represents the variance of all elements of the decomposition residual matrix, represents the Frobenius norm of the decomposition residual matrix, the function is used to ensure that the denominator is not zero;
[0122] Updating step of the feature weight vector and the low-rank approximation bias are negatively correlated, satisfying , is the original three-order tensor, and both use normalized dimensionless values;
[0123] wherein, represents the updating step of the feature weight vector, represents the numerical value of the low-rank approximation bias, represents the Frobenius norm of the original three-order tensor;
[0124] The calculation of the sample confidence score combines the synergistic expression intensity value and the decomposition residual matrix : ;
[0125] wherein, represents a three-dimensional matrix of the synergistic expression intensity value and has a dimension of , represents the number of immune subtypes, represents the number of initial markers, represents a three-dimensional matrix of the decomposition residual matrix and has the same dimension as , represents the sample confidence score, L1 norm of a matrix, i.e., the sum of the absolute values of all elements, The function is used to take the larger value of the two, denotes the Hadamard product operation.
[0126] First, the module will cooperate the matrix of expression intensity values (One reflects the core correlation between markers That is, a 3x49x20 three-order matrix) and the decomposition residual matrix (A three-order matrix of the same dimension containing secondary information ignored by the core model) to perform the Hadamard product operation. This operation is element-wise multiplication, denoted as The purpose of this operation is to integrate the main mode and residual information. Those features with large values in both the core mode and the residual (which may represent nonlinear or unique interactions) will be amplified, thereby providing more abundant features for subsequent classifiers.
[0127] Next, the result of the Hadamard product is sent as input to a support vector machine (SVM) classifier for discrimination, aiming to select high-specificity immune markers from the 20 initial markers. The goal of the classifier is to distinguish, for example, the immune subtype state that can effectively predict drug response. The performance of the SVM classifier depends largely on the kernel function and its parameters. The present embodiment uses a Gaussian kernel (RBF kernel), whose key parameter is dynamically adapted according to the characteristics of the decomposition residual matrix.
[0128] The kernel function parameter is dynamically adapted by the formula The innovation of this formula is that it directly links the width of the kernel function (determined by the reciprocal of ) to the statistical characteristics of the input data. The in the denominator is a scaling factor commonly used in Hebb learning rules and random matrix theory, which is used to standardize the data scale according to the number of features and variance. The numerator is a robust normalization term. When the Frobenius norm of the residual matrix is very small (close to zero, which may lead to numerical instability), this term is approximately 1, avoiding the explosive growth of ; otherwise, the term is 1. This design allows to be adaptively adjusted according to the variance, norm, and feature dimension of the data itself, avoiding time-consuming manual parameter tuning or cross-validation processes. The constant is a small positive number set to prevent division by zero, and its value is selected through thousands of numerical experiments to ensure that it is safe enough under double-precision floating-point operations and does not affect most calculation results. In the present embodiment, The residual matrix of decomposition is calculated The variance of all elements The Frobenius norm of the residual matrix These values are brought into the formula: This value is used to configure the SVM classifier.
[0129] In the training process of SVM, the update step of feature weight vector is also dynamically adjusted if gradient descent or other methods are used for optimization. The update step is determined by the formula . Wherein, is the update step, is the low-rank approximation deviation (i.e. the reconstruction error of Tucker decomposition), is the Frobenius norm of the original third-order tensor. Both and used here are dimensionless values after standardization, for example, obtained by dividing the original value by . The innovation of this formula is to associate the learning rate with the goodness of fit of the model (reflected by the relative reconstruction error ). When the model fits well (small ), it means that the current weight is close to the optimal, and a smaller step should be used for fine-tuning; when the model does not fit well (large ), a larger step can be used to speed up convergence. The coefficient 0.01 is an initial learning rate benchmark set based on experience. Through convergence tests on multiple standard machine learning datasets, it is found that the effect is good in the range of 0.005 to 0.02, and 0.01 is one of the robust choices. In this embodiment, the low-rank approximation deviation is calculated, and the Frobenius norm of the original tensor . This update step is used for weight adjustment in the classifier training process.
[0130] After training, the SVM model generates a classification decision surface and a feature weight vector for each marker. In addition, to evaluate the reliability of the classification result of each sample, the system also calculates the sample confidence score . The sample confidence score is calculated by the formula . Wherein, is the synergistic expression intensity value matrix, is the decomposition residual matrix, is the Hadamard product, is the L1 norm (the sum of the absolute values of all elements of the matrix), The function takes the larger of the two values. The innovation of this formula lies in its use of a normalized metric to quantify the consistency or resonance strength between "core information" and "residual information." (Molecule) This amplifies signals that are in the same direction (both positive or both negative) and have large absolute values. The denominator uses... Normalization, compared to using the sum or product of the two, is more efficient for... and It is more robust to cases with large norm differences. A high The values indicate that the information contained in the residuals is highly correlated with the core patterns extracted by tensor decomposition, which enhances the confidence in classification decisions. Taking a specific sample as an example, the L1 norm of its co-expression strength matrix was calculated. L1 norm of the decomposed residual matrix The L1 norm of the Hadamard product of the two . The result of 0.536 indicates that the classification confidence of this sample is moderately high. A threshold (e.g., 0.75) can be set for the confidence score to filter out classification results with high confidence. Finally, based on the feature weight vector output by the SVM, the 20 biomarkers are ranked, and the top 5 biomarkers with the highest absolute weight values are identified as a set of high-specificity immune biomarkers, such as {PD-L1, GZMB, Ki-67, CD8, CTLA-4}. This set is the final output of this system.
[0131] The above embodiments illustrate preferred embodiments of the present invention. Any equivalent adjustments to the technical solution based on software engineering methods are within the scope of protection, including but not limited to: implementing algorithm logic using different programming languages, refactoring functional modules into services, adjusting data interaction protocols, and optimizing resource scheduling strategies. Any implementation scheme derived from reasonable modifications to the data processing flow, service call chain, or system architecture layer without departing from the core technology of the present invention should be considered within the scope of protection defined by the claims of the present invention.
Claims
1. A machine learning based immune signature analysis system characterized in that, The system comprises: An entropy density rheological module for processing immune marker three-dimensional data by a kernel density estimation algorithm, calculating spatial neighborhood information entropy, generating a marker entropy density gradient map using a sliding average method, and transmitting the marker entropy density gradient map to a boundary dynamic labeling module; The boundary dynamic labeling module is used for calculating directional derivatives based on the marker entropy density gradient map, detecting 8-neighborhood gradient mutation points, and generating an immune marker boundary topology map using a morphological closing operation of a dynamic structure element, and transmitting the immune marker boundary topology map to a collaborative tensor modeling module; The collaborative tensor modeling module is used for cropping low-density data according to the immune marker boundary topology map, constructing an immune subtype-time slice-marker third-order tensor, extracting a core factor matrix using a dimension-adaptive Tucker decomposition algorithm, calculating Manhattan distance to generate a collaborative expression intensity value, and transmitting the collaborative expression intensity value and a decomposition residual matrix to a marker discrimination module; The marker discrimination module is used for performing Hadamard product operation on the collaborative expression intensity value and the decomposition residual matrix, inputting a support vector machine classifier for discrimination, and generating a high-specificity immune marker set; The entropy density gradient map specifically includes spatial entropy distribution characteristics, a gradient vector field, and a density contour line, the boundary topology map contains boundary curvature parameters, a topological connected domain, and a morphological closed contour, the collaborative expression intensity value covers core factor load, a time evolution mode, and a Manhattan distance matrix, the decomposition residual matrix includes an orthogonal residual component, a tensor reconstruction error, and a low-rank approximation deviation, and the high-specificity immune marker set specifically refers to a classification decision surface, a feature weight vector, and a sample confidence score; a radius r of the dynamic structuring element and a standard deviation of the gradient vector field module length satisfies , is obtained by calculating the standard deviation of all module length values of the gradient vector field wherein, represents the radius of the morphological closing structuring element, represents the standard deviation of all the modulus values in the gradient vector field, denotes the logarithm function with base 2, denotes the ceiling operation; The interval Δd of the density contour line is obtained by calculating the moving average value of the Euclidean distance between the centers of adjacent contour lines, and the moving average window size is the square root of the number of three-dimensional data voxels. The calculation of said directional derivative employs a modified Sobel-Feldman operator whose convolution kernel size is dynamically adapted to the average spacing of the density contour lines : ; wherein is the average Euclidean distance of the density contours in the gradient vector field, obtained by averaging the L2 norm of the difference in the coordinates of the mass centers of adjacent contours; the radius of the structuring element of the morphological closing operation satisfies ; wherein, represents the radius of the morphological closing structuring element, represents the set of all modulus values in the gradient vector field, represents the arithmetic mean of the modulus values of the gradient vector field, represents the maximum value in the set of modulus values of the gradient vector field, denotes the ceiling operation; The core tensor dimension setting in the Tucker decomposition algorithm follows: the first dimension ; Second dimension ; Third dimension ; The Manhattan distance matrix is calculated using a sliding window mechanism, the window width The rank of the core factor matrix satisfies , is selected by cross-validation method in the preset interval [3, 8]. wherein, represents the first dimension size of the core tensor, represents the total number of classes of immune subtype classification, denotes a logarithm function with base 2, denotes a ceiling operation, represents the second dimension size of the core tensor, represents the number of time slices, denotes a floor operation, represents the third dimension size of the core tensor, represents the total number of initial markers; Kernel function parameters of the support vector machine classifier Dynamic adaptation of the Frobenius norm of the decomposition residual matrix : ; updating step of the feature weight vector with the low rank approximation bias is negatively correlated, satisfying , is the original three-order tensor, and both use normalized dimensionless values; wherein, a kernel function parameter representing a support vector machine classifier, a total number of initial markers, a variance of all elements of the decomposition residual matrix, a Frobenius norm of the decomposition residual matrix, a function for ensuring that the denominator is not zero, an update step size of the feature weight vector, a numerical value of the low-rank approximation bias, a Frobenius norm of the original third-order tensor.
2. The machine learning based immune signature analysis system of claim 1, wherein, The extraction process of the orthogonal residual component includes: Performing Gram-Schmidt orthogonalization on the core factor matrix to generate a standard orthogonal basis vector set; Calculating the projection component of the decomposition residual matrix on each orthogonal basis vector; Retaining components of the projection coefficients that exceed a threshold wherein is the i-th singular value of the core factor matrix, obtained by singular value decomposition; Threshold value The calculation result of the threshold value is kept to three significant digits after the decimal point. wherein, a screening threshold representing a projection coefficient, a first dimension size of a core tensor, an i-th singular value obtained by singular value decomposition of the core factor matrix, is an integer index from 1 to .
3. The machine learning based immune signature analysis system of claim 2, wherein, The determination of the topological connected domain uses an improved 8-neighborhood connectivity algorithm, and the merging threshold is set to: ; The area screening condition of the shape-closed contour is , , ; wherein, a threshold value representative of connected component merging, an arithmetic mean representative of all values of the border curvature parameter, a standard deviation representative of all values of the border curvature parameter, a total number representative of initial connected components, denotes the natural exponential function, an area representative of the morphologically closed contour, denotes the mathematical constant pi, denotes the minimum structuring element radius employed by the morphological closing operation and takes the value 1, denotes the maximum structuring element radius employed by the morphological closing operation and is obtained by rounding up to the nearest integer after dividing by 2, denotes the average Euclidean distance between adjacent density contour centroids.
4. The machine learning based immune signature analysis system of claim 3, wherein, The extraction of the time evolution mode uses a constrained dynamic time warping algorithm, and the slope limit of the curved path is: ; The calculation of the sample confidence score incorporates the co-expression intensity values and the decomposition residual matrix : ; wherein, represents an index increment in the direction of time series i, represents an index increment in the direction of time series j, represents the number of time slices, denotes a logarithm function with base 2, represents a three-dimensional matrix of co-expression intensity values and has dimensions , represents the number of immune subtypes, represents the number of initial markers, represents a three-dimensional matrix of decomposition residual matrices and has the same dimensions as , represents a sample confidence score, denotes the L1 norm of a matrix, i.e. the sum of the absolute values of all elements, the function returns the larger of the two, denotes the Hadamard product operation.
5. The machine learning based immune signature analysis system of claim 4, wherein, The spatial neighborhood information entropy is calculated using an adaptive bandwidth kernel density estimation: bandwidth with the local variance of the immunological marker three-dimensional data satisfies ; The window length of the moving average method Dynamic adaptation of the number of data voxels , satisfying ; wherein, a bandwidth parameter representative of kernel density estimation, a variance representative of data within a 5x5x5 cube neighborhood centered at the current voxel, a total sample size representative of the immunological marker three-dimensional data, representative of -1 / 5th power of a length of a sliding average window, a total number of data voxels.
6. The machine learning based immune signature analysis system of claim 5, wherein, The kernel density estimation algorithm uses an improved Epanechnikov kernel function: ; Streamline tracking step size for gradient vector fields With three-dimensional data each axial resolution Satisfies , All converted to millimeter units involved in the calculation; wherein, a kernel function representing a kernel density estimation, a normalized distance parameter and , is a current voxel coordinate, is a neighborhood voxel coordinate, is a step size representing a streamline tracking, is a resolution of the three-dimensional data in the x-axis direction, is a resolution of the three-dimensional data in the y-axis direction, is a resolution of the three-dimensional data in the z-axis direction, the function min takes the minimum value of the three resolutions.
Citation Information
Patent Citations
Method of training machine learning model for analyzing immunohistochemical staining images and computing system executing same
CN118985027A
Mathematical algorithm calculation system and optimization method based on big data
CN120371256A