Data visualization method, medium and system for stream detection

Through time segment management and parallel computing of fast dimensionality reduction equations, the problems of data processing delay and unreasonable resource scheduling in flow cytometry are solved, and efficient and real-time flow data visualization is achieved.

CN120804208APending Publication Date: 2025-10-17QINGDAO RAISECARE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511033634.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing flow cytometry has problems such as insufficient real-time performance, high computational complexity, unreasonable resource scheduling, lack of utilization of data timing characteristics, and static configuration of visualization parameters that cannot adapt to dynamic changes when processing real-time streaming data, resulting in slow system response and poor visualization effect.

Method used

It adopts a time-slice-based data flow management mechanism, combined with a data flow cache queue, performs parallel computing through a fast dimensionality reduction equation group, builds visual mapping rules and performs resource allocation to achieve real-time data processing and visualization.

Benefits of technology

It improves data processing efficiency, reduces calculation time, optimizes resource utilization, realizes adaptive adjustment and smooth transition of visualization results, and ensures real-time response of the system and clear and intuitive visualization effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804208A_ABST
    Figure CN120804208A_ABST
Patent Text Reader

Abstract

The invention provides a data visualization method for stream detection, a medium and a system, and belongs to the technical field of data visualization for stream detection.The method comprises the steps that firstly, the length of a data stream collection time slice is formulated, a cache queue is established, and multi-dimensional data streams are collected from the queue and subjected to feature extraction; performing parallel computing processing on the multi-dimensional feature data, dividing the multi-dimensional feature data into a plurality of data blocks, and performing parallel dimension reduction by using a graphics processor; constructing a visual mapping rule for the feature data after dimension reduction, establishing a rendering buffer area, and performing batch processing by adopting a resource allocation equation set; establishing a data updating queue in the memory partition and performing real-time dimension reduction processing; calculating visual expression accuracy based on the feature data after dimension reduction and optimizing a visual mapping rule; and finally, the rendering parameters are updated by adopting the optimized rule, and real-time visual rendering of the data stream is completed in the graphics processor, so that the problem that massive streaming data are difficult to process in real time and are difficult to display visually in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of data visualization for flow detection, and in particular, relates to a data visualization method, medium and system for flow detection. BACKGROUND

[0002] Flow cytometry is an indispensable analytical tool in modern biomedical research and clinical diagnosis, which can produce thousands to tens of thousands of multi-parameter data points per second. With the increase of detection channels and the improvement of sampling frequency, the data generated by flow cytometry presents the characteristics of high dimension, large volume and rapid flow. Traditional flow data visualization methods mainly use static display methods such as two-dimensional scatter plots and histograms. This method faces many challenges when dealing with real-time data streams.

[0003] In the field of biomedical research, researchers need to monitor the dynamic changes of cell populations in real time, including cell number, morphological characteristics, fluorescence intensity and other multi-dimensional information. The existing flow data visualization system mainly exists in the following application scenarios: (1) real-time analysis of patient blood samples in clinical diagnosis, which requires rapid identification of abnormal cell populations; (2) dynamic monitoring of cell response in drug development process, which requires the system to capture small changes in cell state; (3) classification analysis of immune cell subpopulations in immunology research, which requires simultaneous processing of multiple surface marker expression data; (4) tracking of cell differentiation process in stem cell research, which requires continuous recording of the gradual change process of cell characteristics.

[0004] However, the existing technology has significant limitations when dealing with real-time flow data: first, the traditional batch processing mode cannot meet the real-time requirements, and the data processing delay causes important biological events to be missed; second, the commonly used dimensionality reduction algorithms (such as PCA, t-SNE) have high computational complexity, making it difficult to deal with high-speed data streams; third, the resource scheduling in the visualization rendering process is unreasonable, causing system response delay or display lag; fourth, the existing system lacks sufficient utilization of data timing characteristics, reducing the biological interpretability of the visualization results; fifth, the static configuration method of visualization parameters cannot adapt to the dynamic changes of data distribution, affecting the display effect of data characteristics. Especially in high-throughput screening and clinical diagnosis scenarios with high time efficiency requirements, the shortcomings of existing technology are more prominent. When the system needs to process data from multiple fluorescence channels and display the dynamic changes of cell populations in real time, traditional methods often encounter processing bottlenecks and cannot provide timely and accurate visualization feedback to researchers. These problems seriously restrict the application of flow cytometry in real-time monitoring and rapid diagnosis fields. SUMMARY

[0005] Therefore, the application provides a data visualization method, medium and system for flow detection, which can solve the problem that the prior art cannot process massive flow data in real time and visualize display.

[0006] The application is implemented as follows:

[0007] The first aspect of the application provides a data visualization method for flow detection, comprising the following steps.

[0008] S10, the data stream collection time slice length is set as a preset time length, and a data stream cache queue is established based on the preset time length;

[0009] S20, multi-dimensional data stream is collected from the data stream cache queue, feature extraction is performed on the multi-dimensional data stream, and multi-dimensional feature data is obtained;

[0010] S30, a fast dimension reduction equation set is constructed to perform parallel calculation and processing on the multi-dimensional feature data, and reduced feature data is generated;

[0011] S40, the multi-dimensional feature data is divided into multiple data blocks, and a graphics processing unit is used for parallel dimension reduction calculation;

[0012] S50, a visual mapping rule is constructed for the reduced feature data, and a rendering buffer is established in a memory partition based on the visual mapping rule;

[0013] S60, a resource allocation equation set is used to process the data in the rendering buffer in batches;

[0014] S70, a data update queue is established in the memory partition, new data stream is written into the data update queue, and real-time dimension reduction processing is performed on the data in the data update queue;

[0015] S80, the visual expression accuracy is calculated based on the reduced feature data, and the visual mapping rule is optimized online according to the visual expression accuracy;

[0016] S90, the rendering parameter is updated by using the optimized visual mapping rule, and data stream real-time visualization rendering is completed in the graphics processing unit.

[0017] The step S10 specifically comprises: step 101, determining a data stream collection time slice length according to a system memory capacity and a data stream rate; step 102, calculating a cache queue size based on the data stream collection time slice length, the cache queue size being determined by the data stream rate, a data dimension number and a buffer coefficient; step 103, establishing a data stream cache queue according to the cache queue size, for temporarily storing data streams; step 104, dynamically monitoring the data stream cache queue, and automatically expanding when a data volume exceeds a preset threshold; and step 105, adjusting the data stream collection time slice length in real time through a data stream rate monitoring module.

[0018] The step S20 specifically comprises: step 201, reading multi-dimensional data streams from the data stream cache queue; step 202, performing linear transformation on the multi-dimensional data streams by using a feature extraction matrix; step 203, performing normalization processing on the feature data after linear transformation; step 204, calculating statistical features of the feature data, including mean and standard deviation; and step 205, performing standardization processing on the feature data according to the statistical features, to obtain multi-dimensional feature data.

[0019] Further, the step S30 specifically comprises: step 301, calculating a conditional probability distribution between data points in the multi-dimensional feature data; step 302, constructing a high-dimensional space distance matrix according to the conditional probability distribution; step 303, establishing a mapping weight matrix based on the high-dimensional space distance matrix; step 304, constructing constraint conditions by using a local preservation term, a topological preservation term and a global preservation term; and step 305, performing dimension reduction processing on the multi-dimensional feature data by parallel computing, to obtain reduced-dimensional feature data.

[0020] Further, the step S40 specifically comprises: step 401, determining a data block size according to a graphic processor video memory capacity and a parallel computing thread number; step 402, dividing the multi-dimensional feature data into a plurality of data blocks according to the data block size; step 403, allocating an independent computing thread to each data block; step 404, establishing a parallel computing task queue in a graphic processor; and step 405, performing parallel dimension reduction computing processing on the data blocks.

[0021] Further, the step S50 specifically comprises: step 501, calculating a rendering buffer size based on a video memory capacity and a buffer area number; step 502, establishing a plurality of rendering buffers in a video memory partition; step 503, allocating an independent rendering task to each rendering buffer; step 504, constructing visual mapping rules including position mapping rules, color mapping rules, size mapping rules and shape mapping rules; and step 505, applying the visual mapping rules to the rendering buffers.

[0022] The step S60 specifically comprises: step 601, calculating the total memory capacity and the number of memory partitions; step 602, determining the upper limit of the capacity of each memory partition; step 603, establishing a resource allocation equation set; step 604, performing batch processing on the data in the rendering buffer according to the resource allocation equation set; and step 605, performing rendering calculation tasks on each batch of data.

[0023] Further, the step S70 specifically comprises: step 701, calculating the data update queue capacity; step 702, establishing a data update queue in the memory partition; step 703, writing the new data stream to the data update queue; step 704, performing dimension reduction processing on the data in the data update queue in a batch processing manner; and step 705, updating the data after dimension reduction to the rendering buffer.

[0024] Further, the step S80 specifically comprises: step 801, calculating the position accuracy index, the color accuracy index, the size accuracy index, and the shape accuracy index; step 802, calculating the visual expression accuracy according to the weight coefficients of the accuracy indexes; step 803, calculating the gradient of the visual expression accuracy to the mapping rule; step 804, updating the visual mapping rule using an exponentially decaying learning rate; and step 805, verifying the effectiveness of the updated visual mapping rule.

[0025] Further, the step S90 specifically comprises: step 901, applying the optimized visual mapping rule to the rendering parameters; step 902, establishing a rendering task queue in the graphics processor; step 903, performing real-time rendering on the data in the rendering task queue; step 904, monitoring the rendering performance indicators; and step 905, dynamically adjusting the rendering parameters according to the performance indicators.

[0026] Further, the fast dimension reduction equation set comprises a similarity distribution equation, a distance mapping equation, and an error optimization equation.

[0027] The similarity distribution equation is used to calculate the distribution relationship of data points in a high-dimensional space, and the inputs include a feature vector matrix, a distance weight matrix, a similarity threshold, a data point sampling weight coefficient, a Gaussian kernel function parameter, a local neighborhood size parameter, a feature importance weight coefficient, and a time series correlation coefficient, and the output is a conditional probability distribution matrix between data points.

[0028] The distance mapping equation is used to construct the mapping relationship from a high-dimensional space to a low-dimensional space, and the inputs include a conditional probability distribution matrix, a mapping weight matrix, a projection dimension parameter, a neighborhood preservation coefficient, a topological structure preservation coefficient, a local linear preservation coefficient, a global structure preservation parameter, a feature space deformation coefficient, and a distance metric matrix, and the output is a feature vector matrix after dimension reduction.

[0029] The error optimization equation is used to minimize the information loss in the dimensionality reduction process. The input includes the actual dimensionality reduction result matrix, the ideal dimensionality reduction result matrix, the gradient learning rate, the number of optimization iterations, the convergence threshold parameter, the regularization coefficient, the error weight matrix, the local loss function parameters, the global loss function parameters and the dynamic balance coefficient. The output is the optimal mapping parameter matrix.

[0030] The resource allocation equation group includes a memory partition equation and a rendering partition equation;

[0031] The memory partition equation is used to calculate the data block allocation strategy. The input includes the system memory capacity, data block size and data update frequency. The output is the number of memory partitions and the upper limit of the data capacity of each partition.

[0032] The rendering partition equation is used to calculate the video memory allocation strategy. The input includes the graphics processor video memory capacity, rendering frame rate requirements and data complexity. The output is the number of video memory partitions and the rendering task allocation plan.

[0033] The equations or formulas involved in the present invention are described in detail below:

[0034] 1. The similarity distribution equation of the fast dimensionality reduction equation group is specifically expressed as follows:

[0035]

[0036] Where, P ij is the conditional probability between data points i and j; x i ,x j is the eigenvector in the high-dimensional space; w d is the distance weight matrix; σ is the Gaussian kernel function parameter, and its value range is (0,∞); α s is the data point sampling weight coefficient, ranging from [0,1]; β t is the time series correlation coefficient, and its value range is [0,1].

[0037] Parameter acquisition method:

[0038] 1)w d Obtained through feature importance analysis, the calculation method is:

[0039] Where, f i is the value of the i-th feature; is the average value of all features; f ij is the i-th eigenvalue of the j-th sample; n is the number of features; m is the number of samples.

[0040] 2) σ is determined by binary search method, so that: ∑ j≠i P ij=log(k);

[0041] Where k is the desired neighborhood size.

[0042] 2. The distance mapping equation is specifically expressed as follows:

[0043] Y=W·X+λ1L p +λ2L t +λ3L g +∈;

[0044] Where Y is the feature matrix after dimensionality reduction; W is the mapping weight matrix; X is the original feature matrix; L p is the local preservation term; L t is a topology preserving term; L g is the global maintenance term; λ1, λ2, λ3 are balance coefficients; ∈ is the error term.

[0045] in:

[0046]

[0047] Where N(i) represents the neighborhood set of point i; M ij is the neighborhood relationship matrix; D ij is the high-dimensional space distance; d ij is the low-dimensional space distance; Q ij is a low-dimensional probability distribution.

[0048] 3. The error optimization equation is specifically expressed as follows:

[0049]

[0050] Where Y * is the ideal dimension reduction result matrix; ||·|| F is the Frobenius norm; α is the regularization coefficient; β, γ are the local and global loss function parameters; L local ,L global are the local and global loss functions respectively.

[0051] 4. The memory partition equation in the resource allocation equation group is specifically expressed as follows:

[0052]

[0053] Where N m is the number of memory partitions; M total is the total system memory; S block is the data block size; η is the reserved system overhead ratio; C i is the upper limit of the capacity of the i-th partition; δ is the capacity reduction coefficient.

[0054] 5. The rendering partition equation is specifically expressed as follows:

[0055]

[0056] In the formula, N v is the number of memory partitions; V total is the total memory capacity; R fps is the target frame rate; D complexity is the data complexity coefficient; N max is the maximum partition number limit; T i is the task amount of the ith partition; and θ is the task increment coefficient.

[0057] Original explanation:

[0058] 1. The similarity distribution equation is constructed based on a Gaussian kernel function, adopts an exponential form to effectively capture the nonlinear relationship between data points, and the introduction of a weight term considers the importance difference of features, and a time correlation coefficient reflects the time continuity feature of data.

[0059] 2. The distance mapping equation is based on linear mapping, and three regular terms are used to maintain local structure, topological relationship and global distribution respectively, so that the balanced maintenance of multi-scale features is realized.

[0060] 3. The error optimization equation comprehensively considers the deviation between the actual dimensionality reduction result and the ideal result, the complexity of the mapping matrix, and the local and global structure maintenance performance, and the error size is measured by Frobenius norm;

[0061] 4. The memory partition equation considers the system overhead and the unevenness of data distribution, and uses a decreasing coefficient to realize dynamic balance.

[0062] 5. The rendering partition equation is based on the memory capacity and performance demand, and realizes the dynamic adjustment of the calculation load through the task increment coefficient.

[0063] In addition, the present application also relates to the following some calculation processes:

[0064] 1. The data flow collection time slice length calculation formula is:

[0065]

[0066] In the formula, T segment is the time slice length; M buffer is the cache queue size; R data is the data flow rate; N dim is the data dimension; and θ is a buffer coefficient, and the value range is (0, 1).

[0067] 2. The feature calculation formula of the feature extraction process is:

[0068] F = W f · X + b f ;

[0069]

[0070] where F is the extracted feature; W is the feature extraction weight matrix; X is the original data; b is the bias term; F is the normalized feature; μ is the feature mean; σ is the feature standard deviation. f f norm

[0071] 3. Data block division strategy calculation formula:

[0072] where S is the data block size; M is the GPU memory capacity; N is the number of parallel blocks; D is the total data volume; N is the number of GPU threads. block gpu block total thread

[0073] 4. Render buffer allocation calculation formula:

[0074] where B is the buffer size; V is the memory capacity; α is the memory utilization rate, with a value range of (0, 1); N is the number of buffers; β is the buffer overhead coefficient. size mem buffer

[0075] 5. Data update queue capacity calculation formula:

[0076] Q = max(R · T, D · N); size update delay batch batch

[0077] where Q is the queue capacity; R is the data update rate; T is the maximum processing delay; D is the batch data volume; N is the batch number. size update delay batch batch

[0078] 6. Visual expression accuracy calculation formula:

[0079] A = α1A + α2A + α3A + α4S; visual position color size shape ;​​​​​​​​​​​​​​​​​​​​​​​​​

[0080] Where A visyal A is the accuracy of visual expression; position ,A color ,A size ,A shape are the accuracy indicators of position, color, size, and shape respectively; α1, α2, α3, and α4 are weight coefficients, and they satisfy

[0081] The accuracy indicators are calculated as follows:

[0082]

[0083] Where, p i ,c i ,s i They are the actual position, color, and size values ​​respectively; are the ideal position, color, and size values ​​respectively; d max ,c max ,s max The maximum difference values ​​of position, color and size respectively; IoU i is the intersection-and-union ratio of the shapes.

[0084] 7. Visual mapping rule optimization update formula:

[0085]

[0086] η=η0·exp(-λ·t);

[0087] Where, R new ,R old are the mapping rules before and after updating respectively; η is the learning rate; η0 is the initial learning rate; λ is the decay coefficient; t is the number of iterations; is the gradient of accuracy with respect to the mapping rule.

[0088] Principle explanation:

[0089] 1. The time segment length calculation takes into account the data flow rate and dimension, and introduces a buffer factor to cope with data bursts;

[0090] 2. Feature extraction uses linear transformation and normalization to ensure feature comparability;

[0091] 3. Data block division is based on the dual constraints of GPU resources and data volume to achieve optimal resource utilization;

[0092] 4. Render buffer allocation takes into account video memory utilization efficiency and buffer overhead;

[0093] 5. Update queue capacity design takes into account both real-time performance and batch processing efficiency;

[0094] 6. The visual expression accuracy is comprehensively evaluated by multi-dimensional indicators, and normalization processing is adopted to make the indicators comparable;

[0095] 7. The mapping rule optimization adopts gradient-based iterative update, and introduces an exponentially decaying learning rate to ensure convergence.

[0096] These calculation steps are interrelated and collectively constitute a complete streaming data visualization system, in which: 1) the time segment length affects the time granularity of feature extraction;2) the feature extraction result determines the data block division strategy;3) the buffer allocation affects the rendering performance;the accuracy evaluation and optimization ensure the dynamic optimization of the visualization effect.

[0097] The second aspect of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores program instructions, and the program instructions are used to execute the above-mentioned data visualization method for streaming detection when running in a computer.

[0098] The third aspect of the present application provides a data visualization system for streaming detection, which comprises the above-mentioned computer readable storage medium.

[0099] Compared with the prior art, the data visualization method, medium and system for streaming detection provided by the present application have the following beneficial effects:

[0100] Firstly, in terms of data processing efficiency, the present application adopts a time segment-based data stream management mechanism, combined with the design of the data stream buffer queue, to realize smooth reception and processing of data. Secondly, in terms of dimension reduction calculation performance, the fast dimension reduction equation set designed by the present application significantly improves the dimension reduction efficiency through parallel computing architecture. Through the synergistic effect of the similarity distribution equation, the distance mapping equation and the error optimization equation, the introduction of the time correlation coefficient is more suitable for the characteristics of streaming data. Compared with the existing principal component analysis dimension reduction or t-SNE or UMAP, it is more targeted for streaming data, dynamically updates the distance weight matrix, reduces repeated calculation, and quickly reduces the calculation time. Thirdly, in terms of resource utilization efficiency, the present application realizes the dynamic optimization allocation of computing resources through the resource allocation equation set. Fourthly, in terms of visualization effect, the present application realizes the adaptive adjustment of the visualization result through the dynamic optimization of the visual mapping rule. By calculating the visual expression accuracy in real time, the system can automatically adjust the display parameters according to the changes of data characteristics, so that the visualization result is more clear and intuitive. Fifthly, in terms of system stability, the present application realizes the smooth transition of data processing and display through the collaborative design of the data update queue and the rendering buffer.

[0101] In summary, the present application solves the problem that the prior art cannot process massive streaming data in real time and visualize the display. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] Figure 1 A flow chart of the method provided by the present invention;

[0103] Figure 2 This is a data flow collection and analysis diagram in an embodiment, including two sub-graphs, the upper sub-graph showing the fluctuation of data flow rate over time, and the lower sub-graph showing the dynamic changes of cache usage;

[0104] Figure 3 1 is a diagram of the feature extraction process in the embodiment, including three sub-graphs, which are schematic diagrams of original data, data after feature transformation, and normalized data respectively;

[0105] Figure 4 This is a GPU performance analysis chart in an embodiment, including two sub-charts. The left sub-chart shows how the calculation time changes with the data block size, and the right sub-chart shows the usage of video memory.

[0106] Figure 5 A render buffer allocation map for an embodiment;

[0107] Figure 6 It is a time series analysis diagram of subgroups in the embodiment;

[0108] Figure 7 is a feature correlation diagram in the embodiment;

[0109] Figure 8 is a density distribution diagram in the embodiment;

[0110] Figure 9 This is a feature projection diagram in the embodiment, including 4 sub-diagrams. The upper three sub-diagrams show the 2D projections between the three dimensions, and the lower sub-diagram shows the complete 3D projection.

[0111] Figure 10 It is a state transition diagram in the embodiment;

[0112] Figure 11 Graph showing immune responses in the Examples. DETAILED DESCRIPTION

[0113] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0114] like Figure 1 FIG. 1 is a flow chart of a data visualization method for flow cytometry provided by the present invention. The specific implementation of each step is described in detail below:

[0115] The specific implementation of step S10 is to set the data stream collection time slice length as a preset time length, and establish a data stream cache queue based on the preset time length.

[0116] First, according to the system memory capacity M total and the data stream rate R data , the data stream collection time slice length T segment can be determined, and the specific formula is as follows: Where M buffer is the cache queue size, N dim is the data dimension, and θ is the buffer coefficient, which is in the range of (0, 1). The time slice length thus determined takes into account the system memory capacity, as well as the data stream rate and dimension, ensuring that the cache queue does not overflow prematurely.

[0117] Next, based on the determined T segment , the cache queue size M buffer can be calculated. The cache queue size is determined by R data , N dim and θ, and can be represented as M buffer = R data · N dim · T segment · (1+θ). Setting the cache queue size in this way can both temporarily store enough data for subsequent processing and dynamically adjust θ according to system resources to cope with data bursts.

[0118] Then, according to the calculated M buffer , a data stream cache queue of the corresponding size is established in the system memory for temporarily storing the collected data stream.

[0119] In actual operation, dynamic monitoring of the data stream cache queue is also required, and when the data volume exceeds the preset threshold, automatic expansion operation is performed. At the same time, R data can be monitored in real time to dynamically adjust T segment to ensure that the cache queue can effectively store the data stream.

[0120] In summary, the main purpose of step S10 is to establish a data stream cache queue of reasonable size, laying the foundation for subsequent feature extraction and visualization processing.

[0121] The specific implementation of step S20 is to collect multi-dimensional data stream from the data stream cache queue, perform feature extraction on the multi-dimensional data stream, and obtain multi-dimensional feature data.

[0122] First, read the multi-dimensional data X from the data stream cache queue. Then, use the feature extraction matrix W fLinearly transforming the multi-dimensional data stream to obtain preliminary feature data F = W f · X + b f , where b f is a bias term. Next, the linearly transformed feature data is normalized, i.e. , where μ is the feature mean and σ is the feature standard deviation. The purpose of this is to make features of different dimensions comparable.

[0123] Next, the statistical features of the extracted feature data, including the mean μ and the standard deviation σ, are calculated. Finally, the feature data is standardized according to the calculated statistical features to obtain the final multi-dimensional feature data F norm . The purpose of this step is to eliminate dimensional differences, making each dimension of the feature have the same dimension and distribution range, laying the foundation for subsequent dimensionality reduction calculations.

[0124] The specific implementation of step S30 is to construct a fast dimensionality reduction equation system to perform parallel calculation and processing on the multi-dimensional feature data, generating reduced dimension feature data.

[0125] First, the conditional probability distribution P ii between each data point in the multi-dimensional feature data is calculated. The similarity distribution equation can be used: , where x i , x j is a high-dimensional feature vector, w d is a distance weight matrix, σ is a Gaussian kernel function parameter, α s is a data point sampling weight coefficient, and β t is a time series correlation coefficient. This step aims to capture the nonlinear relationship between data points.

[0126] Then, based on the calculated P ij , a high-dimensional space distance matrix is constructed, and a mapping weight matrix W is established based on this. Next, a local preservation term L p , a topological preservation term L t , and a global preservation term L g are used to construct constraint conditions, where: Through parallel calculation, the multi-dimensional feature data is reduced to obtain reduced dimension feature data Y, which is expressed as: Y = W·X + λ1L p + λ2L t + λ3L g + ∈. Where λ1, λ2, λ3 are balance coefficients, and ∈ is an error term.

[0127] The main purpose of this step is to map high-dimensional features to low-dimensional space while maintaining local structure, topological relationship and global distribution. The similarity distribution equation, distance mapping equation and error optimization equation used can effectively capture the intrinsic structural characteristics of the data, laying a good foundation for subsequent visualization processing.

[0128] The specific implementation of step S40 is to divide the multi-dimensional feature data into multiple data blocks and use a graphics processor to perform parallel dimension reduction calculation.

[0129] First, according to the graphics processor memory capacity M gpu and the number of parallel computing threads N thread , determine the size S block of the data block, which can be calculated using the formula: where N block is the number of parallel blocks, and D total is the total data volume. This setting of data block size takes into account the limitations of GPU resources and also considers the overall data volume, ensuring optimal use of resources.

[0130] Next, the multi-dimensional feature data is divided into multiple data blocks according to S block . Each data block is assigned an independent computing thread, and a parallel computing task queue is established in the graphics processor. Finally, the data blocks are subjected to parallel dimension reduction calculation and processing.

[0131] The main purpose of this step is to utilize the powerful parallel computing capabilities of the graphics processor to significantly improve the efficiency of dimension reduction calculation. Through a reasonable data block division strategy, the GPU resources are fully utilized, and the computing tasks are evenly distributed, thereby achieving efficient dimension reduction of streaming data.

[0132] The specific implementation of step S50 is to construct a visual mapping rule for the dimension-reduced feature data and establish a rendering buffer in the memory partition based on the visual mapping rule.

[0133] First, according to the memory capacity V mem and the number of buffers N buffer , the size B size of the rendering buffer can be calculated using the formula: where α is the memory utilization rate and β is the buffer overhead coefficient. This setting of rendering buffer size takes into account the memory capacity and also considers the buffer overhead, ensuring reasonable allocation of resources.

[0134] Next, multiple rendering buffers are established in the video memory partition, and each buffer is assigned an independent rendering task. Then, visual mapping rules R including position mapping rules, color mapping rules, size mapping rules and shape mapping rules are constructed and applied to the rendering buffers. These mapping rules convert the reduced dimension feature data into geometric properties required for visualization expression, laying a foundation for subsequent rendering processing.

[0135] In general, the main purpose of step S50 is to map the reduced dimension feature data to geometric properties required for visualization expression, and to establish corresponding rendering buffers in the video memory to support real-time rendering processing. Through a reasonable buffer allocation strategy, the system's video memory resources are maximized.

[0136] The specific implementation of step S60 is to perform batch processing on the data in the rendering buffer using a resource allocation equation set.

[0137] First, the total system memory capacity M total and the number of memory partitions N m are calculated to determine the upper limit C i of the capacity of each memory partition. The specific formula is: and where η is the reserved system overhead ratio and δ is the capacity reduction coefficient. This memory partition strategy takes into account the total system memory and also considers the unevenness of data distribution, ensuring the rational use of memory resources.

[0138] Next, a resource allocation equation set is established, including a memory partition equation and a rendering partition equation. The memory partition equation is used to calculate the allocation strategy of data blocks, and the rendering partition equation is used to calculate the allocation scheme of the video memory. The specific expressions are: and where N v is the number of video memory partitions, V total is the total video memory capacity, R fps is the target frame rate, D complexity is the data complexity coefficient, N max is the maximum partition number limit, T i is the task amount of the i-th partition, and θ is the task increment coefficient.

[0139] According to the constructed resource allocation equation set, batch processing is performed on the data in the rendering buffer, and rendering calculation tasks are executed for each batch of data.

[0140] The main purpose of this step is to reasonably allocate system memory and video memory resources to ensure real-time rendering processing of streaming data. Through the design of the resource allocation equation set, dynamic balance scheduling of memory and video memory is achieved, improving the overall system performance.

[0141] The specific implementation of step S70 is: establishing a data update queue in the memory partition, writing the new data stream into the data update queue, and performing real-time dimension reduction processing on the data in the data update queue.

[0142] Firstly, the capacity Q of the data update queue can be calculated size , and the formula is: Q size = max(R update ·T delay , D batch ·N batch ). Wherein, R update is the data update rate, T delay is the maximum processing delay, D batch is the batch data volume, and N batch is the batch number. By setting the update queue capacity, both real-time requirements and batch processing efficiency are considered, and timely processing of data is ensured.

[0143] Next, the data update queue is established in the memory partition, and the new data stream is written into it. Then, the data in the data update queue is processed by dimension reduction in a batch manner, and the processing result is updated to the rendering buffer.

[0144] In general, the main role of step S70 is to maintain the data update queue, timely incorporate the new data stream into the visualization processing flow, and ensure that the entire system can respond to data changes in real time and provide the latest visualization effect.

[0145] The specific implementation of step S80 is: calculating the visual expression accuracy based on the reduced feature data, and online optimizing the visual mapping rule according to the visual expression accuracy.

[0146] Firstly, the position accuracy index A position , the color accuracy index A color , the size accuracy index A size and the shape accuracy index A shape are calculated. The specific formula is as follows: Wherein, p i , c i , s i are the actual position, color and size values respectively; are the ideal position, color and size values respectively; d max , c max , s max are the maximum difference values of position, color and size respectively; and IoU i is the shape intersection over union.

[0147] Next, according to the weight coefficient α of each accuracy index i , the visual expression accuracy A is calculated visual = α1A position + α2A color + α3A size + α4A shape , wherein

[0148] Then, the gradient of the visual expression accuracy A visual to the mapping rule R is calculated The visual mapping rule is updated using an exponentially decaying learning rate η = η0·exp(-λ·t) Finally, the effectiveness of the updated visual mapping rule is verified.

[0149] Overall, the main purpose of step S80 is to establish a visual expression accuracy evaluation system and optimize the visual mapping rule online based on it to ensure that the visualization effect continuously meets user needs. Through multi-dimensional accuracy indicators and gradient descent update strategies, dynamic improvement of the visualization effect is achieved.

[0150] The specific implementation of step S90 is to update the rendering parameters using the optimized visual mapping rule and complete real-time visualization rendering of the data stream in the graphics processor.

[0151] First, the optimized visual mapping rule R new is applied to the rendering parameters. Then, a rendering task queue is established in the graphics processor, and real-time rendering is performed on the data in the rendering task queue.

[0152] During the rendering process, it is necessary to monitor rendering performance indicators such as frame rate R fps , delay, etc. If it is found that the performance indicators exceed expectations, the rendering parameters such as projection mode, material properties, etc. can be dynamically adjusted according to the performance indicators to ensure smooth visualization effect.

[0153] Overall, the main purpose of step S90 is to use the optimized visual mapping rule in combination with the powerful rendering capabilities of the graphics processor to achieve real-time visualization display of the streaming data. Through dynamic adjustment of the rendering parameters, the real-time and smoothness of the visualization effect are ensured, meeting the user's interactive needs.

[0154] The second aspect of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores program instructions, and the program instructions are used to execute the above-mentioned data visualization method for streaming detection when running in a computer.

[0155] The third aspect of the present application provides a data visualization system for flow detection, wherein the computer readable storage medium is as described above.

[0156] Specifically, the principle of the present application is that: first, in terms of data stream management, the present application adopts a time slice-based processing strategy, and the rationality of this design is reflected in that: (1) by setting the preset time length, the uniformity of the data block size is ensured, which is conducive to subsequent parallel processing; (2) the design of the data stream buffer queue takes into account the randomness of data arrival, and by dynamically adjusting the buffer size, smooth processing of the data stream is achieved; (3) the sliding window mechanism is adopted in the feature extraction process, which not only ensures the time sequence continuity of the data, but also improves the processing efficiency. This design can theoretically prove that it can complete data preprocessing within O(n) time complexity, where n is the number of data points.

[0157] Secondly, in terms of dimensionality reduction calculation, the fast dimensionality reduction equation designed by the present application has a deep mathematical foundation: the similarity distribution equation is constructed based on the Gaussian kernel function, and its theoretical basis is the local linear hypothesis of data in high-dimensional space. By introducing the distance weight matrix and the time correlation coefficient, this equation can effectively capture the spatial structure and temporal evolution characteristics of the data. The distance mapping equation adopts a multi-objective optimization framework, and through the balance of three regularization terms, the local structure and global topology are jointly maintained. The error optimization equation is solved by the gradient descent method, and its convergence can be proved by the Lyapunov stability theory.

[0158] In terms of resource scheduling, the resource allocation equation set of the present application adopts the idea of heuristic algorithm: the memory partition equation realizes dynamic allocation of storage space through a decreasing coefficient, and its principle is based on statistical analysis of data access frequency; the rendering partition equation adopts a task increment coefficient for load balancing, and by monitoring GPU utilization in real time, the task allocation strategy is dynamically adjusted. This design ensures the optimal utilization of system resources, and theoretical analysis shows that it can improve the resource utilization rate to close to the theoretical upper limit.

[0159] In terms of visualization optimization, the present application adopts an adaptive visual mapping mechanism: by calculating the visual expression accuracy in real time, the system can evaluate the quality of the current visualization effect. This evaluation is based on the viewpoint of information theory, and by calculating the information entropy and mutual information, the degree of preservation of the original data features by the visualization result is quantified. Based on the evaluation results, the system updates the visual mapping rules in an online learning manner, and its optimization process can be expressed as a convex optimization problem, which guarantees the existence and uniqueness of the solution.

[0160] In addition, adaptive data acquisition and caching mechanism, fast and efficient dimensionality reduction calculation, and dynamic adaptive visualization rendering are also used for optimization, which are as follows:

[0161] 1. Adaptive data acquisition and caching mechanism:

[0162] To ensure the real-time and integrity of streaming data, the application first establishes an adaptive data acquisition and caching mechanism. By monitoring the system memory capacity M total and the data flow rate R data in real time, the data acquisition time slice length T segment is dynamically adjusted to meet the relationship , where M buffer is the cache queue size, N dim is the data dimension, and θ is the buffer coefficient. In this way, system resources can be fully utilized, and sudden changes in data flow can be handled. At the same time, the cache queue size M buffer is determined by R data , N dim and θ, which can be dynamically adjusted to meet the actual needs.

[0163] Through the above adaptive mechanism, the application ensures real-time acquisition and complete storage of data streams, laying a solid foundation for subsequent feature extraction and dimensionality reduction processing. At the same time, dynamic monitoring and adjustment of data acquisition parameters also ensure that the entire system can adapt to different data environments at any time, improving flexibility and robustness.

[0164] 2. Fast and efficient dimensionality reduction calculation:

[0165] To meet the needs of real-time visualization, the application uses a fast dimensionality reduction algorithm based on parallel computing. First, the similarity distribution equation is used to calculate the conditional probability distribution P ij between high-dimensional feature data x i to capture the nonlinear relationship between data points.

[0166] Next, the distance matrix in high-dimensional space is constructed based on the obtained P ij , and the mapping weight matrix W is established based on this. Then, the local preservation term L p , the topological preservation term L t and the global preservation term L g are used to construct the constraint conditions, and the multi-dimensional feature data X is processed by parallel computing to obtain the dimensionality reduced feature data Y = W·X + λ1L p + λ2L t + λ3L g + ∈. In this way, not only can the local structure, topological relationship and global distribution of the data be maintained, but also the parallel computing advantages of GPU can be fully utilized, greatly improving the efficiency of dimensionality reduction calculation.

[0167] 3. Dynamic and adaptive visualization rendering:

[0168] To ensure the real-time and interactivity of the visualization effect, the application designs a dynamic self-adaptive rendering mechanism. First, according to the video memory capacity V mem and the buffer number N buffer , the rendering buffer size S is calculated Wherein, alpha is the video memory utilization rate, and beta is the buffer overhead coefficient. A plurality of such rendering buffers are established in the video memory, and each buffer is allocated an independent rendering task.

[0169] Meanwhile, the application designs the memory partition equation and and the rendering partition equation and to realize the dynamic allocation of memory and video memory resources. In this way, not only the data processing and rendering tasks can fully utilize the system resources, but also the allocation strategy can be adjusted in real time according to the actual demand, thereby improving the overall performance.

[0170] In addition, the application also establishes a visual expression accuracy A visual evaluation system, which realizes the automatic optimization of the visualization effect by calculating multi-dimensional indexes such as position accuracy A position , color accuracy A color , size accuracy A size and shape accuracy A shape . Specifically, A visual is improved by updating the visual mapping rule R, so as to ensure the continuous improvement of the visualization effect.

[0171] The workflow of the whole system follows strict mathematical logic: from data input to final display, each link is established on the basis of reliable theory. Especially in the connection process of data processing and visualization, the system adopts a double buffering mechanism, and through an asynchronous updating strategy, the display problems such as screen tearing are effectively avoided. The stability of the system can be analyzed by the method of control theory, and it is proved that under the condition of bounded input, the system state always remains in the stable region.

[0172] In addition, the application adopts a series of optimization techniques in algorithm implementation: (1) sparse matrix storage is used to reduce memory overhead; (2) GPU acceleration is used to realize parallel computing; (3) data prefetching mechanism is used to reduce IO delay; (4) locality principle is used to optimize cache usage. The comprehensive use of these technologies ensures the efficiency and reliability of the system in practical application.

[0173] In general, the technical scheme of the application successfully solves the core problem of real-time processing of streaming data through rigorous theoretical analysis and optimized engineering implementation.

[0174] An embodiment of a specific application scenario of the present application is provided as follows: A biomedical institute is conducting a research project on immune cell function. The research team uses advanced flow cytometry to detect patient blood samples, hoping to provide a basis for disease diagnosis and treatment by analyzing the proportion and functional indicators of different immune cell subgroups. However, due to the large number of patient samples and the complex feature data of several dimensions contained in each sample, researchers face huge data processing and analysis challenges. Traditional visualization methods have been difficult to effectively present the internal structural characteristics of these high-dimensional data, and a new data visualization method is urgently needed to support the research project.

[0175] To meet the above needs, the research team decides to use the data visualization method for flow detection proposed by the present application. The specific implementation process is as follows:

[0176] 1. Data acquisition and caching

[0177] First, the research team determines the basic parameters of data acquisition. According to the performance indicators of the flow cytometer, the data volume of each blood sample is about 1GB, and the acquisition frequency is 1Hz. Considering the possibility of local data burst during the experiment, the researchers choose a larger buffer coefficient θ = 0.2. Thus, the data stream acquisition time segment length is calculated as:

[0178] In order to ensure sufficient data cache space, the researchers further calculate the cache queue size:

[0179] M buffer = R data · N dim · T segment · (1 + θ) = 1GB / s · 50 · 20.48s · 1.2 = 1228.8GB

[0180] Such a cache queue can store about 20 minutes of data stream, fully meeting the research needs.

[0181] During the actual acquisition process, the research team monitors the data stream rate R data and memory usage in real time, and finds that the sample acquisition speed increases significantly in some periods, and the cache queue capacity approaches the upper limit. Therefore, they dynamically adjust T segment and M buffer to ensure that the data stream can be completely and timely acquired and stored. Through this adaptive mechanism, the entire data acquisition process runs stably and there is no data loss.

[0182] Figure 2This is a data flow collection and analysis chart: It shows the rate changes and cache usage during the data flow collection process. The upper part shows the fluctuation of the data flow rate over time, and the lower part shows the dynamic changes in cache usage and indicates the cache limit.

[0183] 2. Feature extraction and dimensionality reduction

[0184] After completing data collection, the researchers performed feature extraction and dimensionality reduction on the data in the cache queue. First, they read 50-dimensional raw cell feature data X from the cache queue, including forward scattered light (FSC), side scattered light (SSC), and eight fluorescence signals (FL1-FL8).

[0185] Using feature extraction matrix W f Perform linear transformation on the original data to obtain preliminary 50-dimensional feature data F=W f ·X+b f The researchers normalized F and calculated the mean μ and standard deviation σ of each dimension, and finally obtained the standardized multidimensional feature data. The purpose of this step is to eliminate the dimensional differences between features of different dimensions and lay the foundation for subsequent dimensionality reduction calculations.

[0186] Next, the researchers began to perform fast dimensionality reduction calculations. First, based on F norm Calculate the conditional probability distribution P between each data point ij , the formula used is: Among them, w d is the distance weight matrix obtained based on feature importance analysis; σ = 0.1 is the Gaussian kernel function parameter; α s =0.8 is the data point sampling weight coefficient; β t =0.6 is the time series correlation coefficient.

[0187] Then, the researchers used P ij A high-dimensional space distance matrix was constructed, and based on it, a mapping weight matrix W was established. Next, they set the following dimensionality reduction constraints: the local preservation term λ1L p =0.6, topology preservation term λ2L t =0.3, global maintenance term λ3L g =0.1. Finally, through parallel calculation, F norm Perform dimensionality reduction processing to obtain 10-dimensional dimensionality reduction feature data Y=W·F norm +λ1L p +λ2L t +λ3L g +∈.

[0188] The entire dimensionality reduction process is executed on the GPU. Researchers use NVIDIA RTX 3090 graphics cards, which allocate 8 parallel computing blocks in total. The data block size S block According to the GPU memory capacity and the number of parallel threads, it is determined to be 1GB. The entire dimensionality reduction process only takes 3.2 seconds, greatly improving the processing efficiency.

[0189] Figure 3 The feature extraction process diagram: shows the three stages of feature extraction: raw data, feature-transformed data, and normalized data. The scatter plot intuitively shows the distribution changes of data at each processing stage.

[0190] Figure 4 GPU performance analysis diagram: shows the relationship between data block size and computation time and memory usage. The left figure shows the change of computation time with data block size, and the right figure shows the memory usage.

[0191] 3. Visualization rendering

[0192] With the 10-dimensional feature data Y after dimensionality reduction, researchers begin to work on visualization rendering. First, according to the memory capacity V mem = 24GB and the number of buffers N buffer = 6, the size of each rendering buffer is calculated Six such rendering buffers are established in the memory, and each buffer is assigned an independent rendering task.

[0193] Figure 5 Rendering buffer allocation diagram: shows the usage of the six rendering buffers in the form of a bar chart, and marks the upper limit of the buffer size

[0194] At the same time, researchers construct visual mapping rules R containing position mapping rules, color mapping rules, size mapping rules, and shape mapping rules, and apply them to the rendering buffer. Specifically:

[0195] Position mapping rule: map 10-dimensional feature data Y to two-dimensional plane coordinates (x, y);

[0196] Color mapping rule: map the first 3 dimensions of 10-dimensional feature data Y to RGB color space;

[0197] Size mapping rule: map the 4th dimension of 10-dimensional feature data Y to the size of the point;

[0198] Shape mapping rule: map the 5th-10th dimensions of 10-dimensional feature data Y to the shape of the point.

[0199] To allocate system resources reasonably, the researchers established a memory partition equation group and a rendering partition equation group. The memory partition equation group determined N m = 4 memory partitions, each with an upper capacity limit C i adjusted dynamically according to data unevenness. The rendering partition equation group determined N v = 3 GPU memory partitions, each with a rendering task volume T i adjusted in real time according to data complexity and target frame rate.

[0200] In addition, the researchers also established a visual expression accuracy A visual evaluation system. Through calculation, the position accuracy A position = 0.92, color accuracy A color = 0.88, size accuracy A size = 0.85, and shape accuracy A shape = 0.91, the comprehensive A visual = 0.89.

[0201] In the actual rendering process, the researchers monitored indicators such as frame rate R fps and rendering delay. It was found that in some complex data areas, the frame rate would decrease, so the projection method and material properties were dynamically adjusted, and finally stabilized at around 60 FPS. At the same time, the visual mapping rule R was continuously optimized, A visual was improved to 0.92, further improving the visualization effect.

[0202] Through the above steps, the research team successfully visualized the flow cytometry data using the method of the present invention. Taking a typical patient sample as an example, the results of data visualization are introduced in detail as follows:

[0203] Table 1 Characteristic data of immune cell samples of a certain patient

[0204] Number FSC SSC FL1 FL2 FL3 FL4 FL5 FL6 FL7 FL8 1 102 245 345 215 163 287 231 176 290 192 2 189 301 412 278 211 342 289 231 344 265 3 134 265 376 231 185 309 260 202 315 228 4 167 289 398 255 203 326 276 215 331 246 5 145 273 385 241 194 317 268 209 323 237 … … … … … … … … … … …

[0205] Table 2 10-dimensional characteristic data after dimension reduction

[0206]

[0207] According to the above 10-dimensional reduced characteristic data Y, the researchers used the visualization rendering method of the present invention to map it to a two-dimensional plane. As shown in the figure, different colored points represent the positions of cell samples in the characteristic space, and the size and shape of the points reflect the changes in other dimensional characteristics. Overall, these immune cell samples show obvious clustering distribution in the characteristic space, indicating that there is a certain functional correlation between them.

[0208] Further observation can find that samples 1, 3 and 5 are gathered in the lower left corner, and their FSC, SSC and part of the fluorescence signals (FL1, FL3, FL5, FL7) are relatively low, while FL2, FL4, FL6 and FL8 are relatively high, which may represent a relatively weak functional immune cell subpopulation. On the contrary, samples 2 and 4 are located in the upper right corner, and their FSC, SSC and most of the fluorescence signals are relatively high, which may belong to a relatively strong functional immune cell subpopulation.

[0209] Through observation of the visualization result, the researchers preliminarily judged that there were two types of immune cell subpopulations with obvious functional differences in the patient's body. In order to further verify this conjecture, they calculated the relative proportion of each subpopulation: the lower left corner cluster accounted for 40%, and the upper right corner cluster accounted for 60%. This result is consistent with the patient's clinical manifestations, such as decreased immune function and increased risk of infection.

[0210] The following provides multiple chart cases of the present application for data visualization for flow detection:

[0211] Figure 6 The subpopulation time series analysis chart: shows the proportion change of two main cell subpopulations over time. The blue line represents subpopulation A, and the red line represents subpopulation B, and their proportion changes dynamically. The peak fluctuation range is marked in the figure, and the filled area is used to represent the relationship between the two subpopulations. Helps researchers understand the time dynamics of the proportion of cell subpopulations; assess the stability of cell subpopulations through fluctuation amplitude (ΔP_max=10%); provide time reference for clinical diagnosis; can be used to predict the dynamic response characteristics of the immune system.

[0212] Figure 7 The feature correlation chart: uses a heat map to show the correlation between FSC, SSC and 6 fluorescence channels. The color changes from deep blue to deep red to represent the correlation coefficient from -1 to 1, and the specific correlation coefficient value is marked in each grid. Reveals the internal relationship between different features; helps identify redundant features and optimize feature selection; guides the development of dimensionality reduction strategies; provides feature basis for cell classification.

[0213] Figure 8 The density distribution chart: shows the cell population distribution density in the CD45+ and CD3+ expression feature space. The density size is represented by the color depth, and the main density peak is marked. The scatter plot superimposes the original data distribution. Directly show the aggregation characteristics of the cell population; help identify the boundaries of cell subpopulations; evaluate the heterogeneity of the cell population; provide density basis for cell classification.

[0214] Figure 9Feature projection plot: Multi-view displays the feature distribution of CD4+, CD8+ T cells and B cells in three dimensions. The upper three subplots show the 2D projection between two of the three dimensions, and the lower part shows the complete 3D projection. The color represents the integrated activity index. Provides intuitive visualization of multi-dimensional data; helps understand the spatial relationship between cell types; reveals cell functional status through integrated activity index; supports cell classification and phenotype analysis.

[0215] Figure 10 State transition plot: Displays the transition relationship diagram of cell functional status. The transition direction between the five state nodes is indicated by arrows, and the transition probability is labeled. The horizontal axis represents cell activity, and the vertical axis represents functional index. Display the dynamic process of cell state transition; quantify the transition probability between different states; predict the evolution trend of cell functional status; provide time window reference for therapeutic intervention.

[0216] Figure 11 Immune response plot: Displays the response intensity curve of T cells, B cells and NK cells over time. The peak time point of T cell response is labeled, and the three cell types are distinguished by different colors. Compare the response characteristics of different immune cells; determine the optimal time window of immune response; evaluate the overall functional status of the immune system; guide the development of immunotherapy programs.

[0217] These charts together constitute a complete immune cell function analysis system, which can comprehensively support researchers to understand and analyze complex immune cell data, and provide strong support for clinical diagnosis and treatment decisions. Each chart is designed for specific analysis needs and can be used independently or in combination to provide more comprehensive data insights.

[0218] It should be noted that the variables involved in the present application are shown in Table 3.

[0219] Table 3 Variable explanation table

[0220]

[0221]

[0222] The above description is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A data visualization method for flow cytometry, characterized in that: The following steps are involved: S10, setting the length of the data stream collection time segment to a preset duration, and establishing a data stream cache queue based on the preset duration; S20, collecting a multidimensional data stream from the data stream cache queue, performing feature extraction on the multidimensional data stream, and obtaining multidimensional feature data; S30, constructing a fast dimensionality reduction equation group to perform parallel calculation processing on the multi-dimensional feature data to generate feature data after dimensionality reduction; S40, dividing the multidimensional feature data into a plurality of data blocks, and performing parallel dimensionality reduction calculation using a graphics processor; S50, constructing a visual mapping rule for the feature data after dimensionality reduction, and establishing a rendering buffer in a video memory partition based on the visual mapping rule; S60, processing the data in the rendering buffer in batches using a resource allocation equation group; S70: Establish a data update queue in the memory partition, write the newly added data stream into the data update queue, and perform real-time dimensionality reduction processing on the data in the data update queue; S80, calculating a visual expression accuracy based on the dimensionally reduced feature data, and performing online optimization on the visual mapping rule according to the visual expression accuracy; S90: Update rendering parameters using the optimized visual mapping rule, and complete real-time visual rendering of the data stream in a graphics processor.

2. A data visualization method for flow cytometry according to claim 1, characterized in that: The step S10 specifically includes: Step 101: Determine the length of a data stream acquisition time segment based on system memory capacity and data stream rate; Step 102: Calculate the size of a buffer queue based on the length of the data stream collection time segment, where the size of the buffer queue is determined by the data stream rate, the number of data dimensions, and the buffer coefficient. Step 103: establishing a data stream cache queue according to the cache queue size for temporarily storing data streams; Step 104: Dynamically monitor the data stream cache queue and automatically expand the capacity when the data volume exceeds a preset threshold; Step 105: adjusting the length of the data stream collection time segment in real time through the data stream rate monitoring module; The step S20 specifically includes: Step 201: Read the multidimensional data stream from the data stream cache queue; Step 202: Performing linear transformation on the multidimensional data stream using a feature extraction matrix; Step 203: normalize the linearly transformed feature data; Step 204: Calculate the statistical characteristics of the characteristic data, including the mean and standard deviation; Step 205: Standardize the feature data according to the statistical features to obtain multi-dimensional feature data.

3. A data visualization method for flow cytometry according to claim 2, characterized in that: The step S30 specifically includes: Step 301: Calculate the conditional probability distribution between each data point in the multidimensional feature data; Step 302: construct a high-dimensional space distance matrix according to the conditional probability distribution; Step 303: Establish a mapping weight matrix based on the high-dimensional space distance matrix; Step 304: construct constraint conditions using local preservation items, topology preservation items, and global preservation items; Step 305: Perform dimensionality reduction processing on the multi-dimensional feature data through parallel computing to obtain feature data after dimensionality reduction.

4. A data visualization method for flow cytometry according to claim 3, characterized in that: The step S40 specifically includes: Step 401: Determine the data block size based on the graphics processor memory capacity and the number of parallel computing threads; Step 402: Divide the multidimensional feature data into multiple data blocks according to the data block size; Step 403: Allocate an independent computing thread to each data block; Step 404: Establish a parallel computing task queue in the graphics processor; Step 405: Perform parallel dimension reduction calculation processing on the data block.

5. A data visualization method for flow cytometry according to claim 4, characterized in that: The step S50 specifically includes: Step 501: Calculate the rendering buffer size based on the video memory capacity and the number of buffers; Step 502: Create multiple rendering buffers in the video memory partition; Step 503: assign an independent rendering task to each rendering buffer; Step 504: constructing a visual mapping rule including a position mapping rule, a color mapping rule, a size mapping rule, and a shape mapping rule; Step 505: Apply the visual mapping rule to the rendering buffer; The step S60 specifically includes: Step 601: Calculate the total system memory capacity and the number of memory partitions. Step 602: Determine the upper limit of the capacity of each memory partition; Step 603: Establish resource allocation equations; Step 604: batch-process the data in the rendering buffer according to the resource allocation equation group; Step 605: Execute rendering calculation tasks on each batch of data.

6. A method for visualizing data for flow cytometry according to claim 5, characterized in that: The step S70 specifically includes: Step 701: Calculate the data update queue capacity; Step 702: Establish a data update queue in the memory partition; Step 703: Write the newly added data stream into the data update queue; Step 704: perform dimensionality reduction processing on the data in the data update queue using a batch processing method; Step 705: Update the data after dimensionality reduction processing to the rendering buffer.

7. A method for visualizing data for flow cytometry according to claim 6, characterized in that: The step S80 specifically includes: Step 801: Calculate the position accuracy index, color accuracy index, size accuracy index, and shape accuracy index; Step 802: Calculate the visual expression accuracy based on the weight coefficient of each accuracy index; Step 803: Calculate the gradient of the visual expression accuracy with respect to the mapping rule; Step 804: Update the visual mapping rule using an exponentially decaying learning rate; Step 805: Verify the validity of the updated visual mapping rules.

8. A method for visualizing data for flow cytometry according to claim 7, characterized in that: The step S90 specifically includes: Step 901: Apply the optimized visual mapping rules to the rendering parameters; Step 902: Establish a rendering task queue in the graphics processor; Step 903: Perform real-time rendering on the data in the rendering task queue; Step 904: monitor rendering performance indicators; Step 905: Dynamically adjust rendering parameters according to performance indicators.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the data visualization method for flow cytometry according to any one of claims 1 to 8.

10. A data visualization system for flow cytometry, characterized in that: Contains the computer-readable storage medium of claim 9.