A multi-dimensional health data processing system and method for internal medicine patients
By performing dimensional filtering, correlation learning, and sparse reconstruction on multidimensional health data of internal medicine patients, a derived data view is generated, which solves the problems of data redundancy and waste of computing resources in existing technologies, and improves the scenario adaptability and decision support accuracy of internal medicine health data processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies cannot effectively distinguish data from different sources, natures, and clinical value densities when processing multidimensional health data of internal medicine patients, leading to data redundancy and wasted computing resources, and reducing the sensitivity and specificity of decision support.
By acquiring multidimensional health data streams, performing dimensional filtering and semantic enhancement, and utilizing correlation learning and sparse reconstruction, a time-varying nonlinear strong coupling matrix is generated, a dimensional interaction graph is constructed, and derived data views are generated in response to external query conditions.
It improves the scenario adaptability of internal medicine health data processing, enhances computing efficiency and the accuracy of clinical decision-making, and achieves efficient transformation from raw data to scenario-based insights.
Smart Images

Figure CN121565477B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a multidimensional health data processing system and method for internal medicine patients. Background Technology
[0002] Internal medicine patients refer to individuals who undergo diagnosis, treatment, and long-term health management in the internal medicine field of medical institutions due to non-traumatic, non-surgical dysfunction or disease of internal organs or systems. The health status of internal medicine patients is usually driven by complex physiological and pathological processes, manifesting as a dynamic evolution of multiple systems, multiple indicators, and long cycles, requiring comprehensive management through the integrated use of drugs, lifestyle interventions, and continuous monitoring.
[0003] Current technologies for processing multi-source heterogeneous health data collect, store, and present all data dimensions with equal weight, or perform simple linear overlay and statistical summarization. This indiscriminately places data from different sources, with different properties, and varying clinical value densities onto the same analytical plane. A large amount of data with low timeliness, high variability, or weak correlation to the current analytical objective is mixed with a few key, high signal-to-noise ratio clinical core signals. This not only causes significant data redundancy and waste of computational resources, but also causes truly indicative early abnormal signals and interaction patterns to be ignored amidst background noise and irrelevant information, interfering with the focus and accuracy of subsequent analysis, and reducing the sensitivity and specificity of decision support. Therefore, how to reconstruct the dimensional relationships of multi-dimensional health data to improve the scenario adaptability of internal medicine health data processing has become a challenge for the industry. Summary of the Invention
[0004] This application provides a multidimensional health data processing system and method for internal medicine patients, which can realize the dimensional correlation reconstruction of multidimensional health data, thereby improving the scenario adaptability of internal medicine health data processing.
[0005] In a first aspect, this application provides a method for processing multidimensional health data of internal medicine patients, comprising the following steps:
[0006] Acquire a multidimensional health data stream of internal medicine patients, the multidimensional health data stream containing four mutually orthogonal independent data dimensions;
[0007] The multidimensional health data stream is filtered by priority scores of each independent data dimension to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement.
[0008] Based on the first intermediate data, the correlation degree of the orthogonal coupling relationship between each independent data dimension is learned to obtain a time-varying nonlinear strength coupling matrix. Then, based on the strength coupling matrix, the correlation topology between each independent data dimension is sparsely reconstructed to obtain the second intermediate data of dimensional interaction in the multidimensional health data stream.
[0009] Based on the second intermediate data, in response to the external query conditions of the current application scenario, a derived data view of the current application scenario is obtained, and then the derived data view is used to visualize the common feature patterns of the internal medicine patient group.
[0010] In some embodiments, the first intermediate data for structural dimensionality reduction and semantic enhancement is obtained by performing dimensionality filtering on the multidimensional health data stream using priority scores of each independent data dimension. Specifically, this includes:
[0011] Priority scores are calculated for each independent data dimension based on a pre-defined clinical scenario task;
[0012] Based on the priority scores, the data channels of the corresponding independent data dimensions in the multidimensional health data stream are weighted, amplified, and suppressed and filtered to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement.
[0013] In some embodiments, learning the correlation degree of the orthogonal coupling relationship between each independent data dimension based on the first intermediate data to obtain a time-varying nonlinear strength coupling matrix specifically includes:
[0014] Within the set sliding time window, extract the temporal feature vectors of each independent data dimension from the first intermediate data;
[0015] The correlation strength between temporal feature vectors of any two independent data dimensions is calculated in parallel using a correlation calculation module based on mutual information and nonlinear kernel functions.
[0016] Fill all the instantaneous correlation strengths into the corresponding positions of a fourth-order square matrix to obtain the time-varying nonlinear strength coupling matrix.
[0017] In some embodiments, the second intermediate data of dimensional interaction in the multidimensional health data stream is obtained by sparsely reconstructing the correlation topology between each independent data dimension based on the strength coupling matrix. Specifically, this includes:
[0018] Determine the sparsity association threshold for dimensional interactions in a multidimensional health data stream;
[0019] Based on the sparsification correlation threshold, the strength coupling matrix is sparsified to obtain a sparsified correlation matrix.
[0020] The sparsed correlation matrix is used as an adjacency matrix of a weighted directed graph to obtain the dimensional interaction graph of the multidimensional health data stream.
[0021] Based on the dimensional interaction graph, each independent data dimension is mapped to a new low-dimensional potential space, thereby obtaining the second intermediate data of dimensional interaction in the multidimensional health data stream.
[0022] In some embodiments, obtaining a derived data view of the current application scenario based on the second intermediate data in response to external query conditions of the current application scenario specifically includes:
[0023] Obtain the external query conditions for the current application scenario, and then identify the clinical intent, data dimension focus, and expected granularity from each external query condition;
[0024] Based on the clinical intent and data dimension focus, relevant dimensional subsets of data and interaction relationship subgraphs are selected from the second intermediate data;
[0025] Based on the desired granularity and the interaction relationship subgraph, the dimensional subset data is subjected to time-series slicing to obtain an intermediate dataset that conforms to the query semantics;
[0026] The intermediate dataset is encapsulated and formatted according to a predefined view template to obtain a derived data view that can be directly used for display in the current application scenario.
[0027] In some embodiments, using the derived data view to visualize common characteristic patterns of the medical patient population specifically includes:
[0028] Cluster analysis was performed on the derived data view to obtain multiple patient subgroups that exhibited similar patterns in the interaction relationships of the selected dimensions;
[0029] Extract the dimensional feature contours of each patient subgroup, and then generate corresponding feature pattern labels;
[0030] The characteristic pattern labels of each patient subgroup and their distribution proportion in the internal medicine patient population are combined and visualized.
[0031] In some embodiments, the independent data dimensions include physiological, biochemical, behavioral, and subjective reporting dimensions.
[0032] Secondly, this application provides a multidimensional health data processing system for internal medicine patients, used to execute a multidimensional health data processing method for internal medicine patients, which includes a visualization unit, the visualization unit comprising:
[0033] The acquisition module is used to acquire a multidimensional health data stream of internal medicine patients, wherein the multidimensional health data stream contains four mutually orthogonal independent data dimensions;
[0034] The processing module is used to perform dimensional filtering on the multidimensional health data stream by using the priority scores of each independent data dimension to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement.
[0035] The processing module is further configured to perform correlation learning on the orthogonal coupling relationship between each independent data dimension based on the first intermediate data to obtain a time-varying nonlinear strength coupling matrix, and then perform sparsification reconstruction on the correlation topology between each independent data dimension according to the strength coupling matrix to obtain the second intermediate data of dimensional interaction in the multidimensional health data stream.
[0036] The execution module is used to obtain a derived data view of the current application scenario based on the second intermediate data and respond to the external query conditions of the current application scenario. Then, the derived data view is used to visualize the common feature patterns of the internal medicine patient group.
[0037] Thirdly, this application provides a computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described multidimensional health data processing method for internal medicine patients.
[0038] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for processing multidimensional health data for internal medicine patients.
[0039] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:
[0040] This application provides a multidimensional health data processing system and method for internal medicine patients. The system acquires a multidimensional health data stream containing four mutually orthogonal independent data dimensions. Dimensional filtering is performed on the multidimensional health data stream using priority scores of each independent data dimension to obtain first intermediate data for structural dimensionality reduction and semantic enhancement. Based on the first intermediate data, correlation learning is performed on the orthogonal coupling relationships between the independent data dimensions to obtain a time-varying, nonlinear strength coupling matrix. Then, based on the strength coupling matrix, the correlation topology between the independent data dimensions is sparsely reconstructed to obtain second intermediate data on dimensional interactions in the multidimensional health data stream. Using the second intermediate data as a basis, a derived data view of the current application scenario is obtained in response to external query conditions. Finally, the derived data view is used to visualize the common characteristic patterns of the internal medicine patient group.
[0041] Therefore, in this application, based on the second intermediate data, a derived data view of the current application scenario is obtained in response to the external query conditions of the current application scenario. This derived data view is then used to visualize the common characteristic patterns of the internal medicine patient group. First, determining the first intermediate data yields a feature-enhanced data set optimized for clinical tasks, thus constructing a high-quality data foundation. Through a dynamic priority mechanism, intelligent filtering and semantic enhancement are performed on the multidimensional data stream, completing a data space dimensionality reduction and feature selection for the clinical scenario. While retaining key task information, the interference of irrelevant dimensional noise is significantly suppressed, allowing the subsequent association learning module to focus on data dimensions and change patterns with high clinical value. This not only improves the computational efficiency of the overall data processing flow but also ensures that the subsequently mined relationships closely align with clinical intent, enhancing the system's ability to analyze and adapt to diverse clinical scenarios. Then, determining the second intermediate data... Intermediate data can be used to obtain low-dimensional feature representations containing key dimensional interaction information, thus providing a structured and interpretable data core for scenario-driven visualization. Through time-varying nonlinear correlation learning and sparse topology reconstruction, the complex dimensional interactions implicit in the first intermediate data are extracted into a clear and quantifiable correlation network. The high-dimensional, dynamic health data stream is transformed into a graph representation with interaction relationships as the hub and corresponding low-dimensional embeddings, constructing an abstract model that reflects the inherent correlation structure of the data. This not only compresses the data dimensions but also retains the key interaction semantics, enabling flexible responses to different external queries and rapid generation of derived views focusing on the correlation patterns of specified dimensions. This achieves efficient and targeted transformation from raw data to scenario-based insights, significantly improving the fit between data processing results and clinical decision-making needs. In summary, based on the above scheme, dimensional correlation reconstruction of multidimensional health data can be achieved, thereby improving the scenario adaptability of internal medicine health data processing. Attached Figure Description
[0042] Figure 1 This is an exemplary flowchart of a multidimensional health data processing method for internal medicine patients, as shown in some embodiments of this application;
[0043] Figure 2 This is an exemplary flowchart illustrating the determination of a multi-scale digital surface model according to some embodiments of this application;
[0044] Figure 3 These are schematic diagrams of the structure of visual units shown in some embodiments of this application;
[0045] Figure 4 This is a schematic diagram of the structure of a computer device that implements a multidimensional health data processing method for internal medicine patients, according to some embodiments of this application. Detailed Implementation
[0046] To better understand the technical solution of this application, the technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0047] refer to Figure 1 The figure is an exemplary flowchart of a multidimensional health data processing method for internal medicine patients according to some embodiments of this application. The figure mainly includes the following steps:
[0048] In step 101, a multidimensional health data stream of an internal medicine patient is obtained, the multidimensional health data stream containing four mutually orthogonal independent data dimensions.
[0049] It should be noted that, in this application, the independent data dimensions include physiological, biochemical, behavioral, and subjective reporting dimensions; multidimensional health data streams are multi-source heterogeneous data sets that characterize the health status of internal medicine patients and are generated continuously or periodically over time; independent data dimensions are data categories that are used to characterize health status from a specified perspective, are non-overlapping in information, and can be independently measured; physiological dimensions are continuous quantitative indicator data reflecting the real-time status of the patient's basic bodily functions; biochemical dimensions are discrete laboratory test results data that assess the patient's metabolic and organ function status; behavioral dimensions are habitual data that record the patient's daily health-related activity patterns and routines; and subjective reporting dimensions are descriptive data used to obtain the patient's subjective experience and assessment of their symptoms, feelings, and quality of life.
[0050] In practice, through integrated data interfaces, data is collected and aggregated from multiple sources, including hospital information systems, laboratory information systems, wearable devices, and patient mobile terminal applications, according to preset cycles or event triggers. For example, heart rate and step count data are continuously read from wearable devices as physiological dimension data, the latest blood glucose, creatinine and other test reports are periodically extracted from the hospital database as biochemical dimension data, medication time and diet content recorded in patient application logs are extracted as behavioral dimension data, and pain scores or fatigue feelings filled in by patients through scales are collected simultaneously as subjective reporting dimension data. Data with different formats and update frequencies are aligned and integrated through patient unique identifiers and a unified timeline to form a structured multidimensional health data stream.
[0051] In step 102, the multidimensional health data stream is dimensionally filtered by the priority scores of each independent data dimension to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement.
[0052] In some embodiments, the first intermediate data for structural dimensionality reduction and semantic enhancement can be obtained by performing dimensionality filtering on the multidimensional health data stream using priority scores of each independent data dimension, which can be achieved through the following steps:
[0053] Priority scores are calculated for each independent data dimension based on a pre-defined clinical scenario task;
[0054] Based on the priority scores, the data channels of the corresponding independent data dimensions in the multidimensional health data stream are weighted, amplified, and suppressed and filtered to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement.
[0055] It should be noted that in this application, the first intermediate data is the basic dataset used for deep correlation analysis; the clinical scenario task is a task description used to clarify the specific clinical goals to be achieved in this multidimensional health data processing; and the priority score is a quantitative indicator that quantifies the importance of each independent data dimension to completing the current clinical scenario task.
[0056] In practice, firstly, a clinical scenario knowledge base is pre-built and maintained. This knowledge base associates different clinical tasks (e.g., "early warning of acute heart failure" or "assessment of long-term glycemic control in diabetes") with a set of prior knowledge of dimensional importance determined by expert experience or machine learning models. When the processing flow is initiated, the current clinical scenario task is first determined based on user input or automatic interpretation. Then, the baseline importance vector corresponding to the task is retrieved from the knowledge base. This baseline importance vector assigns an initial importance base value to each independent data dimension. Simultaneously, the historical data characteristics of the current specific patient (e.g., a dimension with extremely high variability in the patient's historical data) are introduced as an adjustment factor. Through weighted comprehensive calculation, the product of the baseline importance base value and the adjustment factor is used as the priority score of the corresponding independent data dimension. The priority scores of each independent data dimension can be obtained in the above way. Then, for independent data dimensions with priority scores greater than a preset priority threshold, the entire data channel corresponding to the independent data dimension will undergo enhancement processing, for example... If the biochemical dimension has the highest priority score, the data values of all laboratory indicators (e.g., troponin, B-type natriuretic peptide) under that dimension will be multiplied by a gain coefficient greater than 1, or more feature expression bit widths will be allocated during the feature extraction stage, thus giving them a more dominant position in subsequent analysis. Conversely, for independent data dimensions with priority scores less than or equal to a preset priority threshold, their corresponding data channels will undergo simplification. For example, if the behavioral dimension has the lowest priority score, the complex behavioral sequence data under that dimension may be compressed into a few key statistics (e.g., average daily activity level, medication adherence rate), while discarding its high-frequency detailed fluctuations. Priority-based differential processing can achieve a purposeful, non-uniform compression and reshaping of the high-dimensional data space. After performing such priority-based weighted amplification or suppression filtering operations in parallel on all data channels, the processed independent dimension data are re-integrated and encapsulated to obtain a new dataset with a reduced overall data volume but enhanced key clinical semantic information. This new dataset is used as the first intermediate data.
[0057] In step 103, the correlation degree of the orthogonal coupling relationship between each independent data dimension is learned based on the first intermediate data to obtain a time-varying nonlinear strength coupling matrix. Then, the correlation topology between each independent data dimension is sparsely reconstructed according to the strength coupling matrix to obtain the second intermediate data of dimensional interaction in the multidimensional health data stream.
[0058] In some embodiments, the correlation degree learning of the orthogonal coupling relationship between each independent data dimension based on the first intermediate data to obtain the time-varying nonlinear strength coupling matrix can be achieved by the following steps:
[0059] Within the set sliding time window, extract the temporal feature vectors of each independent data dimension from the first intermediate data;
[0060] The correlation strength between temporal feature vectors of any two independent data dimensions is calculated in parallel using a correlation calculation module based on mutual information and nonlinear kernel functions.
[0061] Fill all the instantaneous correlation strengths into the corresponding positions of a fourth-order square matrix to obtain the time-varying nonlinear strength coupling matrix.
[0062] It should be noted that, in this application, the sliding time window is an analysis range with a fixed length that can move as time progresses, used to extract consecutive adjacent time segments from the first intermediate data; the time series feature vector is a set of numerical features that characterize the dynamic change patterns and statistical characteristics of each independent data dimension within a specified time window; the correlation calculation module is a calculation unit used to quantify the degree of mutual dependence or synergistic change between two different data dimensions; the instantaneous correlation strength is a scalar value describing the degree of correlation between two independent data dimensions within a specified sliding time window; the fourth-order square matrix is a four-row, four-column square matrix data structure used to store the pairwise instantaneous correlation strength between the four independent data dimensions; and the strength coupling matrix is a matrix representation used to characterize the dynamic correlation between the dimensions within the multidimensional health data under a specified time window.
[0063] In practical implementation, firstly, a fixed-length time window, such as 24 hours, is preset and used as the basic analysis unit. Starting from the initial time point of the first intermediate data, all data within the first 24-hour period are extracted. For each independent data dimension (e.g., physiological dimension) within this time window, a series of feature extraction operations are performed. These operations include calculating the statistical characteristics of the data for that dimension within this window, such as mean, variance, and extreme values, as well as time-domain characteristics, such as the slope of the trend and the intensity of periodicity. They may also include obtaining the main frequency components through simple transformations (e.g., Fast Fourier Transform). All the aforementioned feature values calculated for each independent data dimension are arranged in a predetermined order to form a fixed-dimensional numerical array. This numerical array is the time-series feature vector of that dimension within the current time window. The above process is repeated for the four independent data dimensions (physiological, biochemical, behavioral, and subjective reports) to obtain the time-series feature vectors for each independent data dimension within the current time window. Then, a dedicated correlation calculation module is deployed. The core algorithm of this module combines mutual information and a nonlinear kernel function. Mutual information is used to measure the time-series characteristics of a given dimension. After vectorization, how much is the uncertainty of the other dimension of time-series feature vectors reduced? Mutual information can capture linear and nonlinear dependencies. To efficiently calculate the mutual information between high-dimensional feature vectors, the correlation calculation module adopts an estimation method based on kernel functions (e.g., Gaussian kernels). This method maps feature vectors to a high-dimensional regenerative kernel Hilbert space and calculates the eigenvalues of its covariance operator in this space to approximate the mutual information. During calculation, the module receives the time-series feature vectors of the four dimensions under the current time window; it calculates all dimension pairings (6 pairs in total) in parallel, such as simultaneously calculating the "physiological dimension and biochemical dimension". The correlation between the pairing of "physiological dimension and behavioral dimension" is calculated as follows: For each pair of dimensions, the module inputs the temporal feature vectors of the two dimensions into the algorithm above and outputs a scalar value between 0 and 1. This scalar value is the instantaneous correlation strength between the two dimensions within this time window. The closer the correlation strength is to 1, the closer the two dimensions change in synergy within this window; the closer it is to 0, the more independent they are. Finally, a 4x4 matrix template is maintained, where the rows and columns represent the four independent data dimensions in a fixed order (e.g., the first row / column corresponds to the physiological dimension, the second row / column corresponds to the biochemical dimension, and so on).After calculating the instantaneous correlation strength of all dimensional pairs within a sliding time window, the system fills these strength values into the corresponding positions of this matrix template. Specifically, the instantaneous correlation strength value of "physiological dimension and biochemical dimension" is simultaneously filled into the first row, second column and the second row, first column of the matrix (because the correlation is usually considered symmetrical). Similarly, the strength value of "physiological dimension and behavioral dimension" is filled into the first row, third column and the third row, first column, and so on, until all 6 pairs of correlation strength values are filled into the matrix. The elements on the diagonal of the matrix (the correlation between each dimension and itself) are fixed at 1 or set according to the self-information of the eigenvector of that dimension. The complete fourth-order square matrix obtained after filling is the strength coupling matrix under the current specified sliding time window. This strength coupling matrix comprehensively reflects the complex (non-linear) interaction strength between the data of each dimension within the current time period. As the sliding window moves forward with time and repeats the above process, a series of matrix sequences that may change numerically will be generated, thereby realizing the capture and expression of the "time-varying non-linear" characteristics of the correlation between dimensions.
[0064] In some embodiments, the correlation topology between each independent data dimension is sparsely reconstructed based on the strength coupling matrix to obtain second intermediate data on dimensional interactions in the multidimensional health data stream, as referenced. Figure 2 The figure described above is a flowchart illustrating the process of determining a multi-scale digital surface model in some embodiments of this application. In this embodiment, determining the multi-scale digital surface model can be achieved using the following steps:
[0065] In step 1031, the sparsity association threshold of dimensional interactions in the multidimensional health data stream is determined;
[0066] In step 1032, the strength coupling matrix is sparsified based on the sparsification correlation threshold to obtain a sparsified correlation matrix;
[0067] In step 1033, the sparsed correlation matrix is used as an adjacency matrix of a weighted directed graph to obtain the dimensional interaction graph of the multidimensional health data stream;
[0068] In step 1034, each independent data dimension is mapped to a new low-dimensional potential space based on the dimensional interaction graph, thereby obtaining the second intermediate data of dimensional interaction in the multidimensional health data stream.
[0069] It should be noted that, in this application, the sparsity association threshold is a critical value used to distinguish between strong and weak associations; the sparsity association matrix is a simplified association strength matrix that retains only the strong interaction relationships between key dimensions; the adjacency matrix of the weighted directed graph is a matrix representation used to describe whether there are connections between nodes in the graph and the connection weights; the dimensional interaction graph is a graph model that intuitively represents the interaction relationships and interaction strengths between various independent data dimensions in a multidimensional health data stream; the low-dimensional latent space is a vector space used to simultaneously encode the attributes of independent data dimensions themselves and their key interaction relationships at a low-dimensional level; and the second intermediate data is a low-dimensional feature dataset that has been integrated with key dimension interaction information to support the generation of the final derived view.
[0070] In practice, firstly, the strength coupling matrix generated under the current sliding time window is read. This matrix contains the instantaneous correlation strength values between all pairs of dimensions (excluding autocorrelation values on the diagonal). Then, the statistical distribution characteristics of all off-diagonal elements (i.e., the values representing correlations between different dimensions) are calculated, such as the mean and standard deviation of these values. The sparsity correlation threshold can be dynamically generated according to the rule of "mean plus N times the standard deviation," where N is a configurable parameter (default set to 1.5 or 2.0). The purpose of this calculation is to filter out strength values that are statistically significantly higher than the average correlation level, and to treat the corresponding interactions as critical and non-critical. The process involves several steps: First, random dimensional interactions are implemented. The system calculates a dynamic value and sets it as the sparsity correlation threshold for the current window. Next, this threshold is applied to the original strong coupling matrix, performing element-level filtering. Specifically, each off-diagonal element in the strong coupling matrix (representing the correlation strength between two different dimensions) is traversed. If the element's value is greater than or equal to the sparsity correlation threshold, its value is retained in the matrix. If the element's value is less than the threshold, its value is set to zero. Elements on the matrix's diagonal (representing the correlation between a dimension and itself) are typically retained, with their values remaining unchanged or processed according to specified rules. After this traversal and conditional replacement, most elements representing weak correlations in the original strong coupling matrix are cleared to zero, leaving only a few strongly correlated elements exceeding the threshold and the diagonal elements with non-zero values. This results in a matrix with most elements being zero—the sparsity correlation matrix.This matrix only depicts the interactions between dimensions considered most important within the current time window. Then, the sparse association matrix is interpreted as a mathematical representation of a weighted directed graph. In this graph model, each independent data dimension (physiological, biochemical, behavioral, subjective report) is considered a node. The row and column indices in the sparse association matrix correspond to the starting and target nodes of the graph, respectively. If the element in the i-th row and j-th column of the matrix is non-zero, then there exists a directed edge from node i to node j in the graph. The magnitude of this non-zero value serves as the weight of this directed edge, used to quantify the influence of dimension i on dimension j. Since the calculation of association strength may be asymmetric, the graph can be directed. If the calculation of association strength in the technical implementation is symmetric, then the matrix is symmetric, and the graph can be considered undirected. Based on the positions and values of all non-zero off-diagonal elements in the sparse association matrix, the corresponding set of nodes, set of edges, and set of weights for each edge are constructed, together forming a complete dimensional interaction graph. This dimensional interaction graph reveals the relatively strong interactions between data dimensions within the current time window. The algorithm considers strong direct interactions and their relative strength. Finally, taking each node in the dimensional interaction graph (i.e., each independent data dimension) as input, the graph embedding algorithm aims to learn a mapping function such that nodes closely related (connected by edges with large weights) in the original dimensional interaction graph have similarly close vector representations in the new low-dimensional latent space after mapping. Conversely, nodes with distant or no direct relationship have much farther vector representations. A message-passing mechanism similar to that in graph neural networks can be used. Each node's initial feature can be a summary of its corresponding temporal feature vector. Then, the node iteratively updates its feature representation based on the features of its neighboring nodes (i.e., other related dimensions) and the weights of the interaction edges. After iteration, each node is transformed into a fixed-length low-dimensional real vector. This vector not only contains the temporal feature information of the dimension itself but also incorporates information about its interactions with other dimensions through the interaction graph. The low-dimensional vectors of the four independent data dimensions are sequentially concatenated as the second intermediate data for dimensional interactions in the multidimensional health data stream.
[0071] In step 104, based on the second intermediate data, in response to the external query conditions of the current application scenario, a derived data view of the current application scenario is obtained, and then the derived data view is used to visualize the common feature patterns of the internal medicine patient group.
[0072] In some embodiments, obtaining a derived data view of the current application scenario based on the second intermediate data in response to external query conditions of the current application scenario can be achieved through the following steps:
[0073] Obtain the external query conditions for the current application scenario, and then identify the clinical intent, data dimension focus, and expected granularity from each external query condition;
[0074] Based on the clinical intent and data dimension focus, relevant dimensional subsets of data and interaction relationship subgraphs are selected from the second intermediate data;
[0075] Based on the desired granularity and the interaction relationship subgraph, the dimensional subset data is subjected to time-series slicing to obtain an intermediate dataset that conforms to the query semantics;
[0076] The intermediate dataset is encapsulated and formatted according to a predefined view template to obtain a derived data view that can be directly used for display in the current application scenario.
[0077] It should be noted that, in this application, external query conditions are user input parameters or preset instruction sets used to trigger and limit the generation of this data view; clinical intent is an abstract expression used to describe the specified clinical activities or decision-making goals to be served or supported by this data query; data dimension focus is used to clarify the independent data dimensions or combinations thereof that are the main focus and need to be presented in this data query; expected granularity is a parameter used to specify the data time resolution and level of detail required by this data query; dimension subset data is a set of low-dimensional feature vectors corresponding to some independent data dimensions selected from the second intermediate data to meet the current specified query requirements; interaction relationship subgraph is a local graph structure used to describe the key interaction relationships and their strengths between the dimensions in the dimension subset data; intermediate dataset is a standardized data set after spatiotemporal and relational reconstruction used directly for view template rendering; derived data view is a visual graphical interface or structured report used to ultimately present multidimensional health data insights to clinical users.
[0078] In practice, the system first receives externally input query requests via an application programming interface (API) or user interface. These requests contain structured query parameters or natural language descriptions. The system then parses the query conditions, extracting the core purpose from them using a pre-built rule engine or intent recognition model. For example, for the query condition "show the patient's blood glucose control trend and influencing factors over the past week," the system identifies its clinical intent as "long-term blood glucose control effect assessment and attribution analysis," with data dimensions focusing on "biochemical dimensions (blood glucose)" and "potentially related dimensions (e.g., behavioral dimensions, physiological dimensions)," and an expected granularity of "daily trend" and "correlation of influencing factors." The unstructured query conditions are then transformed into structured target instructions, clarifying the target intent of this view generation, the key data scope, and the required time. First, the analysis granularity is determined. Second, based on the clinical intent and data dimension focus, the second intermediate data is filtered. The second intermediate data consists of low-dimensional vector representations of all four dimensions and implicitly contains the interaction graph structure between them. First, the dimensions that need to be focused on are selected based on the data dimension focus, such as the "biochemical dimension" and the "behavioral dimension". The corresponding low-dimensional vectors are extracted from the second intermediate data to form a dimension subset data. At the same time, the dimension interaction graph used to generate the second intermediate data is reviewed. Based on the selected dimension focus, a subgraph containing these focus dimension nodes and the edges (interaction relationships) connecting these focus dimension nodes is extracted from the whole graph. For example, if the focus dimensions are "biochemical dimension A" and "behavioral dimension B," and there is a strong correlation edge between "A and B" in the original interaction graph, then this edge and its weight will be included in the interaction relationship subgraph. If there are other dimensions (e.g., "physiological dimension C") that are strongly correlated with A or B and have contextual value for understanding the relationship between A and B, it can also be decided according to rules whether to include node C and its related edges in the subgraph, ultimately obtaining an interaction relationship subgraph closely related to the query intent. Then, based on the identified expected granularity, the dimensional subset data is reorganized according to the time dimension. The second intermediate data itself has a time attribute (corresponding to the sliding time window that generated it).If the desired granularity is "weekly trend" and the original data is a "daily window" sequence, the system will merge the dimensional subset data of seven consecutive daily windows (each data point is a low-dimensional vector representation of that dimension within a day) in chronological order. During merging, the system can perform averaging or other aggregation operations on these seven vectors (e.g., taking the dominant pattern) to generate an aggregated vector representing the overall state of the week. At the same time, the interaction relationship subgraph may also change with the time window. The system can merge the subgraphs of each window within a week, for example, by calculating the average weight of the edges, to obtain an aggregated subgraph representing the typical interaction pattern of the week. The system integrates the aggregated dimensional subset data (aggregated vectors of each focus dimension) with the corresponding aggregated interaction relationship subgraph information (node and edge weights) to form an intermediate dataset that conforms to the query semantics. Finally, a view template library is maintained, storing pre-designed templates for different clinical intents and data dimension focus combinations. Each template defines the view's constituent elements (e.g., trend charts, radar charts, relationship network diagrams, tables), the position of each element, color mapping rules, axis labels, and how to extract data from the intermediate dataset to populate these elements. Based on the clinical intent and data dimension focus of the query, the system selects the most matching view template from the template library and populates the data in the intermediate dataset according to the template's binding rules. For example, the time-series aggregated vector of the "blood glucose" dimension in the dimensional subset data is bound to the Y-axis data of the trend line in the template; the aggregated vector of the "behavioral dimension (exercise volume)" is bound to another trend line; and the edge weight information in the interaction relationship subgraph is bound to the thickness of the lines in the relationship network diagram. After data binding is completed, the system encapsulates the data according to the template format specifications (e.g., JSON, XML, or the configuration file format of the visualization engine) to generate a content object containing complete view definitions and data. This content object is the derived data view that can be directly loaded and displayed by the front-end rendering engine or report generator.
[0079] In some embodiments, the visualization of common characteristic patterns of a group of internal medicine patients using the derived data view can be achieved through the following steps:
[0080] Cluster analysis was performed on the derived data view to obtain multiple patient subgroups that exhibited similar patterns in the interaction relationships of the selected dimensions;
[0081] Extract the dimensional feature contours of each patient subgroup, and then generate corresponding feature pattern labels;
[0082] The characteristic pattern labels of each patient subgroup and their distribution proportion in the internal medicine patient population are combined and visualized.
[0083] It should be noted that in this application, a patient subgroup is a subset of patients that has a high degree of consistency in interaction patterns on a specified dimension; a dimensional feature profile is a multidimensional statistical summary that provides a general description of the typical state value range and the typical interaction intensity range of the patient subgroup in the selected focal dimension; and a feature pattern label is a short text identifier that characterizes and distinguishes the core feature profiles of different patient subgroups.
[0084] In practice, the system first collects intermediate datasets corresponding to derived data views of a group of patients within a certain period (e.g., the past six months) that focus on the same clinical intent and similar data dimensions. These intermediate datasets contain key dimension state vectors and interaction relationship information for each patient within the corresponding time frame, which have undergone spatiotemporal and relational reconstruction. The system first extracts features for clustering from these intermediate datasets. These features include not only the aggregated state values of each focus dimension but also graph structure indicators such as edge weights and node centrality in the interaction relationship subgraph, collectively forming a multidimensional feature vector. Subsequently, an unsupervised clustering algorithm (e.g., density-based DBSCAN algorithm or hierarchical clustering algorithm) is used to analyze these feature vectors. This algorithm calculates the distance between patient feature vectors (e.g., Euclidean distance or cosine distance) and, based on preset clustering parameters (e.g., distance threshold, minimum number of neighbors), automatically groups patients with close distances into the same cluster, while classifying patients with greater distances into different clusters. After clustering, all patients belonging to the same cluster are considered to form a patient subgroup, exhibiting similar patterns in the dimensions of interest and interactions. Then, for each patient subgroup identified in step one, feature profiles are extracted. Specifically, for a given patient subgroup, the feature vectors of all patients within that group (i.e., the features used for clustering in the previous step) are iterated. For each dimension of the feature vector (e.g., average blood glucose level, strength of association between exercise and blood glucose), the central tendency (e.g., median) and dispersion (e.g., interquartile range) of the subgroup patients in that dimension are calculated. These statistics (e.g., "median blood glucose: 8.5 mmol / L, IQR: [7.2, 9.8]; median strength of exercise-blood glucose association: 0.75, IQR: [0.68, 0.82]") are summarized to form the dimensional feature profile of the subgroup, which is a structured data summary. Based on this feature profile, the system automatically generates an easy-to-understand feature pattern label through a rule engine or a lightweight natural language generation module. The generation logic typically identifies the most salient and distinctive features in the profile. For example, if a subgroup's profile shows "high and volatile blood glucose levels, strongly negatively correlated with diet adherence," the system might generate "high glucose volatility - diet sensitive" as its feature pattern label. Each subgroup receives a feature pattern label that describes its core pattern. Finally, the size (number of patients included) of each patient subgroup is calculated as a proportion of the total number of internal medicine patients analyzed. A combined visualization view is designed and generated. This view typically consists of two main parts: the first part uses a Sankey diagram or a modified pie / ring chart to visually represent the distribution proportion of each patient subgroup within the entire patient population.In the Sankey diagram, the left side represents a single source stream of all patients, while the right side branches into multiple subgroups. Each branch represents a subgroup, and its width is proportional to the number of patients in that subgroup (i.e., its distribution proportion). The branches are labeled with the characteristic pattern tags of that subgroup. The second part uses parallel coordinate plots or multidimensional radar charts to display the dimensional feature profiles of each subgroup in detail. In the parallel coordinate plot, the vertical axis represents different feature dimensions (e.g., blood glucose level, association strength, etc.), and each subgroup corresponds to a line segment. The position of this line segment on each dimension axis reflects the typical value (e.g., median) of that subgroup in that dimension, and the line color or style corresponds to the subgroup in the Sankey diagram. In this way, observers can clearly see the proportion of different pattern subgroups in the population and simultaneously compare the differences and commonalities of each subgroup in specific dimensional features. The system combines the calculated proportion data and feature profile data with predefined combined visualization templates, executes the graphics rendering engine, and finally outputs a complete and information-rich combined visualization chart as a visual presentation of the population's common feature pattern analysis results.
[0085] Furthermore, in another aspect of this application, in some embodiments, this application provides a multidimensional health data processing system for internal medicine patients. This multidimensional health data processing system for internal medicine patients includes a visualization unit, as referenced... Figure 3 The figure is a schematic diagram of the structure of a visualization unit according to some embodiments of this application. The visualization unit includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described below:
[0086] The acquisition module 201 in this application is mainly used to acquire the multidimensional health data stream of internal medicine patients. The multidimensional health data stream contains four mutually orthogonal independent data dimensions.
[0087] Processing module 202, in this application, is used to perform dimensional filtering on the multidimensional health data stream through the priority scores of each independent data dimension to obtain the first intermediate data of structural dimensionality reduction and semantic enhancement.
[0088] It should be noted that the processing module 202 is also used to learn the correlation degree of the orthogonal coupling relationship between each independent data dimension based on the first intermediate data to obtain a time-varying nonlinear strength coupling matrix, and then to sparsely reconstruct the correlation topology between each independent data dimension according to the strength coupling matrix to obtain the second intermediate data of dimensional interaction in the multidimensional health data stream.
[0089] The execution module 203 in this application is mainly used to obtain a derived data view of the current application scenario based on the second intermediate data and respond to the external query conditions of the current application scenario, and then use the derived data view to visualize the common feature patterns of the internal medicine patient group.
[0090] The foregoing has detailed examples of a multidimensional health data processing system and method for internal medicine patients provided in the embodiments of this application. It is understood that the corresponding apparatus, in order to achieve the above functions, includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specified application, but such implementation should not be considered beyond the scope of this application.
[0091] In some embodiments, this application also provides a computer device, the computer device including a memory and a processor, the memory for storing a computer program, and the processor for calling and running the computer program from the memory, causing the computer device to perform the above-described multidimensional health data processing method for internal medicine patients.
[0092] In some embodiments, reference Figure 4 The dashed lines in the figure indicate that the unit or module is optional. This figure is a schematic diagram of the structure of a computer device implementing a multidimensional health data processing method for internal medicine patients according to an embodiment of this application. The multidimensional health data processing method for internal medicine patients described in the above embodiments can be... Figure 4 The computer device shown is used to implement this, and the computer device includes at least one processor 301, a memory 302 and at least one communication unit 305. The computer device may be a terminal device, a server or a chip.
[0093] Processor 301 can be a general-purpose processor or a special-purpose processor. For example, processor 301 can be a central processing unit (CPU), which can be used to control computer devices, execute software programs, and process data from software programs. The computer device may also include a communication unit 305 for inputting (receiving) and outputting (transmitting) signals.
[0094] For example, the computer device may be a chip, and the communication unit 305 may be the input and / or output circuit of the chip, or the communication unit 305 may be the communication interface of the chip, which may be a component of a terminal device, network device or other device.
[0095] For example, the computer device may be a terminal device or a server, and the communication unit 305 may be a transceiver of the terminal device or the server, or the communication unit 305 may be a transceiver circuit of the terminal device or the server.
[0096] The computer device may include one or more memories 302 storing a program 304. The program 304 can be executed by a processor 301 to generate instructions 303, causing the processor 301 to execute the method described in the above method embodiments according to the instructions 303. Optionally, the memory 302 may also store data (such as a target audit model). Optionally, the processor 301 may also read data stored in the memory 302, which may be stored at the same storage address as the program 304, or it may be stored at a different storage address than the program 304.
[0097] The processor 301 and memory 302 can be configured separately or integrated together, for example, integrated on the system on chip (SOC) of the terminal device.
[0098] It should be understood that each step of the above method embodiment can be completed by hardware logic circuits or software instructions in the processor 301. The processor 301 can be a CPU, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.
[0099] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] For example, in some embodiments, this application also provides a computer-readable storage medium storing instructions or code that, when executed on a computer, cause the computer to implement the above-described multidimensional health data processing method for internal medicine patients.
[0101] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0102] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for processing multidimensional health data for internal medicine patients, characterized in that, Includes the following steps: Acquire a multidimensional health data stream of internal medicine patients, the multidimensional health data stream containing four mutually orthogonal independent data dimensions; The multidimensional health data stream is filtered by priority scores of each independent data dimension to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement. Based on the first intermediate data, the correlation degree of the orthogonal coupling relationship between each independent data dimension is learned to obtain a time-varying nonlinear strong coupling matrix. Then, the correlation topology between each independent data dimension is sparsely reconstructed according to the strong coupling matrix to obtain the second intermediate data of dimensional interaction in the multidimensional health data stream. The strong coupling matrix is a matrix used to characterize the dynamic correlation between each dimension within the multidimensional health data under a specified time window. Based on the second intermediate data, in response to the external query conditions of the current application scenario, a derived data view of the current application scenario is obtained, and then the derived data view is used to visualize the common feature patterns of the internal medicine patient group. Specifically, the correlation degree learning of the orthogonal coupling relationship between each independent data dimension based on the first intermediate data to obtain the time-varying nonlinear strength coupling matrix includes: Within the set sliding time window, extract the temporal feature vectors of each independent data dimension from the first intermediate data; The correlation strength between temporal feature vectors of any two independent data dimensions is calculated in parallel using a correlation calculation module based on mutual information and nonlinear kernel functions. Fill all the instantaneous correlation strengths into the corresponding positions of a fourth-order square matrix to obtain the time-varying nonlinear strength coupling matrix.
2. The method as described in claim 1, characterized in that, The multidimensional health data stream is dimensionally filtered by priority scores of each independent data dimension to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement, specifically including: Priority scores are calculated for each independent data dimension based on a pre-defined clinical scenario task; Based on the priority scores, the data channels of the corresponding independent data dimensions in the multidimensional health data stream are weighted, amplified, and suppressed and filtered to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement.
3. The method as described in claim 1, characterized in that, Based on the strength coupling matrix, the correlation topology between each independent data dimension is sparsely reconstructed to obtain the second intermediate data of dimensional interaction in the multidimensional health data stream, which specifically includes: Determine the sparsity association threshold for dimensional interactions in a multidimensional health data stream; Based on the sparsification correlation threshold, the strength coupling matrix is sparsified to obtain a sparsified correlation matrix. The sparsed correlation matrix is used as an adjacency matrix of a weighted directed graph to obtain the dimensional interaction graph of the multidimensional health data stream. Based on the dimensional interaction graph, each independent data dimension is mapped to a new low-dimensional potential space, thereby obtaining the second intermediate data of dimensional interaction in the multidimensional health data stream.
4. The method as described in claim 1, characterized in that, Based on the second intermediate data, in response to the external query conditions of the current application scenario, the derived data view of the current application scenario is obtained, specifically including: Obtain the external query conditions for the current application scenario, and then identify the clinical intent, data dimension focus, and expected granularity from each external query condition; Based on the clinical intent and data dimension focus, relevant dimensional subsets of data and interaction relationship subgraphs are selected from the second intermediate data; Based on the desired granularity and the interaction relationship subgraph, the dimensional subset data is subjected to time-series slicing to obtain an intermediate dataset that conforms to the query semantics; The intermediate dataset is encapsulated and formatted according to a predefined view template to obtain a derived data view that can be directly used for display in the current application scenario.
5. The method as described in claim 1, characterized in that, The visualization of common characteristic patterns of the internal medicine patient population using the derived data view specifically includes: Cluster analysis was performed on the derived data view to obtain multiple patient subgroups that exhibited similar patterns in the interaction relationships of the selected dimensions; Extract the dimensional feature contours of each patient subgroup, and then generate corresponding feature pattern labels; The characteristic pattern labels of each patient subgroup and their distribution proportion in the internal medicine patient population are combined and visualized.
6. The method as described in claim 1, characterized in that, The independent data dimensions include physiological, biochemical, behavioral, and subjective reporting dimensions.
7. A multidimensional health data processing system for internal medicine patients, used to execute the multidimensional health data processing method for internal medicine patients as described in any one of claims 1 to 6, wherein the multidimensional health data processing system for internal medicine patients includes a visualization unit, characterized in that, The visualization unit includes: The acquisition module is used to acquire a multidimensional health data stream of internal medicine patients, wherein the multidimensional health data stream contains four mutually orthogonal independent data dimensions; The processing module is used to perform dimensional filtering on the multidimensional health data stream by using the priority scores of each independent data dimension to obtain the first intermediate data for structural dimensionality reduction and semantic enhancement. The processing module is further configured to perform correlation learning on the orthogonal coupling relationship between each independent data dimension based on the first intermediate data to obtain a time-varying nonlinear strength coupling matrix, and then perform sparsification reconstruction on the correlation topology between each independent data dimension according to the strength coupling matrix to obtain the second intermediate data of dimensional interaction in the multidimensional health data stream. The execution module is used to obtain a derived data view of the current application scenario based on the second intermediate data and respond to the external query conditions of the current application scenario. Then, the derived data view is used to visualize the common feature patterns of the internal medicine patient group.
8. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing code, and the processor being configured to retrieve the code and execute the multidimensional health data processing method for internal medicine patients as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multidimensional health data processing method for internal medicine patients as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Statistical analysis method, device and equipment for clinical patients and storage medium
CN119004036A