A clinical data collection method and system based on artificial intelligence

Through the clinical data acquisition method based on artificial intelligence, the sampling ratio and frequency are dynamically adjusted, and the missing data is completed in combination with timing correlation, which solves the problems of insufficient data acquisition and insufficient sampling frequency in the existing technology, and achieves more efficient and accurate clinical data acquisition.

CN119851847BActive Publication Date: 2025-06-06SHENZHEN XIGUWEI BIOMEDICAL TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510329247.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-06
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

In the prior art, it is difficult to conduct unified standards analysis based on the diverse characteristics of different patients in clinical data collection, resulting in data confusion and inaccurate analysis results. In addition, the sampling frequency is difficult to adjust dynamically, resulting in missing key data during the high-frequency change period or collecting a large amount of irrelevant data during the low-frequency change period, increasing the burden of storage and analysis. In the completion of missing data, it is also difficult for the prior art to accurately correct the timing correlation of the data based on the prior art, which affects treatment decisions.

Method used

Using a clinical data collection method based on artificial intelligence, a multi-dimensional feature space is constructed by collecting patient physiological parameters and personal information, a multi-dimensional feature space is constructed, and a patient is clustered according to the characteristic distance, the data sampling ratio is initialized, and the screened patient feature clustering cluster is obtained. Then, the sampling ratio is dynamically adjusted to expand the coverage range, infer the probability value of patient status changes, adjust the sampling frequency, and generate a segmented sampling priority plan. Finally, data acquisition is carried out based on the segmented sampling priority plan, and the missing data is weighted and weighted mean interpolation method is completed.

Benefits of technology

By dynamically adjusting the sampling ratio and frequency, the accuracy and coverage of data acquisition are improved, the processing burden of redundant information is reduced, and the data resource allocation efficiency is optimized. Missing data completion based on timing correlation significantly improves the integrity and reliability of data acquisition results, supporting more accurate treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851847B_ABST
    Figure CN119851847B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of medical information management technology, specifically to a clinical data collection method and system based on artificial intelligence, comprising the following steps: collecting patient physiological parameters and corresponding personal information, constructing a multidimensional feature space with patient physiological parameters, clustering patients according to the feature distance between patients in the multidimensional feature space, initializing the data sampling ratio with reference to the information entropy value of the patient feature distribution in each cluster in the cluster grouping, and obtaining the screened patient feature cluster clusters. In the present invention, by normalizing the patient physiological parameters, different patient features can be compared under the same standard, thereby improving the accuracy of the analysis. Patient clustering based on feature distance classifies patient features with high similarity into the same cluster cluster, avoiding data confusion caused by individual feature differences, and the calculation of information entropy and the screening of low entropy cluster clusters further improve the pertinence of the collected data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical information management, and in particular to a clinical data collection method and system based on artificial intelligence. Background Art

[0002] Medical information management is the technology that uses information technology to collect, store, process, analyze and apply various types of information in the medical field. The field covers standardized management of medical data, electronic health records (EHR), medical system integration, telemedicine and intelligent diagnosis.

[0003] Among them, clinical data collection methods refer to the way of extracting and collecting relevant data from the patient's diagnosis and treatment process through technical means and standardized processes. Its main purpose is to obtain real and complete clinical information to provide support for medical research, disease management, treatment effect evaluation and medical system optimization.

[0004] When collecting patient data, the existing technology is difficult to perform unified standard analysis based on the diverse characteristics of different patients. For example, in parameters such as heart rate and blood pressure, the differences between different units or magnitudes can easily cause data confusion, resulting in inaccurate analysis results. When grouping patient feature data, the existing method cannot quantify the similarity through the distance between features, and it is difficult to achieve accurate patient grouping. In this case, it is easy to cause highly similar patients to belong to different groups, further affecting the effect of data screening. In addition, the existing technology mostly uses fixed frequency sampling when collecting data, and it is difficult to optimize the sampling frequency according to the dynamics of patient feature changes. For example, the time period of high-frequency changes may miss key data due to insufficient frequency, while the time period of low-frequency changes will collect a large amount of irrelevant data, increasing the burden of storage and analysis. For the missing data in the collection process, the existing technology mainly relies on simple interpolation or static completion methods, which is difficult to make accurate corrections based on the temporal correlation of the data. For example, in dynamic monitoring of patients, the lack of data in certain key time periods may lead to misjudgment of state fluctuations, thereby affecting treatment decisions. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a clinical data collection method and system based on artificial intelligence.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: a clinical data collection method based on artificial intelligence, comprising the following steps:

[0007] S1: Collect the physiological parameters and corresponding personal information of patients, construct a multidimensional feature space with the physiological parameters of patients, cluster the patients according to the feature distance between patients in the multidimensional feature space, initialize the data sampling ratio with reference to the information entropy value of the patient feature distribution in each cluster in the cluster grouping, and obtain the screened patient feature clusters;

[0008] S2: obtaining the number of patients in the screened patient feature cluster and the current sampling ratio, comparing the current data sampling ratio with the initialization data sampling ratio, updating the current data sampling ratio according to the comparison result to expand the coverage of the screened patient feature cluster, and generating a patient feature cluster adjustment result;

[0009] S3: obtaining the timeline information of each cluster in the patient feature cluster adjustment result, dividing the timeline into multiple time periods, inferring the probability value of the patient status change within the target time period, comparing the probability value with the preset change threshold, adjusting the sampling frequency according to the comparison result and the current data sampling ratio, and obtaining a segmented sampling priority plan;

[0010] S4: Collect clinical data according to the segmented sampling priority plan, assign weights to the missing data generated during the sampling process, complete the missing data according to the assigned weights, and generate patient clinical data collection results.

[0011] As a further solution of the present invention, the steps of clustering patients according to feature distance are specifically as follows:

[0012] S111: Acquire the patient's physiological parameter data, clean the collected data, remove abnormal values ​​and perform normalization processing, map to a unified range, map the normalized data to a multidimensional feature space, integrate the patient data point information in the multidimensional feature space, and obtain the physiological parameter mapping result;

[0013] S112: Based on the physiological parameter mapping result, the formula is used:

[0014] ;

[0015] Calculate the number of patients in the multidimensional feature space and patients The characteristic distance between ;

[0016] in, is the dimension index of the multidimensional feature space, is the total number of dimensions of the multidimensional feature space, Is a patient In the The normalized eigenvalues ​​of dimensions, Is a patient In the Normalized eigenvalues ​​of dimensions;

[0017] S113: constructing a characteristic distance matrix between patients based on the characteristic distance, and by comparing the characteristic distance between patients with a preset distance threshold, grouping patients with matching comparison results into the same group to obtain a patient clustering grouping result.

[0018] As a further solution of the present invention, the step of obtaining the patient feature cluster after screening is specifically as follows:

[0019] S121: According to the patient clustering grouping result, the formula is used:

[0020] ;

[0021] Calculate the information entropy value of the patient characteristics distribution in each cluster in the cluster grouping ;

[0022] in, ' is the total number of feature categories within the cluster, It is the first The probability of feature categories, yes The binary logarithm of The logarithmic weight of the feature class probability;

[0023] S122: Based on the information entropy value of each cluster, compare it with a preset information entropy threshold, initialize the data sampling ratio of the corresponding cluster according to the comparison result, and extract feature data in proportion to obtain the screened patient feature cluster.

[0024] As a further solution of the present invention, the step of obtaining the patient feature cluster adjustment result is specifically:

[0025] S211: clustering the patient characteristics according to the screened clusters, counting the number of patients in each cluster, and obtaining the current sampling ratio, comparing the current sampling ratio with the initial sampling ratio one by one, and adjusting according to the comparison results to generate an adjusted sampling ratio result;

[0026] S212: Based on the adjusted sampling ratio result, the updated ratio is applied to the data sampling of each cluster, covering all cluster data, and obtaining the patient characteristic cluster adjustment result.

[0027] As a further solution of the present invention, the step of obtaining the probability value of the patient's state change within the inferred target time period is specifically:

[0028] S311: Based on the patient characteristic cluster adjustment result, the time axis information is sorted in chronological order, the change range of the patient characteristic data at each time point is analyzed, the characteristic change of the time period is judged by the change range, and the time period division result is generated;

[0029] S312: Based on the time period division result, the formula is used:

[0030] ;

[0031] Calculate the given observation data Under the condition of Patients in the state Probability , get the probability value of the patient's state change;

[0032] in, Indicates that the current patient is in status When , the observed value is generated The possibility of Indicates the target time period Patients in the state The initial probability of Indicates that the patient is in the target time period The potential state within is the observed value, is the possible patient status, Is the state index variable, used to enumerate all possible states .

[0033] As a further solution of the present invention, the step of obtaining the segmented sampling priority plan is specifically as follows:

[0034] S321: Based on the probability value of the patient status change within the target time period, the probability value of each time period is compared with a preset change threshold, and the data sampling frequency is adjusted according to the comparison result and the change situation to obtain a sampling frequency plan corresponding to each time period;

[0035] S322: Based on the sampling frequency plan corresponding to each time period, all time periods are sorted according to priority to obtain a segmented sampling priority plan.

[0036] As a further solution of the present invention, the steps of obtaining the patient clinical data collection results are specifically as follows:

[0037] S411: Based on the segmented sampling priority plan, referring to the sampling frequency corresponding to the priority allocation, for the data missing situation occurring during the sampling process, determining the time point of the missing data and counting its distribution, assigning weights in combination with the importance of the time period and the number of surrounding sampling points, and obtaining a weight allocation result for each missing data;

[0038] S412: Based on the weight distribution result of each missing data, the adjacent data points in the time period where the missing data is located are screened, and the missing data are supplemented by analyzing the distribution trend and characteristic mean of the data points to generate the patient clinical data collection results.

[0039] A clinical data collection system based on artificial intelligence, the clinical data collection system based on artificial intelligence is used to execute the above-mentioned clinical data collection method based on artificial intelligence, the system comprising:

[0040] The data collection and grouping module collects patients' physiological parameters and personal information, constructs a multidimensional feature space, performs clustering and grouping according to the patient's feature distance, uses the information entropy value to determine the initial sampling ratio, and generates the screened patient feature clusters;

[0041] The data sampling optimization module compares the current sampling ratio with the initial sampling ratio based on the screened patient characteristic clusters, adjusts the sampling range, and generates a patient characteristic cluster adjustment result;

[0042] The state inference module divides the time axis into multiple time periods based on the patient feature clustering adjustment result, infers the probability of patient state change within the target time period and adjusts the sampling frequency to generate time period change priority data;

[0043] The sampling priority planning module optimizes the sampling frequency of the time period based on the time period change priority data, adjusts the sampling resources in combination with the patient distribution, and generates a segmented sampling priority plan;

[0044] The data completion and generation module performs segmented sampling based on the segmented sampling priority plan, uses the weighted mean interpolation method to complete the missing data, combines patient characteristics and credibility scores to improve accuracy, and generates patient clinical data collection results.

[0045] Compared with the prior art, the advantages and positive effects of the present invention are:

[0046] In the present invention, by normalizing the physiological parameters of patients, different patient characteristics can be compared under the same standard, thereby improving the accuracy of the analysis. Patient clustering based on feature distance classifies patient characteristics with high similarity into the same clustering cluster, avoiding data confusion caused by individual feature differences. The calculation of information entropy and the screening of low entropy clustering clusters further improve the pertinence of the collected data and reduce the processing burden of redundant information. The dynamic adjustment of the sampling ratio expands the coverage, solves the problem of feature omissions that may be caused by insufficient sampling data, and optimizes the efficiency of data resource allocation. The feature change amplitude analysis based on the time axis is combined with the hidden Markov model to infer the probability value of the patient's state change, realizing the accurate division of time periods, further focusing high-frequency sampling on high-probability change time periods, and low-frequency sampling on low-probability change time periods, thereby improving the validity of the data while reducing the amount of invalid data collected. Weight allocation and missing data completion significantly improve the integrity and reliability of data collection results by using associated data to correct the completion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of the workflow of the present invention;

[0048] Figure 2 A flow chart of clustering patients according to feature distances in the present invention;

[0049] Figure 3 A flow chart for obtaining the patient feature clusters after screening for the present invention;

[0050] Figure 4 A flow chart for obtaining the patient feature clustering cluster adjustment results for the present invention;

[0051] Figure 5 A flow chart of the present invention for inferring a probability value of a patient's state change within a target time period;

[0052] Figure 6 A flow chart for obtaining a segmented sampling priority plan for the present invention;

[0053] Figure 7 The present invention is a flow chart for obtaining the clinical data collection results of patients. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0055] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.

[0056] See also Figure 1 The present invention provides a technical solution: a clinical data collection method based on artificial intelligence, comprising the following steps:

[0057] S1: Collecting physiological parameters and corresponding personal information of patients, normalizing physiological parameters of clinical patients, constructing a multidimensional feature space with the processed physiological parameters of patients, wherein the dimensions in the multidimensional feature space represent the corresponding physiological parameters of patients respectively, calculating the characteristic distance between patients in the multidimensional feature space, clustering patients according to the characteristic distance, calculating the information entropy value of the characteristic distribution of patients in each cluster in the cluster grouping, marking clusters with information entropy values ​​lower than a preset information threshold, initializing the data sampling ratio of clusters lower than the preset information threshold, and obtaining the screened patient characteristic clusters;

[0058] S2: Obtain the number of patients in the screened patient feature clusters and the current sampling ratio, and compare the current data sampling ratio with the initialization data sampling ratio. If the amount of data sampled according to the current data sampling ratio is less than the amount of data sampled according to the initialization data sampling ratio, make multiple rounds of adjustments to the current data sampling ratio, update the current data sampling ratio to expand the coverage of the screened patient feature clusters, and generate patient feature cluster adjustment results;

[0059] S3: Obtain the timeline information of each cluster in the patient feature cluster adjustment result, analyze the change amplitude of the corresponding patient feature according to the timeline information, divide the timeline into time periods according to the change amplitude, infer the probability value of the patient status change in the target time period through the hidden Markov model, compare the probability value with the preset change threshold, and classify the probability value exceeding the preset change threshold as a high probability change, and if it does not exceed the threshold, classify it as a low probability change. For the time period with high probability change, increase the sampling frequency according to the current data sampling ratio, and for the time period with low probability change, reduce the sampling frequency according to the current data sampling ratio, and obtain a segmented sampling priority plan;

[0060] S4: Collect clinical data according to the segmented sampling priority plan, assign weights to the missing data generated during the sampling process, complete the missing data according to the assigned weights, and generate the patient clinical data collection results;

[0061] The screened patient feature clusters include the patient's multidimensional feature parameters, the corresponding cluster labels, the information entropy value within each cluster and the initialized data sampling ratio. The patient feature cluster adjustment results include the adjusted data sampling ratio, the updated patient number distribution, the expanded patient coverage and the coverage ratio of each cluster. The segmented sampling priority plan includes the sampling frequency within the target time period, the sampling density distribution of the high probability change time period, the sampling frequency reduction plan of the low probability change time period and the sampling priority division of each time period. The patient clinical data collection results include the complete patient clinical feature data set, the missing data weight distribution after completion and the sampling frequency information of the time period corresponding to each patient.

[0062] See also Figure 2 ,The specific steps of clustering patients according to feature distance are:

[0063] S111: Acquire the patient's physiological parameter data, clean the collected data, remove abnormal values ​​and perform normalization processing, map to a unified range, map the normalized data to a multidimensional feature space, integrate the patient data point information in the multidimensional feature space, and obtain the physiological parameter mapping result;

[0064] In clinical data collection, wearable devices or medical monitoring devices (such as dynamic electrocardiographs, sphygmomanometers, thermometers, etc.) are used to obtain the patient's physiological parameter data, including heart rate, systolic and diastolic blood pressure (blood pressure), body temperature and blood oxygen saturation. The monitoring equipment collects multiple data points regularly every day to ensure sufficient sample size, and combines the patient's personal information such as age, gender, and past medical history. The collected raw data may contain abnormal values, such as data points with heart rate exceeding the normal physiological range (40-200bpm), data points with blood pressure exceeding the normal range (systolic blood pressure 80-180mmHg, diastolic blood pressure 50-110mmHg), data points with body temperature below 35°C or above 40°C, and data points with blood oxygen saturation below 80% or above 100%. These abnormal values ​​may be caused by equipment failure, patient movement interference, or clinical operation errors. Abnormal data is eliminated by setting reasonable parameter thresholds. After data cleaning, all valid data are normalized. After normalization, all data are mapped to the [0,1] interval to eliminate the impact of different parameter dimensions on subsequent analysis. After processing, a multidimensional feature space is constructed. Each dimension in the feature space represents a physiological parameter, such as heart rate, blood pressure, and body temperature. Each patient corresponds to a point in the feature space, and the coordinates of the point are determined by the normalized physiological parameter values.

[0065] S112: Based on the physiological parameter mapping results, the formula is used:

[0066] ;

[0067] Calculate the number of patients in the multidimensional feature space and patients The characteristic distance between ;

[0068] in, is a scalar that indicates the overall size of the difference between the two patients in the feature space. The smaller the value, the smaller the difference. It is the dimension index of the multidimensional feature space, which is used to identify the dimension number in the current feature space. For example, heart rate is the first dimension, blood pressure is the second dimension, body temperature is the third dimension, etc. is the total number of dimensions in the multidimensional feature space, indicating the number of types of patient physiological parameters. For example, if heart rate, blood pressure, body temperature, and blood oxygen saturation are monitored, then , Is a patient In the The normalized eigenvalues ​​of dimensions, Is a patient In the The normalized eigenvalues ​​of the dimensions.

[0069] Assume that the normalized eigenvalues ​​of patients A and B are as follows: Patient A: heart rate (0.5), blood pressure (0.7), body temperature (0.6), Patient B: heart rate (0.6), blood pressure (0.8), body temperature (0.5), and the feature space dimension is . Calculate the feature distance between patient A and patient B:

[0070] ;

[0071] The results show that the difference between patient A and patient B in the multidimensional feature space is .

[0072] S113: constructing a characteristic distance matrix between patients based on the characteristic distances, and by comparing the characteristic distances between patients with a preset distance threshold, grouping patients with matching comparison results into the same group to obtain a patient clustering grouping result;

[0073] According to the calculated characteristic distances between patients, a distance matrix is ​​first constructed. , each element in the matrix Indicates patient and patients After completing the distance matrix construction, the hierarchical clustering method is used to cluster the patients. The specific steps are as follows: Set the initial distance threshold , whose value can be obtained by The average value of all distance values ​​in the matrix is ​​determined by statistical analysis, for example, it is set to the average value of all distance values ​​in the matrix According to the distance matrix The characteristic distance in the , compare the distance values ​​between the two patients in turn, and take the value less than or equal to Patients of are grouped into the same initial cluster. For example, if the distance between patient A and patient B is ,and , then patient A and patient B are divided into the same cluster. After traversing and comparing the distance values ​​of all patients, the preliminary cluster division is completed, and the number and distribution of patients in each cluster are recorded. The preliminary divided clusters are further merged and analyzed. If the minimum characteristic distance between two clusters meets the threshold , then merge the two clusters into a larger cluster, and repeat this process until there are no more clusters that meet the conditions to be merged. Output the final clustering grouping results. The characteristics of each group of patients have high similarity, and the characteristics of the groups are quite different.

[0074] See also Figure 3 , the specific steps for obtaining the patient feature cluster after screening are:

[0075] S121: According to the patient clustering results, the formula is used:

[0076] ;

[0077] Calculate the information entropy value of the patient characteristics distribution in each cluster in the cluster grouping ;

[0078] in, It indicates the randomness or uncertainty of the distribution of patient characteristics within the cluster. The lower the value, the more concentrated the distribution of patient characteristics within the cluster. ' is the total number of feature categories within the cluster, indicating the number of categories of patient features within the cluster, such as the range of distribution of feature values ​​such as heart rate and blood pressure, It is the first The probability of the feature category, indicating the Patients in each characteristic category The number of patients in the cluster accounts for the total number of patients in the cluster The ratio is calculated by the following formula: , yes The binary logarithm of The logarithmic weight of the probability of each feature category is used to calculate the contribution value of information entropy.

[0079] Suppose a cluster contains 10 patients, and their heart rate distribution is as follows: 60-80bpm interval: 6 patients; 80-100bpm interval: 3 patients; 100-120bpm interval: 1 patient.

[0080] Determine feature probabilities : ;

[0081] Calculate the logarithm of probability : , , .

[0082] Substitute the formula to calculate information entropy :

[0083] ;

[0084] The results show that the information entropy value calculated .

[0085] S122: comparing the information entropy value of each cluster with a preset information entropy threshold, initializing the data sampling ratio of the corresponding cluster according to the comparison result and extracting feature data in proportion to obtain the screened patient feature cluster;

[0086] According to the calculated cluster information entropy value , first determine the preset information entropy threshold , the threshold can be determined by calculating the average value of the information entropy of all clusters. For example, assuming there are 5 clusters, their information entropy values ​​are , , , , , then the information entropy threshold is calculated as: Information entropy threshold ,in, is the entropy index. Then, traverse the entropy values ​​of all clusters , the information entropy value is lower than The clusters of are marked as low information entropy clusters, for example, and Both are less than the information entropy threshold of 1.1, so cluster 2 and cluster 3 are marked as low information entropy clusters. Next, initialize the data sampling ratio of each low information entropy cluster , the data sampling ratio can be calculated by assigning weights to the inverse of the number of patients. For example, assuming that the number of patients in cluster 2 is 20 people, the number of patients in cluster 3 If there are 10 people, the sampling ratios are: , . Set the sampling ratio and Applied to their respective clusters, patient characteristic data are randomly extracted in proportion to form a new sample set. The screened patient characteristic clusters are obtained by statistically analyzing the final sampling results, among which the proportion of samples in the low information entropy clusters is higher, forming the final screened cluster data set.

[0087] See also Figure 4 , the specific steps for obtaining the patient characteristic cluster adjustment results are:

[0088] S211: clustering the patients according to their characteristics after screening, counting the number of patients in each cluster, and obtaining the current sampling ratio, comparing the current sampling ratio with the initial sampling ratio one by one, and adjusting the ratio according to the comparison result to generate an adjusted sampling ratio result;

[0089] According to the clustering of the screened patient characteristics, each cluster is first counted. Number of patients in , and obtain the current data sampling ratio of each cluster , for example, a cluster Contains 30 patients, with an initial sampling ratio of , the current sampling ratio is Compare the current data sampling ratio with the initialization data sampling ratio one by one to calculate the current sampled data volume and the amount of data initially sampled ,in , Assumptions , then the current sampling data volume , the initial sampling data volume .if , that is, the current sampling ratio cannot meet the initial sampling ratio requirement, then the current data sampling ratio needs to be adjusted. The adjustment method can be gradually increased The value of , for example, the step size is set to 0.01, after adjustment The current data sampling ratio is adjusted and the new sampling data volume is calculated again, and the adjustment process is repeated until the current data sampling ratio is adjusted. until.

[0090] S212: Based on the adjusted sampling ratio result, the updated ratio is applied to the data sampling of each cluster, covering all cluster data, and obtaining the patient characteristic cluster adjustment result;

[0091] Update each cluster After the current data sampling ratio is set, check the adjusted data sampling ratio , which is applied to the data sampling process of the corresponding cluster, for example, a cluster The adjusted data sampling ratio is , by calculating the adjusted sampling data volume . Assume that the number of patients in the cluster is , then the adjusted sampling data volume After confirming that the amount of sampled data meets the requirements of the initial sampling ratio, the sampled data set of each cluster is updated and stored. Through multiple rounds of adjustments, all data of each cluster is covered, the coverage of the patient characteristic cluster is expanded, and the final adjusted cluster data result set is formed, and the results are output for subsequent analysis.

[0092] See also Figure 5 , the specific steps for obtaining the probability value of the patient's state change within the target time period are as follows:

[0093] S311: Based on the patient characteristic cluster adjustment result, the time axis information is sorted in chronological order, the change range of the patient characteristic data at each time point is analyzed, the characteristic change of the time period is judged by the change range, and the time period division result is generated;

[0094] By timestamping the characteristic data of patients in the cluster, the time axis information is arranged in chronological order to ensure that the characteristic data at each time point is complete and available, and the changes of each characteristic data on the time axis are analyzed. For example, for the characteristic data of patients in the cluster, such as heart rate, blood pressure, body temperature, etc., the average value of each characteristic is counted according to the time point, and the data at each time point is compared with the data at the previous time point, so as to analyze the characteristic change amplitude of the corresponding time point. By calculating the change value of each time point, the time point with significant characteristic value change and the time point with gentle change are identified, and the time axis is divided according to the magnitude of the change amplitude, and the judgment condition of the change amplitude is set, and the continuous high change time points in the time axis are divided into one time period, and the continuous low change time points are divided into another time period. For example, for a time axis within five consecutive minutes, the change amplitude of the heart rate characteristic data is large and the blood pressure change is also significant, then these five minutes can be divided into a high change time period, and the characteristic changes of the remaining time points are small, which are divided into low change time periods. After completing the segmentation of the time axis, a set of characteristic change information associated with the time period is formed.

[0095] S312: Based on the time period division results, the formula is used:

[0096] ;

[0097] Calculate the given observation data Under the condition of Patients in the state Probability , get the probability value of the patient's state change;

[0098] in, Indicates that the current patient is in status When , the observed value is generated The possibility of And the corresponding observed value distribution is determined. For example, based on the historical observation data of the patient's heart rate changes, the frequency of observed values ​​under various states (such as stable, fluctuating, and violently fluctuating) is counted to determine the corresponding probability value. Indicates the target time period Patients in the state The initial probability is calculated by counting the state distribution in the historical data. For example, in the monitoring data of the past 30 days, the frequency distribution of patients in different states (such as stable, fluctuating, and violently fluctuating) is counted. Represents the observed value The weighted sum of probabilities appearing in all possible states is used to ensure the final probability value Between 0 and 1, Indicates that the patient is in the target time period The potential state of the inner, such as stability, fluctuation or violent fluctuation, is used as the hidden variable in the hidden Markov model, and the specific state is inferred through the model. is the observed value, indicating the target time period The patient characteristic change data observed during the target time period specifically reflects the change in patient characteristics (such as heart rate, blood pressure) during the target time period. is a possible patient state, representing a state in the set of all possible states, e.g. It may be any of stable, fluctuating, and violently fluctuating, which is determined by state definition and historical data annotation. For example, the heart rate fluctuation range is divided into different state categories. Is the state index variable, used to enumerate all possible states , Indicates the index of the target time period.

[0099] Suppose a cluster The target time period is 1 to 3 minutes, and the observed value is the average value of the patient's heart rate change, the possible state set include: , assuming a priori probability: , , . Conditional probability of observations: , , .

[0100] Calculate the denominator:

[0101] ;

[0102] Calculate the conditional probability:

[0103] ;

[0104] ;

[0105] ;

[0106] The final probability value is: , , .

[0107] The results show that within the target time period, the probability of the patient being in a "stable" state is the highest, at 73.17%, the probability of being in a "fluctuating" state is 21.95%, and the probability of being in a "severely fluctuating" state is lower, at only 4.88%.

[0108] See also Figure 6 ,The specific steps for obtaining the segmented sampling priority plan are:

[0109] S321: Based on the probability value of the patient status change within the target time period, the probability value of each time period is compared with a preset change threshold, and the data sampling frequency is adjusted according to the comparison result and the change situation to obtain a sampling frequency plan corresponding to each time period;

[0110] Compare the corresponding probability value in each time period with the preset change threshold to determine the nature of the change in the time period. The preset change threshold can be set by analyzing the change rules of historical data. For example, the threshold is calculated based on the average probability value of the patient's status change. Assuming the change threshold is 0.6, when the probability value of the time period is higher than 0.6, the time period is divided into a high probability change time period, and when it is lower than 0.6, it is divided into a low probability change time period. Subsequently, the data sampling frequency is adjusted according to the nature of the change in the time period. For the high probability change time period, the sampling frequency is increased in combination with the data sampling ratio, for example, the original sampling frequency is increased by 1.5 times, and the number of sampling points is increased at the same time to ensure that more data is collected in the high change time period; for the low probability change time period, the sampling frequency is reduced in combination with the data sampling ratio, for example, the original sampling frequency is reduced to 70% of the original frequency, and the number of sampling points is reduced, thereby optimizing data collection resources. After completing the sampling frequency adjustment of all time periods, the sampling frequency plan corresponding to each time period is output.

[0111] S322: based on the sampling frequency plan corresponding to each time period, all time periods are sorted according to priority to obtain a segmented sampling priority plan;

[0112] The integrated sorting results in a segmented sampling priority plan. According to the adjusted sampling frequency of the time period, all time periods are sorted by priority. The time period with high probability of change has the highest priority, and the time period with low probability of change has the lowest priority. The data sampling plans for each time period are integrated to form a complete segmented sampling priority plan. The data collection needs of the time period with high probability of change are given priority, and sampling tasks in the time period with low probability of change are allocated resources with a lower priority. The integrated plan is sorted in chronological order, and finally the sampling frequency and priority of each time period are indicated.

[0113] See also Figure 7 , the specific steps for obtaining the patient's clinical data collection results are:

[0114] S411: Based on the segmented sampling priority plan, referring to the sampling frequency corresponding to the priority allocation, for the data missing situation occurring during the sampling process, the time point of the missing data is determined and its distribution is counted, and the weight is allocated in combination with the importance of the time period and the number of surrounding sampling points to obtain the weight allocation result for each missing data;

[0115] According to the adjustment result of the sampling frequency of the time period, the data of each time period is collected according to the priority. The priority is determined by the relationship between the probability value of the state change of the target time period and the change threshold. If the probability value is higher than the change threshold, the priority of the time period is judged to be high; otherwise, it is low. The setting basis of the change threshold is the mean value of the patient's state change in the historical data or according to the actual clinical needs, such as analyzing the mean value of the state change of all patients in the past month as the threshold. Specifically, the time period with a high probability value often indicates that the patient's characteristics change drastically or the state fluctuates significantly. Such time periods are given high priority and high-frequency sampling is used to ensure data integrity; while the time period with a low probability value indicates that the patient's characteristics change steadily or the state does not fluctuate significantly, and is given low priority. Low-frequency sampling is used to save collection resources. During the collection process, multiple groups of data points are collected per second for the high-priority time period to ensure the continuity and integrity of the data; the low-priority time period adopts a low-frequency sampling method to collect data points at longer intervals to optimize the allocation of collection resources. During the collection process, due to the possible data loss caused by sensor failure or external interference, the time point of the missing data is first determined, and the sampling situation at each time point is counted. According to the distribution of missing data and sampling frequency, weights are assigned to missing data. The weights are assigned based on the importance of the time period and the number of sampled data points around the missing data. For example, a higher weight is assigned to missing data in a high-priority time period, while a lower weight is assigned to missing data in a low-priority time period. The weight value of missing data can be determined by calculating the time distance between the missing point and the data points before and after it, as well as the sampling frequency of adjacent data points.

[0116] S412: Based on the weight distribution result of each missing data, the adjacent data points in the time period where the missing data is located are screened, and the missing data are supplemented by analyzing the distribution trend and characteristic mean of the data points to generate the patient clinical data collection results;

[0117] First, the neighboring data points of the time period where the missing data is located are screened, the characteristic mean or trend change of these data points is calculated, and the completion result is adjusted by the weight distribution ratio. For example, for missing data in a high-priority time period, the initial completion value is calculated by linear interpolation of the two groups of data points before and after, and then the interpolation result is corrected according to the weight value. The time period with a higher weight tends to use more data points for interpolation fitting, while the time period with a lower weight directly uses the mean method to complete the data. In the priority judgment process, the influence of neighboring data points on the completion value also needs to be corrected in combination with the importance of the time period. For example, the completion of missing data in high-priority time periods tends to ensure its accuracy through trend analysis combined with the completion value, while the low-priority time periods are given priority to use simple linear methods for rapid completion. After the data completion of all time periods is completed, the consistency of the completed data is checked to ensure that the trend and distribution of the completed data are consistent with the original data, and the final patient clinical data collection results are generated and output as a complete time series data set.

[0118] A clinical data collection system based on artificial intelligence, the clinical data collection system based on artificial intelligence is used to execute the above-mentioned clinical data collection method based on artificial intelligence, the system comprises:

[0119] The data collection and grouping module collects patients' physiological parameters and personal information, constructs a multidimensional feature space, performs clustering and grouping according to the patient's feature distance, uses the information entropy value to determine the initial sampling ratio, and generates the screened patient feature clusters;

[0120] The data sampling optimization module compares the current sampling ratio with the initial sampling ratio based on the screened patient characteristic clusters, adjusts the sampling range, and generates patient characteristic cluster adjustment results;

[0121] The state inference module divides the time axis into multiple time periods based on the cluster adjustment results of patient characteristics, infers the probability of patient state change within the target time period, adjusts the sampling frequency, and generates time period change priority data;

[0122] The sampling priority planning module optimizes the sampling frequency of each time period based on the priority data of time period changes, adjusts the sampling resources according to the patient distribution, and generates a segmented sampling priority plan;

[0123] The data completion and generation module performs segmented sampling based on the segmented sampling priority plan, uses the weighted mean interpolation method to complete the missing data, combines patient characteristics and credibility scores to improve accuracy, and generates patient clinical data collection results.

[0124] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. A clinical data collection method based on artificial intelligence, characterized in that: The following steps are involved: Collect the patient's physiological parameters and corresponding personal information, construct a multidimensional feature space with the patient's physiological parameters, cluster the patients according to the feature distance between the patients in the multidimensional feature space, initialize the data sampling ratio with reference to the information entropy value of the patient's feature distribution in each cluster in the cluster grouping, and obtain the screened patient feature clusters; Obtaining the number of patients and the current sampling ratio in the screened patient feature clustering cluster, comparing the current data sampling ratio with the initialization data sampling ratio, updating the current data sampling ratio according to the comparison result to expand the coverage of the screened patient feature clustering cluster, and generating a patient feature clustering cluster adjustment result; Obtaining the timeline information of each cluster in the patient feature cluster adjustment result, dividing the timeline into multiple time periods, inferring the probability value of the patient status change within the target time period, comparing the probability value with a preset change threshold, adjusting the sampling frequency according to the comparison result and the current data sampling ratio, and obtaining a segmented sampling priority plan; The steps for obtaining the probability value of the patient's state change within the inferred target time period are specifically as follows: Based on the patient characteristic cluster adjustment result, the time axis information is sorted in chronological order, the change range of the patient characteristic data at each time point is analyzed, the characteristic change of the time period is judged by the change range, and the time period division result is generated; Based on the time period division results, the formula is used: ; Calculate the given observation data Under the condition of Patients in the state Probability , get the probability value of the patient's state change; in, Indicates that the current patient is in status When , the observed value is generated The possibility of Indicates the target time period Patients in the state The initial probability of Indicates that the patient is in the target time period The potential state within is the observed value, is the possible patient status, Is the state index variable, used to enumerate all possible states ; Clinical data collection is performed according to the segmented sampling priority plan, weights are assigned to missing data generated during the sampling process, missing data are completed according to the assigned weights, and patient clinical data collection results are generated.

2. The artificial intelligence-based clinical data collection method according to claim 1, characterized in that: The specific steps for clustering patients according to feature distance are as follows: Acquire the patient's physiological parameter data, clean the collected data, remove outliers and perform normalization, map to a unified range, map the normalized data to a multidimensional feature space, integrate the patient data point information in the multidimensional feature space, and obtain the physiological parameter mapping result; Based on the physiological parameter mapping results, the formula is adopted: ; Calculate the number of patients in the multidimensional feature space and patients The characteristic distance between ; in, is the dimension index of the multidimensional feature space, is the total number of dimensions of the multidimensional feature space, Is a patient In the The normalized eigenvalues ​​of dimensions, Is a patient In the Normalized eigenvalues ​​of dimensions; A characteristic distance matrix between patients is constructed based on the characteristic distances, and patients with matching comparison results are classified into the same group by comparing the characteristic distances between patients with a preset distance threshold, thereby obtaining a patient clustering grouping result.

3. The artificial intelligence-based clinical data collection method according to claim 2, characterized in that: The steps for obtaining the patient feature cluster after screening are specifically as follows: According to the patient clustering grouping results, the formula is used: ; Calculate the information entropy value of the patient characteristics distribution in each cluster in the cluster grouping ; in, ' is the total number of feature categories within the cluster, It is the first The probability of feature categories, yes The binary logarithm of The logarithmic weight of the feature class probability; Based on the information entropy value of each cluster, it is compared with a preset information entropy threshold, the data sampling ratio of the corresponding cluster is initialized according to the comparison result, and the feature data is extracted proportionally to obtain the screened patient feature cluster.

4. The artificial intelligence-based clinical data collection method according to claim 3, characterized in that: The steps for obtaining the patient characteristic cluster adjustment result are specifically as follows: According to the screened patient feature clustering clusters, the number of patients in each cluster is counted, and the current sampling ratio is obtained, the current sampling ratio is compared with the initial sampling ratio one by one, and adjustments are made according to the comparison results to generate an adjusted sampling ratio result; Based on the adjusted sampling ratio result, the updated ratio is applied to the data sampling of each cluster cluster, covering all cluster cluster data, and obtaining the patient characteristic cluster cluster adjustment result.

5. The artificial intelligence-based clinical data collection method according to claim 4, characterized in that: The steps for obtaining the segmented sampling priority plan are specifically as follows: Based on the probability value of the patient's state change within the target time period, the probability value of each time period is compared with a preset change threshold, and the data sampling frequency is adjusted according to the comparison result and the change situation to obtain a sampling frequency plan corresponding to each time period; Based on the sampling frequency plan corresponding to each time period, all time periods are sorted according to priority to obtain a segmented sampling priority plan.

6. The artificial intelligence-based clinical data collection method according to claim 5, characterized in that: The steps for obtaining the patient's clinical data collection results are specifically as follows: Based on the segmented sampling priority plan, referring to the sampling frequency corresponding to the priority allocation, for the data missing situation occurring during the sampling process, the time point of the missing data is determined and its distribution is counted, and the weight is allocated in combination with the importance of the time period and the number of surrounding sampling points to obtain the weight allocation result for each missing data; Based on the weight distribution result of each missing data, the adjacent data points in the time period where the missing data is located are screened, and the missing data are supplemented by analyzing the distribution trend and characteristic mean of the data points to generate the patient's clinical data collection results.

7. A clinical data collection system based on artificial intelligence, characterized in that: According to any one of claims 1 to 6, the method for collecting clinical data based on artificial intelligence comprises: The data collection and grouping module collects patients' physiological parameters and personal information, constructs a multidimensional feature space, performs clustering and grouping according to the patient's feature distance, uses the information entropy value to determine the initial sampling ratio, and generates the screened patient feature clusters; The data sampling optimization module compares the current sampling ratio with the initial sampling ratio based on the screened patient characteristic clustering clusters, adjusts the sampling range, and generates a patient characteristic clustering cluster adjustment result; The state inference module divides the time axis into multiple time periods based on the patient feature clustering adjustment result, infers the probability of patient state change within the target time period and adjusts the sampling frequency to generate time period change priority data; The sampling priority planning module optimizes the sampling frequency of the time period based on the time period change priority data, adjusts the sampling resources in combination with the patient distribution, and generates a segmented sampling priority plan; The data completion and generation module performs segmented sampling based on the segmented sampling priority plan, uses the weighted mean interpolation method to complete the missing data, combines patient characteristics and credibility scores to improve accuracy, and generates patient clinical data collection results.

Citation Information

Patent Citations

  • Method for evaluating acute kidney injury induced by rhabdomyolysis

    CN118866388A