Intelligent senile syndrome scientific research data management method and system
By integrating and standardizing medical data for geriatric syndrome, building proprietary field ontology and training high-risk prediction models, the problem of multi-source data integration is solved, research efficiency and prediction accuracy are improved, and interpretability analysis support is provided.
Patent Information
- Application Number
- CN202510756696.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The prior art is difficult to effectively integrate and analyze multi-source heterogeneous medical data for geriatric syndrome, resulting in long research cycles, insufficient prediction accuracy, and lack of interpretability analysis.
By integrating structured and unstructured data from hospital information systems, a proprietary field ontology of geriatric syndrome is constructed, standardized term data is generated, high-risk prediction models are trained using random forest models, and significance analysis and variable contribution measurement are performed.
It has realized efficient integration and standardization of multi-source medical data, quickly constructed a scientific research cohort, improved the efficiency and prediction accuracy of geriatric syndrome research, and provided interpretability analysis support.
Smart Images

Figure CN120260940A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing. Specifically, it particularly relates to an intelligent scientific research data management method and system for geriatric syndromes. Background Art
[0002] In the current environment, the prevention and treatment of geriatric syndromes (such as sarcopenia, cognitive impairment) have become the core challenges in the field of medical health. Although hospital information systems (HIS, EMR, etc.) have accumulated a large amount of patient data, due to multi-source heterogeneity (the mixture of structured and unstructured data), inconsistent terminology, and the lack of temporal characteristics, it is difficult to directly use them for scientific research and clinical decision-making. Existing technologies rely on manual data cleaning and static statistical models, resulting in a long research cycle and insufficient prediction accuracy. There is still a need for intelligent tools to achieve data fusion, standardization, dynamic risk prediction, and interpretability analysis to improve the efficiency of geriatric medicine research and the accuracy of clinical intervention.
[0003] There are problems in the prior art such as medical research data islands, insufficient processing of unstructured data, inefficient research processes, and poor domain adaptability. Summary of the Invention
[0004] (I) Technical Problems to be Solved In view of the problems in the related art, the present invention provides an intelligent scientific research data management method for geriatric syndromes to overcome the above-mentioned technical problems existing in the related prior art.
[0005] (II) Technical Solutions To solve the above technical problems, the present invention is implemented through the following technical solutions: S1. Integrate the structured and unstructured data of the hospital to obtain multi-source patient structured data and temporal feature data; S2. Construct an ontology for the specific field of geriatric syndromes; based on the ontology for the specific field of geriatric syndromes, perform data cleaning and normalization on the multi-source patient structured data to obtain patient standardized term data; S3. According to the patient standardized term data, temporal feature data, and combined with the requirements of scientific research variables, construct a disease-specific data matrix; Based on the case inclusion rules and the disease-specific data matrix, obtain the target research medical record cohort data; S4. Use historical medical record data combined with an optimization algorithm to train and optimize the initial random forest model to obtain a high-risk medical record prediction model; Input the target research medical record cohort data into the high-risk medical record prediction model to obtain real-time high-risk cases; S5. Perform a significance analysis on the disease-specific data matrix to obtain a significance analysis result; perform an interpretable AI analysis on the real-time high-risk cases to obtain variable contribution degree data; Generate a scientific research analysis report based on the significance analysis results in combination with the variable contribution degree data; Through data processing, analysis and deep learning technology, the present invention realizes the efficient integration and standardization of multi-source medical data, the rapid construction of a scientific research cohort and the training of a high-precision prediction model, as well as the interpretable quantitative analysis of risk factors, thus solving the problem of data islands, improving the research efficiency of geriatric syndromes, and providing a scientific basis with both statistical significance and actual contribution degree for clinical decision-making.
[0006] Preferably, the S1 includes the following steps: S11. Set a hospital information system set, and the hospital information system set includes each information system of the hospital; Connect to each information system in the hospital information system set through a standardized interface, extract the structured data of patients, and obtain the structured data of patients; Perform format conversion on the structured data of patients to obtain the structured data of patients with a unified format; S12. Use natural language processing technology to parse the text data in the unstructured data, and extract key entities and temporal relationships to obtain structured entity data; Scan the picture and paper data in the unstructured data through OCR technology to obtain the original text; call the AI semantic model to perform secondary cleaning and structured processing on the original text to obtain structured field data; The structured data of patients with a unified format, the structured entity data, and the structured field data together constitute the multi-source structured data of patients; S13. Integrate the device data in the unstructured data, store the dynamic monitoring indicators through a time series database, and associate them with the patient ID to obtain time series feature data; The present invention integrates the structured data of each system in the hospital through a standardized interface and uniformly converts it into the JSON / XML format; uses NLP to extract key entities and temporal relationships in the unstructured text, combines OCR and AI models to process image / paper data to generate structured fields; finally, associates the device monitoring indicators with the patient ID through a time series database to form multi-source structured data and time series feature data, solving the problem of data heterogeneity.
[0007] Preferably, the S2 includes the following steps: S21. Construct an ontology for the specific field of geriatric syndromes; S22. Combine the ontology for the specific field of geriatric syndromes to perform data cleaning and normalization on the multi-source structured data of patients to obtain the standardized term data of patients; The present invention constructs an ontology for geriatric syndromes and combines logical rules and edit distance algorithms to clean multi-source data, realizing the standardization of terms and units.
[0008] Preferably, the S21 includes the following steps: S211. Define the core data elements of geriatric syndromes in the ontology of the specialized field of geriatric syndromes; S212. Define the relationships between entities in the ontology of the specialized field of geriatric syndromes through the knowledge graph to obtain logical constraint rules; The logical constraint rules include temporal logic rules, numerical logic rules, and semantic logic rules; The present invention realizes data standardization and logical verification by defining the core data elements of geriatric syndromes and establishing logical constraint rules between entities based on the knowledge graph.
[0009] Preferably, the S22 includes the following steps: S221. Use the temporal logic rules of the ontology of the specialized field of geriatric syndromes to detect the data that does not conform to the temporal logic in the multi-source patient structured data, and trigger automatic correction or manual review; S222. Use the numerical logic rules of the ontology of the specialized field of geriatric syndromes to detect the data that does not conform to the numerical logic in the multi-source patient structured data, and trigger automatic correction or manual review; S223. Use the semantic logic rules of the ontology of the specialized field of geriatric syndromes and combine with the edit distance algorithm to calculate the semantic similarity, and convert the non-standard terms and units in the multi-source patient structured data into standard semantics; S224. Obtain the standardized term data through steps S221, S223, and S223; The present invention detects multi-source data anomalies through temporal, numerical, and semantic logic rules, combines with the edit distance algorithm to convert non-standard terms, and finally generates standardized term data.
[0010] Preferably, the S3 includes the following steps: S31. According to the patient standardized term data and temporal feature data, combine with the research requirement variables, and generate an electronic case report form by dragging variables; S32. Upload the electronic case report form to the Excel template and automatically map it to the system fields to obtain the disease-specific data matrix; S33. Set the case inclusion rules; based on the case inclusion rules and the disease-specific data matrix, use Boolean logic to screen cases and target the research case cohort; The present invention dynamically generates an electronic case report form by dragging variables, automatically maps fields to generate a disease-specific data matrix, and combines with Boolean logic to quickly screen the research case cohort that meets the inclusion rules, improving the research efficiency.
[0011] Preferably, the S4 includes the following steps: S41. Construct an initial random forest model, set the number of trees in the initial random forest model; set the training accuracy of the initial random forest model and the training accuracy threshold. S42. Collect historical medical record data, where the historical case data includes the physical data of each patient and a high-risk label. S43. Use the historical medical record data to train the initial random forest model. During the training process, use an optimization algorithm to find the number of trees in the initial random forest model, and judge the convergence of the initial random forest model according to the training accuracy and the training accuracy threshold to obtain the optimal solution. Take the optimal solution as the number of trees in the initial random forest model to obtain a high-risk medical record prediction model. S44. Input the target research medical record cohort data into the high-risk medical record prediction model to obtain real-time high-risk cases. By constructing an initial random forest model, training based on historical medical record data, dynamically adjusting the number of trees parameter using an optimization algorithm, combining convergence judgment to generate a high-precision prediction model, and finally outputting high-risk cases of the target cohort in real time, the present invention improves the prediction accuracy and clinical decision-making efficiency.
[0012] Preferably, in S43, using an optimization algorithm to find the number of trees in the initial random forest model, and judging the convergence of the initial random forest model according to the training accuracy and the training accuracy threshold to obtain the optimal solution includes the following steps: S431. Construct a wild horse population, set the size of the wild horse population; set the maximum number of optimization iterations. S432. According to the number of trees in the initial random forest, randomly set the initial positions of the wild horse population to obtain an initial set of the wild horse population. S433. According to the training accuracy of the initial random forest and the training accuracy threshold, define the fitness function of the initial positions of the wild horses in the wild horse population. S434. Perform iterative operations on the initial position set of the wild horse population. In each round of iteration, calculate the fitness value of each position in the initial position set of the wild horse population according to the fitness function, update the position concentration of each wild horse in the initial position set of the wild horse population from high to low according to the fitness value, and obtain the best wild horse individual position and the global best wild horse position in the wild horse population in each round of iteration. S435. Repeat S434. When the maximum number of optimization iterations is reached, stop the iteration and take the global best wild horse position as the optimal solution. By simulating the behavior of the wild horse population, iteratively updating the individual position concentration based on the fitness function, dynamically adjusting the number of trees parameter, and finally selecting the global optimal solution as the number of trees in the random forest model, the present invention realizes the automatic optimization of the model hyperparameters and improves the prediction accuracy and convergence efficiency.
[0013] Preferably, S5 includes the following steps: S51. Divide the data in the disease-specific data matrix into continuous data and categorical data according to the data type; S52. Use the t-test algorithm to perform a significance analysis on the continuous data to obtain the analysis result of the continuous data; S53. Use the chi-square test to perform a significance analysis on the categorical data to obtain the analysis result of the categorical data; S54. The analysis result of the continuous data and the analysis result of the categorical data together constitute the significance analysis result; S55. Calculate the variable contribution degree of real-time high-risk cases through the SHAP value to obtain variable contribution degree data; S56. Generate a scientific research analysis report according to the significance analysis result combined with the variable contribution degree data; The present invention generates a significance analysis result by analyzing continuous data through the t-test and categorical data through the chi-square test, quantifies the variable contribution degree by combining the SHAP value, and finally fuses statistical significance and clinical impact weight to generate a scientific research report, providing a decision support basis with both statistical significance and interpretability for the research of geriatric syndromes.
[0014] An intelligent scientific research data management system for geriatric syndromes, used to implement the above-mentioned intelligent scientific research data management method for geriatric syndromes, includes a multi-source medical data integration and feature extraction module, a domain ontology construction and data standardization module, a disease-specific data matrix construction and case screening module, a high-risk prediction model training and optimization module, and a scientific research analysis and report generation module; The multi-source medical data integration and feature extraction module docks with the hospital information system through a standardized interface, integrates the structured data of patients, and uses natural language processing, OCR technology, and a time series database to parse the unstructured data, extracts key entities and time series features, and finally generates multi-source patient structured data and time series feature data with a unified format, providing a standardized input for subsequent analysis; The domain ontology construction and data standardization module constructs a proprietary domain ontology based on the core data elements and knowledge graph of geriatric syndromes, defines the logical constraint rules between entities; through time logic detection, numerical logic correction, and semantic logic mapping, combined with the edit distance algorithm, it converts multi-source data into standardized term data to ensure the consistency and scientific research availability of the data; The disease-specific data matrix construction and case screening module dynamically generates an electronic case report form according to the research requirement variables, constructs a disease-specific data matrix through Excel template mapping, and uses Boolean logic to screen the target research cohort in combination with the case inclusion rules; it converts the standardized data into a research-oriented structured data set to support subsequent statistical analysis and model training; The high - risk prediction model training and optimization module uses the random forest model as the basic model, combines the wild horse optimization algorithm to dynamically adjust the number of tree parameters, trains the model through historical medical record data and evaluates the convergence, and finally generates a high - risk medical record prediction model; the high - risk medical record prediction model realizes the accurate prediction of real - time high - risk cases in the target cohort through feature selection and hyperparameter optimization; The scientific research analysis and report generation module conducts significance analysis and interpretability analysis on the disease - specific data matrix, identifies key variables, and generates a scientific research analysis report by combining statistical results and variable contribution degrees; it reveals the internal relationship of data through quantitative analysis, providing statistical evidence and clinical decision - making support for the research of geriatric syndromes.
[0015] (III) Beneficial effects The present invention has the following beneficial effects: Through data processing and analysis combined with deep - learning technology, the present invention realizes the efficient integration and standardization of multi - source medical data, the rapid construction of a scientific research cohort and the training of a high - precision prediction model, as well as the interpretable quantitative analysis of risk factors, thus solving the problem of data islands, improving the efficiency of geriatric syndrome research, and providing a scientific basis with both statistical significance and actual contribution degree for clinical decision - making.
[0016] The present invention integrates the data of systems such as HIS and EMR through a standardized interface, combines NLP and OCR technologies to parse unstructured texts, and uses a time - series database to store dynamic monitoring indicators to solve the problem of medical data islands; the logical rules of the geriatric syndrome - specific domain ontology and the edit - distance algorithm achieve the standardization of domain - specific terms, improve data quality and consistency, and provide a reliable basis for data analysis.
[0017] Based on the disease - specific data matrix generated by dragging variables and the Boolean logic enrollment rules, the present invention quickly screens target cases, shortens the scientific research data preparation cycle; combines the wild horse optimization algorithm to dynamically adjust the number of trees in the random forest, and optimizes the random forest model parameters through the fitness function, significantly improving the accuracy and convergence efficiency of high - risk prediction.
[0018] The present invention quantifies the variable contribution degree through SHAP values, combines the significance results of t - tests and chi - square tests, and generates a multi - dimensional scientific research report; the scientific research report not only reveals the statistical significance of the data, but also clarifies the influence weights of clinical variables, providing a decision - making basis with both statistical and clinical interpretability for the mechanism research and risk intervention of geriatric syndromes.
[0019] Of course, it is not necessary for any product implementing the present invention to achieve all the above - mentioned advantages simultaneously. Brief description of the drawings
[0020] To more clearly illustrate the technical solutions of the embodiments of the invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the invention. For those of ordinary skill in the art, without creative efforts, additional drawings can be obtained based on these drawings.
[0021] Figure 1 It is a schematic flow chart of a method for intelligent scientific research data management of geriatric syndromes according to the present invention; Figure 2 It is a schematic flow chart of obtaining a high-risk medical record prediction model in a method for intelligent scientific research data management of geriatric syndromes according to the present invention; Figure 3 It is a schematic flow chart of obtaining a significance analysis result in a method for intelligent scientific research data management of geriatric syndromes according to the present invention; Figure 4 It is a schematic module diagram of a system for intelligent scientific research data management of geriatric syndromes according to the present invention. Detailed implementation manners
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the invention with reference to the accompanying drawings in the embodiments of the invention. Obviously, the described embodiments are only some, rather than all, of the embodiments of the invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the invention without creative efforts fall within the scope of protection of the invention.
[0023] In the description of the present invention, it should be understood that the terms "openings", "upper", "lower", "top", "middle", "inner", etc. indicating orientations or positional relationships are only for convenience of describing the invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the invention. Embodiment
[0024] Please refer to Figure 1 、 Figure 2 、 Figure 3 The present invention discloses a method for intelligent scientific research data management of geriatric syndromes, including the following steps: S1. Integrate the structured and unstructured data of the hospital to obtain multi-source patient structured data and time-series feature data; The S1 includes the following steps: S11. Set the hospital information system set , where a i represents the i th information system of the hospital, bRepresents the total number of hospital information systems; the hospital information systems such as HIS, EMR, LIS, PACS; Connect to each information system in the hospital information system concentration through a standardized interface, extract the structured data of patients, and obtain patient structured data; the patient structured data includes basic information, diagnosis records, inspection reports, etc.; Perform format conversion on the patient structured data to obtain patient structured data with a unified format; unify it into a preset JSON or XML format to ensure that the data fields are mapped consistently with the target database; S12. Use natural language processing technology to parse the text data in the unstructured data, and extract key entities and temporal relationships to obtain structured entity data; the text data includes medical record texts, course records, discharge summaries, etc.; the key entities such as symptoms, disease names, drug doses; the temporal relationship such as "blood pressure decreased three days after taking the medicine"; Scan the pictures and paper data in the unstructured data through OCR technology to obtain the original text; call the AI semantic model to perform secondary cleaning and structured processing on the original text to obtain structured field data; the OCR technology such as Tesseract, Alibaba Cloud OCR; The patient structured data with a unified format, structured entity data, and structured field data together constitute multi-source patient structured data; S13. Integrate the device data in the unstructured data, store the dynamic monitoring indicators through a time series database, and associate them with the patient ID to obtain time series feature data; the sensor data such as muscle strength, walking speed, etc.; S2. Construct an ontology for the special field of geriatric syndromes; perform data cleaning and normalization on the multi-source patient structured data based on the ontology for the special field of geriatric syndromes to obtain patient standardized term data; The S2 includes the following steps: S21. Construct an ontology for the special field of geriatric syndromes; The S21 includes the following steps: S21. Define the core data elements of geriatric syndromes in the ontology for the special field of geriatric syndromes; the geriatric syndromes such as sarcopenia, cognitive impairment; the core data elements such as Chinese name, English name, value range, data level; S22. Define the relationships between entities in the ontology for the special field of geriatric syndromes through a knowledge graph to obtain logical constraint rules; the relationships between entities such as "the diagnosis of sarcopenia must include three indicators: muscle mass, grip strength, and walking speed"; The logical constraint rules include time logic rules, numerical logic rules, and semantic logic rules; S22. Combine with the ontology in the specialized field of geriatric syndromes to clean and normalize the multi-source structured patient data, and obtain the standardized term data of the patients; The S22 includes the following steps: S221. Use the time logic rules of the ontology in the specialized field of geriatric syndromes to detect the data in the multi-source structured patient data that does not conform to the time logic, and trigger automatic correction or manual review; for example, detect that "the admission time is later than the discharge time"; S222. Use the numerical logic rules of the ontology in the specialized field of geriatric syndromes to detect the data in the multi-source structured patient data that does not conform to the numerical logic, and trigger automatic correction or manual review; for example, a height of 205 cm exceeds the numerical rules; S223. Use the semantic logic rules of the ontology in the specialized field of geriatric syndromes and combine with the edit distance algorithm to calculate the semantic similarity, and convert the non-standard terms and units in the multi-source structured patient data into the standard semantics; for example, map "dementia" to "Alzheimer's disease" and convert "pound" to "kg"; the formula of the edit distance algorithm is as follows,
[0025] Among them, S i represents the semantic similarity, s 1 represents the non-standard term and unit data in the multi-source structured patient data, s 2 represents the standard semantics in the ontology in the specialized field of geriatric syndromes; ED ( s 1, s 2) represents s the edit distance between 1 and s 2 (that is, the minimum number of edit operations required to convert s 1 into s 2); represents s the maximum value of the data lengths of 1 and s 2; S224. Through steps S221, S223, and S223, obtain the standardized term data; S3. According to the standardized term data of the patients, the time series feature data, and in combination with the research variable requirements, construct a disease-specific data matrix; Based on the case inclusion rules and the disease-specific data matrix, obtain the target research medical record cohort data; The S3 includes the following steps: S31. According to the standardized term data of the patients, the time series feature data, in combination with the research requirement variables, and generate an electronic case report form by dragging variables; research requirement variables such as age, walking speed, and number of falls; S32. Upload the electronic case report form to the Excel template and automatically map it to the system fields to obtain the disease-specific data matrix C. The disease-specific data matrix is as follows:
[0026] where C in represents the data of the i th patient of the n th disease, m represents the total number of disease types; S33. Set the case inclusion rules, such as age ≥ 65 years old, etc.; Based on the case inclusion rules and the disease-specific data matrix, use Boolean logic to screen cases and target the research case cohort; S4. Use historical case data combined with an optimization algorithm to train and optimize the initial random forest model to obtain a high-risk case prediction model; Input the data of the target research case cohort into the high-risk case prediction model to obtain real-time high-risk cases; The S4 includes the following steps: S41. Construct an initial random forest model, set the number of trees of the initial random forest model; set the training accuracy of the initial random forest model to α , and the training accuracy threshold to β ; The single-tree depth of the initial random forest model is 15 layers, the minimum number of samples for node splitting is 3, the minimum number of samples for leaf nodes is 2, and each tree randomly selects √n or log2(n) features; S42. Collect historical case data, where the historical case data contains the physical data and high-risk labels of each patient; S43. Use the historical case data to train the initial random forest model. During the training process, use the optimization algorithm to find the number of trees of the initial random forest model, and judge the convergence of the initial random forest model according to the training accuracy and the training accuracy threshold to obtain the optimal solution; Use the optimal solution as the number of trees of the initial random forest model to obtain a high-risk case prediction model; The steps of using the optimization algorithm to find the number of trees of the initial random forest model and judging the convergence of the initial random forest model according to the training accuracy and the training accuracy threshold to obtain the optimal solution in S43 include the following steps: S431. Construct a wild horse population, set the size of the wild horse population to z , then the wild horse population is represented as , where p i represents the iA wild horse; set the maximum number of optimization iterations; S432. According to the number of trees in the initial random forest, randomly set the initial positions of the wild horse population to obtain an initial set of the wild horse population as , where q i represents the initial position of the i th wild horse in the wild horse population; the position of the wild horse can reflect the distance of the wild horse from the river; S433. According to the training accuracy α of the initial random forest, a training accuracy threshold β , define a fitness function for the initial positions of the wild horses in the wild horse population. The formula of the fitness function is as follows.
[0027] where ,g represents the fitness function, η represents the bias (the bias is used to help the model better fit the data); S434. Perform an iterative operation on the initial position set of the wild horse population. The higher the fitness value, the closer the position of the wild horse is to the river. In each round of iteration, according to the fitness function, calculate the fitness value of each position in the initial position set of the wild horse population, and update the position concentration of each wild horse in the initial position set of the wild horse population from high to low according to the fitness value. And in each round of iteration, obtain the best wild horse individual position and the global best wild horse position in the wild horse population; S435. Repeat S434. When the maximum number of optimization iterations is reached, stop the iteration, and use the global best wild horse position as the optimal solution; S44. Input the target research medical record cohort data into the high-risk medical record prediction model to obtain real-time high-risk cases; S5. Perform a significance analysis on the disease-specific data matrix to obtain a significance analysis result; perform an interpretable AI analysis on the real-time high-risk cases to obtain variable contribution degree data; Generate a scientific research analysis report according to the significance analysis result combined with the variable contribution degree data; The S5 includes the following steps: S51. Divide the data in the disease-specific data matrix into continuous data and categorical data according to the data type; S52. Use the t-test algorithm to perform a significance analysis on the continuous data to obtain a continuous data analysis result. The formula of the t-test algorithm is as follows.
[0028] where, t represents the continuous data analysis result, Xi , X j respectively represent the mean of the i th group of data in continuous data (e.g., the average walking speed of sarcopenia patients is 0.75 m / s) and the j th group of data in continuous data (e.g., the average walking speed of non-sarcopenia patients is 0.95 m / s). k 2 i , k 2 j respectively represent the variance of the i th group of data in continuous data (e.g., the variance of the walking speed of sarcopenia patients is 0.02 m² / s²) and the j th group of data in continuous data (e.g., the variance of the walking speed of non-sarcopenia patients is 0.01 m² / s²). h i , h j respectively represent the total amount of the i th group of data in continuous data (e.g., the number of sarcopenia patients is 100 cases) and the j th group of data in continuous data (e.g., the number of sarcopenia patients is 100 cases); S53. Use the chi-square test to perform a significance analysis on categorical data to obtain the categorical data analysis result; the chi-square test formula is as follows.
[0029] Among them, λ 2 represents the categorical data analysis result; O de represents the observed frequency, such as the actual observed number of "sarcopenia and fall" patients is 45 cases; E de represents the expected frequency, such as the expected observed number of "sarcopenia and fall" patients is 36 cases; S54. The continuous data analysis result and the categorical data analysis result together constitute the significance analysis result; S55. Calculate the variable contribution degree of real-time high-risk cases through SHAP values to obtain variable contribution degree data; for example, the contribution ratio of walking speed to the fall risk is 30%; S56. Generate a scientific research analysis report according to the significance analysis result combined with the variable contribution degree data. Embodiment
[0030] Please refer to Figure 4, a smart scientific research data management system for geriatric syndromes, which is used to implement the above-mentioned smart scientific research data management method for geriatric syndromes, including a multi-source medical data integration and feature extraction module, a domain ontology construction and data standardization module, a disease-specific data matrix construction and case screening module, a high-risk prediction model training and optimization module, and a scientific research analysis and report generation module; The multi-source medical data integration and feature extraction module docks with the hospital information system through a standardized interface, integrates the structured data of patients, and uses natural language processing, OCR technology, and time series databases to parse unstructured data, extracts key entities and time series features, and finally generates multi-source patient structured data and time series feature data in a unified format to provide standardized input for subsequent analysis; The domain ontology construction and data standardization module constructs a proprietary domain ontology based on the core data elements and knowledge graph of geriatric syndromes, defines the logical constraint rules between entities; through time logic detection, numerical logic correction, and semantic logic mapping, combined with the edit distance algorithm, it converts multi-source data into standardized term data to ensure data consistency and scientific research availability; The disease-specific data matrix construction and case screening module dynamically generates an electronic case report form according to the research requirement variables, constructs a disease-specific data matrix through Excel template mapping, and uses Boolean logic to screen the target research cohort in combination with the case inclusion rules; it converts the standardized data into a research-oriented structured data set to support subsequent statistical analysis and model training; The high-risk prediction model training and optimization module uses the random forest model as the basic model, dynamically adjusts the tree number parameter in combination with the wild horse optimization algorithm, trains the model through historical medical record data and evaluates the convergence, and finally generates a high-risk medical record prediction model; the high-risk medical record prediction model realizes the accurate prediction of real-time high-risk cases in the target cohort through feature selection and hyperparameter optimization; The scientific research analysis and report generation module conducts significance analysis and interpretability analysis on the disease-specific data matrix, identifies key variables, and generates a scientific research analysis report in combination with statistical results and variable contribution degrees; it reveals the internal relationship of the data through quantitative analysis to provide statistical evidence and clinical decision-making support for the research of geriatric syndromes.
[0031] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0032] The preferred embodiments of the invention disclosed above are only used to help illustrate the invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principle and practical application of the invention, so that those skilled in the art can well understand and utilize the invention.
Claims
1. An intelligent scientific research data management method for geriatric syndromes, characterized in that, It includes the following steps: S1. Integrate the structured and unstructured data of the hospital to obtain multi-source patient structured data and time-series feature data; S2. Construct an ontology for the specific field of geriatric syndromes; Based on the ontology for the specific field of geriatric syndromes, perform data cleaning and normalization on the multi-source patient structured data to obtain patient standardized term data; S3. According to the patient standardized term data, time-series feature data, and combined with the requirements of research variables, construct a disease-specific data matrix; Based on the case enrollment rules and the disease-specific data matrix, obtain the target research medical record cohort data; S4. Use historical medical record data combined with an optimization algorithm to train and optimize the initial random forest model to obtain a high-risk medical record prediction model; Input the target research medical record cohort data into the high-risk medical record prediction model to obtain real-time high-risk cases; S5. Perform significance analysis on the disease-specific data matrix to obtain a significance analysis result; Perform interpretable AI analysis on real-time high-risk cases to obtain variable contribution degree data; Generate a scientific research analysis report according to the significance analysis result combined with the variable contribution degree data.
2. The intelligent scientific research data management method for geriatric syndrome according to claim 1, characterized in that The S1 includes the following steps: S11. Set a hospital information system set, and the hospital information system set contains each information system of the hospital; Connect to each information system in the hospital information system set through a standardized interface, extract the structured data of patients to obtain patient structured data; Perform format conversion on the patient structured data to obtain patient structured data with a unified format; S12. Use natural language processing technology to parse the text data in the unstructured data, and extract key entities and time-series relationships to obtain structured entity data; Scan the pictures and paper data in the unstructured data through OCR technology to obtain the original text; call the AI semantic model to perform secondary cleaning and structured processing on the original text to obtain structured field data; The patient structured data with a unified format, structured entity data, and structured field data together constitute multi-source patient structured data; S13. Integrate the device data in the unstructured data, store the dynamic monitoring indicators through a time-series database, and associate them with the patient ID to obtain time-series feature data.
3. The intelligent scientific research data management method for geriatric syndrome according to claim 1, wherein, The S2 includes the following steps: S21. Construct an ontology for the specific field of geriatric syndromes; S22. Combine the ontology for the specific field of geriatric syndromes to perform data cleaning and normalization on the multi-source patient structured data to obtain patient standardized term data.
4. The intelligent scientific research data management method for geriatric syndrome according to claim 3, wherein, The S21 includes the following steps: S211. Define the core data elements of geriatric syndromes in the ontology for the specific field of geriatric syndromes; S212. Define the relationships between entities in the ontology for the specific field of geriatric syndromes through a knowledge graph to obtain logical constraint rules; The logical constraint rules include time logic rules, numerical logic rules, and semantic logic rules.
5. The intelligent scientific research data management method for geriatric syndrome according to claim 3, characterized in that, The S22 includes the following steps: S221. Use the time logic rules of the ontology for the specific field of geriatric syndromes to detect the data in the multi-source patient structured data that does not conform to the time logic, and trigger automatic correction or manual review; S222. Use the numerical logic rules of the ontology in the specific field of geriatric syndromes to detect the data that does not conform to the numerical logic in the multi-source patient structured data, and trigger automatic correction or manual review; S223. Use the semantic logic rules of the ontology in the specific field of geriatric syndromes and combine the edit distance algorithm to calculate the semantic similarity, and convert the non-standard terms and units in the multi-source patient structured data into standard semantics; S224. Through steps S221, S223, and S223, obtain the standardized term data.
6. The intelligent scientific research data management method for geriatric syndrome according to claim 1, characterized in that The said S3 includes the following steps: S31. According to the patient standardized term data and the time series feature data, combine the scientific research requirement variables, and generate an electronic case report form by dragging variables; S32. Upload the electronic case report form to the Excel template and automatically map it to the system fields to obtain the disease-specific data matrix; S33. Set the case inclusion rules; based on the case inclusion rules and the disease-specific data matrix, use Boolean logic to screen cases to obtain the target research medical record cohort.
7. A smart scientific research data management method for geriatric syndromes according to claim 1, characterized in that The said S4 includes the following steps: S41. Construct an initial random forest model, set the number of trees in the initial random forest model; set the training accuracy of the initial random forest model and the training accuracy threshold; S42. Collect historical medical record data, and the historical case data contains the physical data and high-risk labels of each patient; S43. Use the historical medical record data to train the initial random forest model. During the training process, find the number of trees in the initial random forest model through an optimization algorithm, and judge the convergence of the initial random forest model according to the training accuracy and the training accuracy threshold to obtain the optimal solution; Take the said optimal solution as the number of trees in the initial random forest model to obtain the high-risk medical record prediction model.
8. The intelligent scientific research data management method for geriatric syndrome according to claim 7, wherein, The steps of finding the number of trees in the initial random forest model through the optimization algorithm and judging the convergence of the initial random forest model according to the training accuracy and the training accuracy threshold to obtain the optimal solution in the said S43 include the following steps: S431. Construct a wild horse population, set the scale of the wild horse population; set the maximum number of optimization iterations; S432. According to the number of trees in the initial random forest, randomly set the initial positions of the wild horse population to obtain the initial set of the wild horse population; S433. According to the training accuracy and the training accuracy threshold of the initial random forest, define the fitness function of the initial positions of the wild horses in the wild horse population; S434. Perform iterative operations on the initial position set of the wild horse population. During each round of iteration, calculate the fitness value of each position in the initial position set of the wild horse population according to the fitness function, update the position concentration of each wild horse in the initial position set of the wild horse population from high to low according to the fitness value, and obtain the best wild horse individual position and the global best wild horse position in the wild horse population during each round of iteration; S435. Repeat S434. When the maximum number of optimization iterations is reached, stop the iteration, and take the global best wild horse position as the optimal solution.
9. The intelligent scientific research data management method for geriatric syndrome according to claim 1, characterized in that The said S5 includes the following steps: S51. Divide the data in the disease-specific data matrix into continuous data and categorical data according to the data type; S52. Use the t-test algorithm to perform significance analysis on continuous data to obtain the continuous data analysis results; S53. Use the chi-square test to perform significance analysis on categorical data to obtain the categorical data analysis results; S54. The continuous data analysis results and the categorical data analysis results together constitute the significance analysis results; S55. Calculate the variable contribution degree of real-time high-risk cases through SHAP values to obtain variable contribution degree data; S56. Generate a scientific research analysis report according to the significance analysis results combined with the variable contribution degree data.
10. An intelligent scientific research data management system for geriatric syndromes, characterized in that, Implement an intelligent scientific research data management method for geriatric syndromes as described in any one of claims 1-9. The system includes a multi-source medical data integration and feature extraction module, a domain ontology construction and data standardization module, a disease-specific data matrix construction and case screening module, a high-risk prediction model training and optimization module, and a scientific research analysis and report generation module.
Citation Information
Patent Citations
Severe infectious disease queue data typing method, typing model and electronic equipment
CN112820416A
Lymphoma research database construction and application method based on real world research
CN115455973A
Special disease standard database automatic construction method based on patient data
CN117542467A
Asphalt pavement rut depth prediction method based on interpretable ensemble learning
CN117874509A
Machine learning-based early warning system for occurrence of acute kidney injury of critically ill patient
CN118299054A
Cited By
Medical teaching and research integrated system based on common bottom layer
CN120766912A
Medical education and research integrated system based on common bottom layer
CN120766912B
Visual data extraction method based on medical information system
CN121117093A