Data modeling system based on research data

By studying and optimizing data acquisition, cleaning, encryption, and model building modules, and combining AES encryption and algorithm matching analysis, the problems of single dependency form and insufficient security in medical research data modeling were solved, realizing flexible and diverse modeling and efficient and secure data processing.

CN121456928APending Publication Date: 2026-02-03HIPOWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511551721.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing medical research data modeling and processing suffers from problems such as a single form of dependency and insufficient attention to data security, which inhibits the intelligence and coordination of model optimization and processing, and makes it impossible to reliably guarantee data security.

Method used

The study employs a data acquisition module, a cleaning and transformation module, an encryption module, and a model building, screening, and optimization module. The structured dataset is encrypted using the AES symmetric encryption algorithm, and the modeling matching degree of each algorithm mechanism is analyzed. The model is then optimized and constructed, and the modeling algorithm library is used to store the machine algorithms and their historical training data.

Benefits of technology

It improves the flexibility and diversity of research data modeling, enhances data security, ensures the rationality, standardization and accuracy of the model, and improves the robustness and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456928A_ABST
    Figure CN121456928A_ABST
Patent Text Reader

Abstract

The invention discloses a data modeling system based on research data, which relates to the technical field of data modeling and comprises a research data acquisition module, a research data cleaning and conversion module, a research data encryption module, a model building, screening and optimizing module and a modeling algorithm library. According to the method, dynamic model construction algorithm selection can be carried out according to research data, model construction processing is carried out, the flexibility and diversity of research data modeling attachment forms are improved, the research data are finely analyzed, reasonable screening of the model construction algorithms can be carried out according to specific data characteristics, and the model construction efficiency is improved. According to the method, the intelligence and coordination of model optimization processing are powerfully supported and guaranteed, higher convenience is also provided for calling and checking of subsequent models, encryption processing is performed on the structured data set by using the AES symmetric encryption algorithm, the attention on data transmission and storage security is improved, and the method is suitable for being used in the field of data transmission and storage. And reliable guarantee can be provided for safe storage of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data modeling technology, specifically to a data modeling system based on research data. Background Technology

[0002] With the rapid development of data informatization and intelligence, hospital information systems cover a large amount of medical prescription data and medical imaging data. The rapid growth of medical data and the continuous expansion of data application needs have made data modeling and processing increasingly important. By modeling medical research data, various types of medical research data can be integrated, processed and stored, providing greater convenience for personnel to access and view the data and for collaborative storage management.

[0003] Currently, there are still many limitations in the modeling and processing of medical research data. Specifically, existing research data processing and modeling mainly adopts a static upload and modeling storage method. On the one hand, the data modeling method is relatively simple and lacks detailed analysis of the research data. This makes it impossible to reasonably select model building algorithms based on specific data characteristics, which inhibits the intelligence and coordination of model optimization to a certain extent. On the other hand, in terms of data security, there is relatively insufficient attention to data storage security, which cannot provide reliable guarantees for the secure storage of data.

[0004] In summary, there are still many aspects of medical research data modeling that need improvement and optimization. Based on this, a comprehensive and complete data modeling system is needed to ensure the rationality, standardization, and accuracy of research data modeling. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a data modeling system based on research data, which can effectively solve the problems mentioned in the background section.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a data modeling system based on research data, comprising: a research data acquisition module for acquiring research data.

[0007] The research data cleaning and transformation module is used to clean and transform the collected research data to obtain structured datasets.

[0008] The research focuses on a data encryption module used to encrypt structured datasets using the AES symmetric encryption algorithm.

[0009] The model building, screening, and optimization module is used to analyze the modeling matching degree of each algorithm mechanism with the corresponding structured dataset, optimize and build models for the structured dataset, and transmit them to the data modeling service processor for storage.

[0010] The modeling algorithm library stores various types of machine algorithms and their historical training data, reference validation ROC curves and reference validation PR curves, and stores the reference clustering metric ranges for the corresponding adapted processing datasets of each type of machine algorithm.

[0011] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects: (1) The present invention provides a data modeling system based on research data, which can select a dynamic model building algorithm based on the research data and perform model building processing. This not only effectively makes up for the defects caused by the use of solid-state upload modeling storage in the prior art, but also improves the flexibility and diversity of the research data modeling dependency form. The present invention has conducted a detailed analysis of the research data and can reasonably screen the model building algorithm according to the specific data characteristics, so that the intelligence and coordination of model optimization processing are strongly supported and guaranteed.

[0012] (2) By setting up a research data encryption module, the present invention uses the AES symmetric encryption algorithm to encrypt the structured dataset, thereby increasing the attention to data transmission and storage security and providing reliable protection for the secure storage of data.

[0013] (3) This invention analyzes the modeling matching degree of each algorithm mechanism with the corresponding structured dataset and optimizes the model construction of the structured dataset. Considering the potential differences in the performance of different algorithms on the same dataset, this invention also conducts in-depth analysis of different algorithm mechanisms based on the analysis of research data. Therefore, it effectively improves the adaptability and fit between the final executed algorithm mechanism and the structured dataset, and also effectively improves the robustness and robustness of the final constructed model, further ensuring the rationality, standardization and accuracy of the modeling of research data.

[0014] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the system module connections of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] In the description of this invention, it should be understood that the terms "opening", "upper", "lower", "thickness", "top", "middle", "length", "inner", "around", etc., which indicate orientation or positional relationship, are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the components or elements referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as limiting this invention.

[0018] Please see Figure 1 As shown, this embodiment of the invention provides a technical solution: a data modeling system based on research data, comprising: a research data acquisition module, a research data cleaning and transformation module, a research data encryption module, a model building, screening and optimization module, and a modeling algorithm library.

[0019] In its embodiments, this invention provides a data modeling system based on research data, which can dynamically select model building algorithms and perform model building processing based on the research data. This not only effectively compensates for the defects caused by the solid-state upload and storage method of existing technologies, but also improves the flexibility and diversity of research data modeling dependency forms. This invention conducts a detailed analysis of the research data and can reasonably select model building algorithms based on specific data characteristics, so that the intelligence and coordination of model optimization processing are strongly supported and guaranteed.

[0020] The research data acquisition module is used to collect research data.

[0021] Specifically, the research data includes electronic medical record information, prescription information, imaging information, and laboratory test information. Electronic medical records refer to digital information such as text, symbols, charts, graphics, numbers, and images generated by medical personnel using information systems during medical activities, and can be stored, managed, transmitted, and reproduced. It is a form of medical record, including outpatient medical records, emergency cases, and inpatient medical records.

[0022] The prescription information specifies the drug name, quantity, dosage form, and usage, ensuring the drug's specifications and safety and effectiveness.

[0023] The image information refers to image data obtained through various medical imaging technologies, typically including X-ray images, magnetic resonance imaging, and PET scan images.

[0024] The laboratory test information refers to data obtained by testing and analyzing samples from patients using laboratory techniques, including physiological, biochemical, and immunological aspects. This includes, for example, blood test information such as blood cell count, hemoglobin count, white blood cell count, and platelet count, as well as biochemical test information such as blood glucose, liver function indicators, and kidney function indicators.

[0025] The research data cleaning and transformation module is used to clean and transform the collected research data to obtain a structured dataset. The research data cleaning and transformation module is designed to address the situation where the formats of data from different sources are not uniform. By setting up the research data cleaning and transformation module, the system uses technologies such as natural language processing and information extraction to transform all research data into a structured dataset as data input for modeling.

[0026] The research data encryption module is used to encrypt structured datasets using the AES symmetric encryption algorithm. AES is a symmetric key algorithm, meaning that the same key is used to encrypt and decrypt data. The encrypted research data is stored to ensure that only authorized users can access the key and decrypt the research data.

[0027] In the embodiments of this invention, by setting up a research data encryption module and using the AES symmetric encryption algorithm to encrypt the structured dataset, the focus on data transmission and storage security is improved, providing reliable protection for secure data storage.

[0028] The model building, screening, and optimization module is used to analyze the modeling matching degree of each algorithm mechanism with the corresponding structured dataset, optimize and build models for the structured dataset, and transmit them to the data modeling service processor for storage. The data modeling service processor is usually a cloud computing platform used to manage and execute data modeling tasks. The functions of the data modeling service processor include data processing, model training, model evaluation, and model deployment. It can adapt to data modeling tasks of different scales and complexities and has good scalability and flexibility.

[0029] In the embodiments of this invention, the modeling matching degree of each algorithm mechanism with the corresponding structured dataset is analyzed, and the model of the structured dataset is optimized and constructed. Considering that different algorithms may have potential differences in performance on the same dataset, in-depth analysis of different algorithm mechanisms is also carried out based on the analysis of research data. Therefore, the adaptability and fit between the final executed algorithm mechanism and the structured dataset are effectively improved. At the same time, the robustness and robustness of the final constructed model are also effectively improved, further ensuring the rationality, standardization and accuracy of the modeling of research data.

[0030] Specifically, the modeling matching degree of each algorithm mechanism corresponding to the structured dataset is analyzed as follows: Based on the structured dataset, the clustering metric value of the structured dataset is obtained through data fitting normalization. In data modeling, data fitting normalization is a common data preprocessing technique used to map data processing of different features to the final values ​​to prevent the value range of some data from being too large, which would lead to unstable model training or bias towards certain features. The clustering metric value of the structured dataset is matched with the reference clustering metric value range of each type of machine algorithm in the modeling algorithm library corresponding to the adapted processing dataset. If the clustering metric value of the structured dataset is within the reference clustering metric value range of a certain type of machine algorithm corresponding to the adapted processing dataset, then that type of machine algorithm is defined as a candidate machine algorithm, thereby obtaining each candidate machine algorithm of the structured dataset.

[0031] The candidate machine learning algorithms in the structured dataset are extracted and combined in various ways to obtain the algorithm mechanisms. The branches of each algorithm mechanism are then statistically analyzed, and each algorithm mechanism is numbered as follows: , , The total number of algorithm mechanisms is given, and each branch algorithm is numbered as follows: , , This represents the total number of branches in the algorithm.

[0032] For example, if the candidate machine learning algorithms in the structured dataset include linear regression, logistic regression, and CNN classification, then by performing various extraction and combination processes, the number of algorithm mechanisms obtained is 7. From the first algorithm mechanism to the seventh algorithm mechanism, the branch algorithms covered are as follows: First algorithm mechanism: linear regression algorithm; second algorithm mechanism: logistic regression algorithm; third algorithm mechanism: CNN classification algorithm; fourth algorithm mechanism: linear regression algorithm and logistic regression algorithm; fifth algorithm mechanism: linear regression algorithm and CNN classification algorithm; sixth algorithm mechanism: logistic regression algorithm and CNN classification algorithm; seventh algorithm mechanism: linear regression algorithm, logistic regression algorithm, and CNN classification algorithm. According to the above combination rules, the branch algorithms in each algorithm mechanism can be counted.

[0033] Based on each branch of the algorithm mechanism, the first modeling fit value and the second modeling fit value of the corresponding structured dataset are obtained sequentially. These values ​​are then incorporated into the modeling matching weight coefficients to measure their impact on the modeling fit degree. Finally, the modeling fit degree of each algorithm mechanism corresponding to the structured dataset is obtained, with the following constraints: ; In the formula, The modeling matching degree of the algorithm mechanism d for the structured dataset. This represents the first modeling fit value for the structured dataset corresponding to each algorithm mechanism. The second modeling fit value for each algorithm mechanism corresponding to the structured dataset. , These are the modeling matching weight coefficients corresponding to the first and second modeling fit values ​​of the set algorithm mechanism, respectively.

[0034] In this embodiment, the modeling matching degree of each algorithm mechanism with the corresponding structured dataset is analyzed. This can evaluate the modeling matching degree of different algorithms, provide support for modeling decisions on the structured dataset, and adjust the selection of algorithm mechanisms or optimize the model building strategy for research data based on the analysis and evaluation results, so as to improve the robustness and adaptability of the final structured dataset model.

[0035] Furthermore, based on the structured dataset, the convergence parameter characterization coefficients and sub-data characterization coefficients of the structured dataset are obtained sequentially. Weighting coefficients are then introduced to quantify the influence of the convergence parameter characterization coefficients and sub-data characterization coefficients on the clustering metrics. Finally, the clustering metrics of the structured dataset are summarized. The specific processing constraints for the clustering metrics of the structured dataset are as follows: ; In the formula, For clustering metrics of structured datasets, , These are the convergence parameter characterization coefficients of the structured dataset and the characterization coefficients of the subsystem data, respectively. , These are the weight coefficients for the structured dataset and the subsystem data, respectively.

[0036] In this embodiment, analyzing the clustering metrics of the structured dataset enables the econometric evaluation of the structured dataset, which helps to provide precise data support for subsequently determining which algorithm to implement, thereby creating a more reliable model.

[0037] Furthermore, the specific processing steps for the convergence parameter characterization coefficients and subsystem data characterization coefficients of the structured dataset include: statistically analyzing the convergence parameters of the structured dataset and the attribute information of each subsystem data, wherein the convergence parameters include transmission timestamp, original spatial payload, data payload, and transmission compression ratio.

[0038] It should be understood that, in this embodiment, the transmission timestamp is the interval timestamp between the time the structured dataset leaves the sending end and the time it arrives at the receiving end. The original spatial payload refers to the complete content carried by the structured dataset during transmission, including the data header and data payload. The data header contains data information such as the source address, destination address, protocol information, and data checksum of the structured dataset. The data payload is the actual valid data portion carried in the structured dataset. The effective data payload refers to the actual valid data transmitted from the structured dataset, excluding the data header or other additional information; it is the true effective payload portion carried in the structured dataset. The transmission compression ratio is the ratio of the compressed size of the structured dataset to the uncompressed size.

[0039] The attribute information of each subsystem data includes the average resolution of the image, the average size of the image, and the average saturation of the image. Each subsystem data is a subset of the structured dataset, such as multiple subfiles in a compressed file.

[0040] It should be understood that in this embodiment, the structured dataset is identified and parsed to distinguish different types of files such as text and graphics, and PC calculations are performed to obtain the final average resolution, average size, and average saturation of the graphics.

[0041] Based on the aggregation parameters of the structured dataset, the transmission timestamp is extracted and scaled up to a set reference transmission timestamp for verification. The ratio between the original spatial payload and the data payload is scaled up to a set spatial definition ratio of the data payload for verification. Simultaneously, the transmission compression ratio is scaled up to a set reference transmission compression ratio for verification. The results of each scaling-up verification are statistically analyzed, and corresponding weights are introduced to aggregate and couple the results one by one to obtain the aggregation parameter characterization coefficients of the structured dataset. The constraints are as follows: ; In the formula, The aggregation parameters of the structured dataset represent the coefficients. , , , These are the transmission timestamp, original spatial payload, data payload, and transmission compression ratio, which are the aggregation parameters of the structured dataset. For the set reference transmission timestamp, To define the spatial definition ratio of the data payload, The set reference transmission compression ratio, , and The weights are set for the transmission timestamp, the percentage of data payload, and the transmission compression ratio.

[0042] In this embodiment, the aggregation parameter characterization coefficients of the structured dataset are analyzed. The transmission timestamp, original spatial payload, data payload and transmission compression ratio are numerically aggregated to analyze information such as the latency and data volume of the structured dataset. This allows for the quantitative evaluation of the characteristics of the structured dataset and provides data support for the clustering measurement of the structured dataset.

[0043] Based on the average resolution, average size, and average saturation of the graphics in the attribute information of each subsystem data, and comparing them with the set reference resolution, reference size, and reference saturation of the subsystem data, a proportional-driven verification process is performed to obtain the proportional-driven verification results of each subsystem data. Corresponding weights are then introduced to aggregate the proportional-driven verification results of each subsystem data, and finally, the subsystem data representation coefficients of the structured dataset are obtained. The constraints are as follows: ; In the above formula, These are the subsystem data representation coefficients of the structured dataset. , , The attributes of the subsystem data f, in order, are the average resolution, average size, and average saturation of the graphics. , , These are the weights corresponding to the set average image resolution, average image size, and average image saturation, respectively. , , These represent the graphic reference resolution, graphic reference size, and graphic reference saturation of the defined subsystem data, respectively, where f is the number of each subsystem data. , This represents the total number of subsystem data.

[0044] In this embodiment, the sub-data representation coefficients of the structured dataset are analyzed. The purpose is that when the number of characters, paragraphs, graphics, average resolution of graphics, average size of graphics, and average saturation of graphics in the sub-data are all in a high value state, it indicates that the structured dataset has a large data volume, strong professionalism, and detailed and complex content. At this time, the requirements for the selection of algorithms are also higher. Numerical processing can further optimize the selection decision of subsequent model building algorithms.

[0045] More specifically, the first modeling fit value of each algorithm mechanism corresponding to the structured dataset is analyzed in the following process: the historical training data of each branch algorithm of each algorithm mechanism is statistically analyzed from the modeling algorithm library, the execution training time of each historical call is extracted from it, and the historical verification accuracy, historical verification precision, historical verification recall and historical verification F1 score are extracted at the same time.

[0046] The execution training time of historical calls serves as a reference indicator for algorithm selection. If the execution training time of an algorithm is too long, or if there are significant differences in the execution training time of each call, it indicates a potential risk of performance degradation. Further optimization or parameter adjustment may be needed to improve the execution efficiency and performance of the algorithm.

[0047] Historical validation accuracy refers to the proportion of correctly predicted samples out of the total number of samples, which can measure the overall accuracy of the algorithm in building the model.

[0048] Historical verification accuracy refers to the proportion of samples predicted as positive that are actually positive. High accuracy indicates that the algorithm is more reliable.

[0049] Historical validation recall refers to the proportion of predicted positive examples to actual positive examples. It measures the algorithm's ability to identify actual positive examples. A high recall indicates that the algorithm is better able to identify actual positive examples.

[0050] Historical validation of the F1 score, which is the harmonic mean of precision and recall, provides a more comprehensive evaluation of the algorithm's performance.

[0051] Extract the historical validation accuracy, historical validation precision, historical validation recall, and historical validation F1 score of each branch of each algorithm mechanism and label them successively. , , , .

[0052] The historical training data of each branch of each algorithm mechanism are averaged to obtain the historical training data mean set of each algorithm mechanism. The historical training data mean set specifically includes the historical verification accuracy mean, historical verification precision mean, historical verification recall mean, and historical verification F1 score mean of each algorithm mechanism. The proportionalized validation results of each branch of each algorithm mechanism are statistically analyzed and compared with the mean set of historical training data of each algorithm mechanism. The influence of the proportionalized validation results on the first model fit value is measured using the model fit weight coefficient. Finally, the proportionalized validation results of each algorithm mechanism are summarized and averaged to obtain the first model fit value of the corresponding structured dataset for each algorithm mechanism. The constraints are as follows: .

[0053] In the above formula, This represents the first modeling fit value for the structured dataset corresponding to each algorithm mechanism. , , , These are the modeling fit weight coefficients for the historical validation accuracy, historical validation precision, historical validation recall, and historical validation F1 score, respectively.

[0054] In the embodiments, the first modeling fit value of the structured dataset corresponding to each algorithm mechanism is analyzed, and the historical verification accuracy, historical verification precision, historical verification recall, and historical verification F1 score of each branch algorithm mechanism are comprehensively evaluated. This can fully evaluate the performance of each algorithm mechanism, increase the credibility of the final model construction, and enhance the sustainability of the model's application.

[0055] More specifically, the second modeling fit value of each algorithm mechanism corresponding to the structured dataset is analyzed in the following ways: based on the historical training data of each branch of the algorithm mechanism, historical validation ROC curves and historical validation PR curves are extracted. The horizontal axis of the ROC curve represents the false positive rate, which is the proportion of negative samples that are incorrectly classified as positive out of all negative samples. The vertical axis represents the true positive rate, which is the proportion of positive samples that are correctly classified as positive out of all positive samples. The ROC curve, by plotting the changes in the true positive rate and false positive rate under different classification thresholds, demonstrates the trade-off between sensitivity and false alarm rate of the algorithm. The horizontal axis of the PR curve represents the recall rate, and the vertical axis represents the precision rate. The PR curve, by plotting the changes in precision and recall under different classification thresholds, demonstrates the trade-off between precision and recall of the algorithm.

[0056] Based on the historical validation ROC curves and historical validation PR curves of each branch algorithm of each algorithm mechanism, and extracting the reference validation ROC curves and reference validation PR curves of each branch algorithm of each algorithm mechanism from the modeling algorithm library, and sequentially performing overlap verification, the overlap lengths of the historical validation ROC curves and historical validation PR curves of each branch algorithm of each algorithm mechanism are extracted and marked respectively. , .

[0057] The overlap lengths of the historical validation ROC curves and historical validation PR curves of each branch of each algorithm mechanism are compared with the set overlap lengths of the historical validation ROC curves and historical validation PR curves to verify their proportional relationship. A correction factor is then used to adjust the influence of the proportional relationship verification results on the second modeling fit value. Finally, the proportional relationship verification results of each branch of each algorithm mechanism after correction are summed and averaged to obtain the second modeling fit value of the corresponding structured dataset for each algorithm mechanism. The constraints are as follows: .

[0058] In the above formula, The second modeling fit value for each algorithm mechanism corresponding to the structured dataset. , These are the reference values ​​for the overlap length of the historical validation ROC curve and the historical validation PR curve, respectively. , These are the correction factors for the set ROC curve and PR curve, respectively.

[0059] In this embodiment, the second modeling fit value of the structured dataset corresponding to each algorithm mechanism is analyzed. Based on the historical validation ROC curve and historical validation PR curve of each branch algorithm of each algorithm mechanism, and through in-depth analysis, the specific performance of the algorithm in terms of classification accuracy, confusion and robustness can be evaluated more comprehensively and intuitively, providing accurate data assurance for the selection of subsequent algorithms.

[0060] More specifically, the process of optimizing and constructing a model for a structured dataset includes: selecting the algorithm mechanism with the highest model matching degree for the structured dataset based on the model matching degree of each algorithm mechanism, using it as the first execution algorithm mechanism for the structured dataset, and having the first execution algorithm mechanism perform the model construction process for the structured dataset.

[0061] More specifically, the model building process of the structured dataset by the first execution algorithm mechanism further includes: before the model building process, sampling pre-modeling processing of the structured dataset is performed. The specific process is: based on the sub-system data representation coefficients of the structured dataset, and input into the preset mapping set between the sub-system data representation coefficients and the modeling sampling ratio in the database, the modeling sampling ratio of the structured dataset is obtained, which is denoted as the specified modeling sampling ratio. Data is drawn from the structured dataset at a specified modeling sampling ratio, and the drawn data is recorded as the modeling sample data; The modeling sample data is dynamically cropped to obtain the target modeling sample data; The target modeling sample data is pre-modeled using the first execution algorithm mechanism. The training convergence rate during the pre-modeling process is obtained and compared with the pre-set verification training convergence rate. If the training convergence rate during the pre-modeling process is greater than or equal to the verification training convergence rate, it is determined that the first execution algorithm mechanism will directly perform the model construction process of the structured dataset. Otherwise, the first execution algorithm mechanism will be dynamically optimized, and the model construction process of the structured dataset will be performed through the dynamically optimized first execution algorithm mechanism.

[0062] More specifically, the process of dynamically optimizing the first execution algorithm mechanism is as follows: the deviation between the training convergence rate and the verification training convergence rate during the pre-modeling process is statistically analyzed and recorded as the training convergence rate deviation value during the pre-modeling process. The training convergence rate deviation value in the pre-modeling process is matched with the learning rate adjustment coefficient corresponding to each preset training convergence rate deviation value interval to obtain the learning rate adjustment coefficient in the pre-modeling process. The algorithm covers several branches of the algorithm under the first execution algorithm mechanism, which are denoted as the target algorithm. The current set learning rate of the target algorithm is calculated and coupled (i.e. multiplied) with the learning rate adjustment coefficient in the pre-modeling process to obtain the expected set learning rate of the target algorithm. The expected learning rate of the target algorithm is compared with the preset upper limit of the target algorithm's learning rate. If the expected learning rate of the target algorithm is less than the upper limit of the target algorithm's learning rate, the subsequent model construction will be carried out using the expected learning rate of the target algorithm; otherwise, the subsequent model construction will be carried out using the upper limit of the target algorithm's learning rate.

[0063] This embodiment performs model building and processing of structured datasets, including but not limited to classification models and image analysis models. It can provide data model visualization, statistics and query functions, support the digital governance of medical information, and fully ensure the effective utilization of the model. For example, in the classification model, electronic medical record information with the same disease is stored in the same type of model, which can provide retrieval and verification functions for information such as disease diagnosis, treatment plan, and medical history. In the image analysis model, it can provide in-depth analysis functions for medical images such as X-ray images, MRI images, and CT scan images.

[0064] The modeling algorithm library is used to store various types of machine algorithms and their historical training data, reference validation ROC curves and reference validation PR curves, and to store the reference clustering metric ranges of the corresponding adaptive processing datasets for each type of machine algorithm.

[0065] In the embodiments, the above-mentioned machine algorithms cover classification and regression algorithms such as linear regression, logistic regression, SVM, RNN, and CNN, and can be further expanded. Model building and training can be performed using single or combined algorithms according to data characteristics or specific requirements.

[0066] In another specific embodiment, the system can automatically iteratively try all algorithms in the modeling algorithm library and models built under different parameter settings, use a multi-fold cross-validation mechanism to evaluate the performance of each built model, and simultaneously and automatically select the model with the best performance.

[0067] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0068] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. The selection and detailed description of these embodiments in this specification are intended to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. Any modifications or variations that do not deviate from the structure of the invention or exceed the scope defined by the invention should fall within the protection scope of the invention.

Claims

1. A data modeling system based on research data, characterized by, include: The research data acquisition module is used to collect research data. The research data cleaning and transformation module is used to clean and transform the collected research data to obtain structured datasets. The research focuses on a data encryption module used to encrypt structured datasets using the AES symmetric encryption algorithm. The model building, screening, and optimization module is used to analyze the modeling matching degree of each algorithm mechanism with the corresponding structured dataset, optimize and build models for the structured dataset, and transmit them to the data modeling service processor for storage. The modeling algorithm library stores various types of machine algorithms and their historical training data, reference validation ROC curves and reference validation PR curves, and stores the reference clustering metric ranges for the corresponding adapted processing datasets of each type of machine algorithm.

2. The data modeling system based on research data according to claim 1, wherein: The research data specifically includes electronic medical records, prescription information, imaging information, and laboratory test information.

3. The data modeling system based on research data according to claim 1, wherein: The modeling matching degree of each algorithm mechanism for the structured dataset is analyzed in detail, including: Based on the structured dataset, the clustering metrics of the structured dataset are obtained through data fitting. These metrics are then matched with the reference clustering metrics ranges of various types of machine algorithms in the modeling algorithm library to obtain candidate machine algorithms for the structured dataset. The candidate machine algorithms in the structured dataset are extracted and combined in various ways to obtain the algorithm mechanisms, and the branches of each algorithm mechanism are statistically analyzed. Based on each branch algorithm in each algorithm mechanism, the first modeling fit value and the second modeling fit value of the structured dataset corresponding to each algorithm mechanism are processed sequentially. The first modeling fit value and the second modeling fit value are respectively introduced into the modeling matching weight coefficient to measure the influence on the modeling matching degree. Finally, the modeling matching degree of the structured dataset corresponding to each algorithm mechanism is obtained by summing them up.

4. The data modeling system based on research data according to claim 3, wherein: The specific processing procedure for the clustering metrics of the structured dataset is as follows: Based on the structured dataset, the clustering parameter characterization coefficients and sub-data characterization coefficients of the structured dataset are processed sequentially. Weight coefficients are then introduced to quantify the influence of the clustering parameter characterization coefficients and sub-data characterization coefficients on the clustering metrics. Finally, the clustering metrics of the structured dataset are obtained by summarizing the results.

5. The data modeling system based on research data according to claim 4, characterized in that: The specific processing steps for the convergence parameter characterization coefficients and subsystem data characterization coefficients of the structured dataset include: The aggregation parameters of the statistical structured dataset and the attribute information of each subsystem; The convergence parameters include transmission timestamp, original spatial payload, data payload and transmission compression ratio, and the attribute information of each subsystem data includes average image resolution, average image size and average image saturation. The transmission time stamp is extracted from the structured data set based on the aggregation parameter, and the transmission time stamp is scaled and driven for checking processing with the set reference transmission time stamp. The ratio between the original space load and the data payload is scaled and driven for checking processing with the set space defined proportion of the data payload. The transmission compression ratio is scaled and driven for checking processing with the set reference transmission compression ratio. The results of the scaling and driving checking processing are counted, and the corresponding weights are introduced to aggregate and couple the results of the scaling and driving checking processing one by one to obtain the aggregation parameter representation coefficient of the structured data set. Based on the average graphic resolution, the average graphic size and the average graphic saturation in the attribute information of each sub-system data, the scaling and driving checking processing is performed with the set graphic reference resolution, the graphic reference size and the graphic reference saturation of the sub-system data respectively. The scaling and driving checking processing results of each sub-system data are obtained, and the corresponding weights are introduced to aggregate the scaling and driving checking processing results of each sub-system data. Finally, the sub-system data representation coefficient of the structured data set is obtained.

6. The data modeling system based on research data according to claim 3, wherein: The first modeling fitting value of the structured data set corresponding to each algorithm mechanism includes the following specific analysis process: The historical training data of each branch algorithm of each algorithm mechanism is counted from the modeling algorithm library, and the historical verification accuracy, the historical verification precision, the historical verification recall rate and the historical verification F1 score are extracted therefrom. The historical training data of each branch algorithm of each algorithm mechanism is processed by mean value to obtain the historical training data mean set of each algorithm mechanism. The historical training data mean set specifically includes the historical verification accuracy mean value, the historical verification precision mean value, the historical verification recall rate mean value and the historical verification F1 score mean value of each algorithm mechanism. The scaling verification result between the historical training data of each branch algorithm of each algorithm mechanism and the historical training data mean set of each algorithm mechanism is counted, and the influence component of the scaling verification result on the first modeling fitting value is measured by the modeling fitting weight coefficient. Finally, the scaling verification results of each algorithm mechanism after measurement are averaged to obtain the first modeling fitting value of the structured data set corresponding to each algorithm mechanism.

7. The data modeling system based on research data according to claim 3, wherein: The second modeling fitting value of the structured data set corresponding to each algorithm mechanism includes the following specific analysis process: According to the historical training data of each branch algorithm of each algorithm mechanism, the historical verification ROC curve and the historical verification PR curve are extracted therefrom, and the reference verification ROC curve and the reference verification PR curve of each branch algorithm of each algorithm mechanism are extracted from the modeling algorithm library. The historical verification ROC curve overlap length and the historical verification PR curve overlap length of each branch algorithm of each algorithm mechanism are extracted by coincidence checking in sequence. The historical verification ROC curve coincidence length and the historical verification PR curve coincidence length of each branch algorithm of each algorithm mechanism are verified in proportional relationship with the coincidence length reference limit value of the historical verification ROC curve and the historical verification PR curve, and the influence component of the proportional relationship verification result on the second modeling fitting value is modified by a correction factor. Finally, the proportional relationship verification results of each branch algorithm of each algorithm mechanism after the modification processing are averaged to obtain the second modeling fitting value of each algorithm mechanism corresponding to the structured data set.

8. The data modeling system based on research data of claim 1, wherein: The model optimization construction of the structured data set includes the following specific processes: According to the modeling matching degree of each algorithm mechanism corresponding to the structured data set, the algorithm mechanism with the highest modeling matching degree is screened as the first execution algorithm mechanism of the structured data set, and the model construction processing of the structured data set is performed by the first execution algorithm mechanism.

9. The data modeling system based on research data of claim 1, wherein: The model construction processing of the structured data set by the first execution algorithm mechanism further includes the following specific processes before the model construction processing: Based on the sub-series data representation coefficients of the structured data set, the mapping set between the sub-series data representation coefficients and the modeling sampling ratio is input into the database to obtain the modeling sampling ratio of the structured data set, which is denoted as the specified modeling sampling ratio. Data is extracted from the structured data set at the specified modeling sampling ratio, and the extracted data is denoted as the modeling sample data. The modeling sample data is dynamically cropped to obtain the target modeling sample data. The target modeling sample data is pre-modeled by the first execution algorithm mechanism to obtain the training convergence rate in the pre-modeling process, and the training convergence rate is compared with the preset verification training convergence rate. If the training convergence rate in the pre-modeling process is greater than or equal to the verification training convergence rate, it is judged that the model construction processing of the structured data set is directly performed by the first execution algorithm mechanism, otherwise, the first execution algorithm mechanism is dynamically optimized, and the model construction processing of the structured data set is performed by the dynamically optimized first execution algorithm mechanism.

10. The data modeling system based on research data of claim 1, wherein: The specific process of dynamically optimizing the first execution algorithm mechanism is as follows: The deviation value between the training convergence rate in the pre-modeling process and the verification training convergence rate is calculated, which is denoted as the training convergence rate deviation value in the pre-modeling process. The training convergence rate deviation value in the pre-modeling process is matched with the learning rate adjustment coefficient corresponding to each training convergence rate deviation value interval to obtain the learning rate adjustment coefficient in the pre-modeling process. The target algorithm covered by the first execution algorithm mechanism is calculated, which is denoted as the target algorithm, the current set learning rate of the target algorithm is calculated, and the learning rate adjustment coefficient in the pre-modeling process is coupled to obtain the predicted set learning rate of the target algorithm. The predicted setting learning rate of the target algorithm is compared with the preset upper limit value of the setting learning rate of the target algorithm, if the predicted setting learning rate of the target algorithm is less than the upper limit value of the setting learning rate of the target algorithm, the subsequent model construction is performed by using the predicted setting learning rate of the target algorithm, otherwise, the subsequent model construction is performed by using the upper limit value of the setting learning rate of the target algorithm.