Software robot system for real world research and operation method
By setting up data flow and comprehensive analysis, key information is extracted from multi-source heterogeneous data using natural language processing and image recognition technologies. Data cleaning and outlier detection are performed, solving the flexibility and adaptability problems of existing software robot systems in medical big data processing. This enables efficient and accurate data analysis and real-world research support.
Patent Information
- Application Number
- CN202511026315.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing software robot systems lack flexibility and adaptability in medical big data collection, processing, and analysis tasks, making it difficult to adapt to the complexity and changes in real-world environments. Furthermore, their ability to process non-standardized health and medical data is limited, failing to effectively support real-world research.
A software robot system for real-world research is employed to extract key information from multi-source heterogeneous data by setting up data flow, data extraction, preprocessing, quality control, and comprehensive analysis. It utilizes natural language processing and image recognition technologies to perform data cleaning, normalization, and outlier detection, and combines machine learning methods for data correction and analysis, including descriptive statistics, propensity score matching, survival analysis, and security assessment, to ensure the accuracy and reliability of the analysis results.
It achieves automated data processing, significantly reduces human resource input, improves data processing efficiency and accuracy, provides comprehensive and reliable analysis results in real-world studies, supports real-time monitoring and instant analysis of complex medical data, reduces error rates, and improves the reliability and comparability of analysis results.
Smart Images

Figure CN120913797A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of software robots, in particular to a software robot system for real-world research and a running method. BACKGROUND
[0002] With the improvement of electronic medical record system and medical informatization level, how to efficiently collect, accurately process and flexibly use these massive data has become a problem to be solved. The conventional data collection and processing method depends on manual operation, which is not only inefficient, but also prone to errors, and is difficult to adapt to the requirements of the big data era. Software robots can simulate manual operation and automatically perform rule-based tasks such as data entry and format conversion. However, the existing software robot applications are mainly concentrated on simple repetitive tasks, and there is insufficient support for complex judgment and medical big data collection, processing and analysis tasks, and there are limitations in data engineering process optimization. The current software robot operation process is usually static, lacking flexibility and adaptability, and cannot be optimized according to real-time feedback and continuously changing medical environment. And its application in the real environment still faces the problems of poor environmental adaptability, limited data processing capacity, etc. The solution commonly used for the existing problems is complex customized development, and the processing capacity for non-standardized health medical data is limited.
[0003] Therefore, it is necessary to provide a process automation and intelligent system that can effectively utilize a large amount of multi-source heterogeneous health medical data and conduct real-world research (RWR) based on these data, and achieve the level of senior experts in the field. Real-world research evaluates the effectiveness of interventions (such as drugs, medical devices, etc.) in routine medical practice by analyzing data collected in non-randomized clinical trial settings. These studies can provide real-world evidence and valuable references for understanding and improving health outcomes. SUMMARY
[0004] The present application provides a software robot system for real-world research and a running method to solve the problems in the prior art.
[0005] To achieve the above-mentioned purpose, the technical solutions adopted by the present application are as follows:
[0006] A running method of a software robot system for real-world research, comprising the following steps: step 1, setting up a data flow and performing data extraction operations: configuring the data flow according to the data type and connecting the original data source; extracting key information from the original data source; and using the extracted information for data preprocessing; step 2, preprocessing the data and converting the preprocessed data, and performing data quality control operations on the converted data; step 3, performing comprehensive analysis on the data processed by the quality control operation, and running according to the analysis result, wherein the first step, descriptive statistical analysis is performed, and the user group is grouped: according to the key information, the users are divided into different subgroups, which is helpful to understand the key information of the data and determine the potential confounding factors that need to be controlled; the second step, using propensity score matching method to control potential confounding factors: controlling confounding factors through propensity score matching method to ensure the balance of baseline characteristics between subgroups; the third step, survival analysis: using Kaplan-Meier survival curve and Cox proportional hazards regression model method to perform survival analysis to evaluate the survival situation or event occurrence rate between different subgroups; the fourth step, safety evaluation combined with integrated data: analyzing the safety data collected in the study to evaluate the safety of different treatments or interventions; the fifth step, sensitivity analysis by changing the matching standard and / or adjusting the model parameters to verify the robustness and reliability of the results; the sixth step, according to the analysis results of each angle from the first step to the fifth step, the comprehensive results of the analysis are obtained.
[0007] Based on the above technical solution, further, in step 1, the data extraction process is: the software robot system extracts key information from the original data source by using natural language processing and image recognition technology, wherein the key information at least includes user information, chief complaint, present illness history, past history, laboratory test results and imaging examination results.
[0008] Based on the above technical scheme, further, in step 2, the preprocessing process is: removing irrelevant data and duplicate records in an automated manner, and performing data normalization processing by converting the data into a consistent format or arrangement. The data conversion process is: converting the preprocessed data into a unified or preset format through the built-in mapping tool. The preset rules include: identifying outliers using statistical methods, determining the box position of the data points, and calculating the absolute value of the data point Z-Score; rule one: if a data point is lower than Q1-1.5IQR or higher than Q3+1.5IQR, it is an outlier, wherein IQR=Q3-Q1, Q3 is the upper quartile, Q1 is the lower quartile, and IQR is the interquartile range; rule two: if the absolute value of the Z-Score of a data point is greater than 2, it is an outlier, wherein Z-Score is the measurement unit; wherein if the result of rule one or rule two is an outlier, the data point is determined to be an outlier.
[0009] Based on the above technical scheme, further, in step 2, the data quality control operation includes outlier detection operation and data correction operation; the outlier detection in the data quality control operation uses preset rules and algorithms to analyze the data and identify outliers, missing values and inconsistent values in the data set that are not logical or do not match known patterns; wherein the detection process of outliers is: using statistical methods or machine learning methods to identify data points that deviate from the preset range; the detection and processing process of missing values is: analyzing the pattern of data missing and using corresponding strategies to handle missing values; the detection process of inconsistent values is: checking the logical errors and inconsistencies in the data by setting rules.
[0010] Based on the above technical scheme, further, in step 2, the data correction process in the data quality control operation includes the following steps: step 21, data cleaning: cleaning the data identified as outliers, missing or inconsistent; when the number of outliers < 10% of the total data, replace the outliers with the mean or median of the entire data set; when the number of outliers ≥ 10% of the total data, remove the current feature column, and perform KNN clustering on the remaining features to replace the outliers with the average value of the K nearest neighbors; step 22, data verification: after cleaning, re-verify the data to ensure the effectiveness of the modification measures and the consistency of the data; step 23, data record: record the version before and after data update to ensure the traceability of the data correction process.
[0011] Based on the above technical scheme, further, in step 3, the comprehensive analysis further includes machine learning analysis, and the machine learning analysis process includes the following steps: step 31, data acquisition, classification and preprocessing are performed; wherein the data preprocessing process includes data cleaning, data standardization and normalization, and data set division; step 32, feature engineering processing is performed: including feature selection processing and feature extraction processing; step 33, model selection and model training are performed, wherein k-fold cross-validation is used to evaluate the model performance, and cross-validation is performed; step 34, model evaluation and optimization: when evaluating, the confusion matrix is drawn, the performance indicators are calculated and analyzed; when optimizing, the feature optimization is performed based on feature selection and engineering, and the model parameter optimization is performed based on model parameter adjustment; step 35, model interpretation according to the evaluation and optimization results.
[0012] A software robot system for real-world research adopts a running method of a software robot system for real-world research, comprising a data input module, a data processing module and a data analysis module; and the data input module transmits the processed data to the data processing module for processing, and the extracted data is transmitted to the data analysis module for analysis; the data input module is used for setting data flow, data extraction and data transmission operation; the data processing module is used for data preprocessing, data conversion and data quality control operation; and the data analysis module is used for comprehensive analysis.
[0013] Compared with the prior art, the present application has the following beneficial effects:
[0014] (1) The present application uses the automatic technology of the software robot system to set the data flow and the preprocessing operation, extracts and processes different source data from various data sources, reduces the errors and time consumption of manual data entry, significantly reduces the investment of human resources, improves the data processing and overall work efficiency, and more comprehensively analyzes the performance of the intervention measures in different populations, and automatically analyzes based on these more comprehensive data and reaches the level of experts in the field of real-world research.
[0015] (2) The present application can identify outliers, missing values and inconsistent values that are not logical or do not match known patterns in the data set through the data quality control operation, monitor user conditions and treatment effects in real time, obtain and perform immediate data analysis, reveal the reactions of different user groups to medical intervention measures, and provide a certain reference for improving the identification of which treatment methods are more effective in actual application. And the preset rules and algorithms are used for quality control operation and error correction of the data, which reduces the error rate, reduces the misoperation, and improves the accuracy and reliability of the data.
[0016] (3) The operation method and software robot system disclosed by the present application perform comprehensive analysis on data, and through the analysis of data by two analysis methods, the medical information related to users collected, analyzed and managed in real world research can be applied; compared with the traditional method, the real world research provides evidence closer to the actual clinical situation, reflects the application effect and safety of the treatment method in the real world, and better reveals the real effect of the treatment method in different populations and environments.
[0017] (4) When the present application performs comprehensive analysis, first, descriptive analysis is performed to ensure the comparability of the analysis and avoid the accumulation of subsequent errors to make the results unusable; then, propensity score matching is performed to control the balance of confounding factors between different subgroups and improve the accuracy of the analysis results; then, survival analysis and safety evaluation are performed to give the analysis results. In addition, sensitivity analysis is also used to verify the robustness and reliability of the results, to avoid unreliable analysis results, further ensure the accuracy and reliability of the analysis results, and also provide a basis for the optimization of the overall analysis method. Using this systematic and comprehensive analysis method can make the analysis operation almost error-free, and can realize full-process automatic data processing and analysis, which can save about 83% of the time cost compared with the related non-systematic statistical analysis software for real world research. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 The simple flowchart of the control method of the present application. DETAILED DESCRIPTION
[0019] The present application will be further described and explained with reference to the accompanying drawings and specific embodiments. The technical features of each embodiment in the present application can be combined accordingly without conflict. The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the scope of the present application, so the present application is not limited to the specific embodiments disclosed below.
[0020] EMBODIMENT
[0021] In combination Figure 1As shown, a method for operating a software robot system for real-world research is implemented, which needs to clearly define the main goal of the research in the early stage, such as evaluating the effect of a treatment plan, identifying potential side effects, or evaluating the response of different populations to treatment, based on the expected goal, and setting verifiable hypotheses. Finally, the clinical evidence of the value, potential benefits or risks of the use of medical products in the real world is obtained, thereby providing strong evidence support for decision makers, which not only helps to optimize the treatment plan, but also identifies potential risks at an early stage, thereby better protecting user safety and improving treatment effect, to ensure the accuracy and scientificity of the analysis.
[0022] The specific method includes the following steps:
[0023] Step 1, set up data flow and perform data extraction operation: according to the required data type, configure the data flow and connect the preset original data source, wherein the data source includes electronic medical record system, laboratory information management system, medical influence database, etc. Specifically, the data type refers to the category or format of the data, which can be structured data (such as database table), semi-structured data (such as XML file) or unstructured data (such as text file, image video, etc.); according to the data type, different extraction and transmission methods are adopted.
[0024] Data flow refers to the transmission path of data in the system, including the entire transmission process from data extraction to the target; the data flow defines the path and flow method between the original data source and the target system. The data source is the source of data, which can be various databases, file systems, network services, etc.
[0025] The following is an example:
[0026] Suppose there is a real-world study on a new type of drug for treating hypertension. The relationship between data type, data flow and data source can be described as follows:
[0027] (1) Data type: structured data includes patient's basic information (such as age, height, weight, etc.), drug treatment plan (such as dose, frequency), physiological indicators (such as blood pressure, heart rate) during treatment, etc. Semi-structured data (such as doctor's prescription information, medical record, laboratory test result, etc.).
[0028] (2) The data flow in this scheme can be understood as extracting the patient's basic information, treatment plan and physiological indicator data from the clinical trial database, extracting the patient's medical records, prescription information and laboratory test results through the hospital information system, collecting the patient's physiological indicator data in daily diagnosis and treatment through the electronic health record system, and collecting the patient's subjective feelings and drug compliance data through questionnaire survey or telephone interview, etc.
[0029] (3) The data source is the source of the data, including clinical trial database, hospital information system, electronic health record system, etc. Among them, the clinical trial database stores patient basic information, treatment plan and physiological index data from clinical trials. The hospital information system stores patient medical records, prescription information and laboratory examination results and other data. The electronic health record system records the physiological index data and other medical information of the patient in daily diagnosis and treatment.
[0030] Further, key information is extracted from the original data source. The data extraction process is as follows: the software robot system extracts key information from complex original data sources by using natural language processing (NLP) and image recognition technology (such as OCR), wherein the key information at least includes user information (such as name, gender, age, hometown, contact information, occupation, etc.), chief complaint, present illness history (such as onset condition, illness time, main symptom characteristics, cause and inducement, disease development and evolution, accompanying symptoms, diagnosis and treatment history, etc.), past history, laboratory test results and imaging examination results, etc.; and the extracted information is transmitted to the next step for preprocessing through an encrypted channel, to ensure the security and privacy protection of the data in transmission.
[0031] The specific implementation process includes:
[0032] 1. Obtain system information: system information can be obtained to add system internal meta information to the input data, such as system time.
[0033] 2. Perform table input operation through corresponding table input component: introduce the data content to be queried and converted based on SQL statement through table input.
[0034] 3. Generate records: convert part of the text data into data rows, and each field is a column of a data row. Converting text data to columns of data rows usually involves text parsing and data structuring, which can be achieved by using string processing functions in programming languages, regular expressions, text parsing libraries, etc.; the specific implementation method depends on the format and structure of the text data.
[0035] Suppose there is a piece of text data representing the basic information of each employee, such as: (1) Name: Zhang San, age: 30, gender: male, position: manager; (2) Name: Li Si, age: 25, gender: female, position: technician; (3) Name: Wang Wu, age: 35, gender: male, position: salesman. The specified fields (name, age, gender, position) can be grabbed as a column of a data row by writing code, and the corresponding data can be converted to columns of data rows according to the requirements, and saved to the Data Frame. In this way, the text data is converted to columns of data rows and stored in the format of Data Frame.
[0036] 4. REST client: Request a service interface, the request body is a JSON, the service interface responds to data, and the data is also in JSON format, and the JSON field of the response body is parsed.
[0037] 5. Web service query: Obtain network information through a Web service.
[0038] 6. JSON Input: JSON object file input, which can read files written in data according to the JSON standard.
[0039] Step 2, preprocess the data, and convert the preprocessed data, and perform data quality control operations on the converted data;
[0040] Specifically, the preprocessing process is: the initially received raw data is preprocessed, irrelevant data and duplicate records are removed in an automated manner, and standardized processing is performed to improve data quality; Specifically, irrelevant data usually refers to data that has no analysis or processing significance, which may be data unrelated to the research or task, or data that does not meet the target;
[0041] For example, in a sales report, employee birthday data unrelated to sales may be considered irrelevant data; In text data, there may be irrelevant data such as nonsense, noise, or annotations that are irrelevant to the topic; Redundant data can also be considered irrelevant data. In operation, irrelevant data is removed in an automated manner, which can include using programming scripts, algorithms, or data cleaning tools to identify and delete irrelevant data, which can identify and delete irrelevant data according to pre-set rules, patterns, or machine learning algorithms.
[0042] Standardized processing refers to converting data into a consistent format or arrangement, making data easier to understand, process and analyze, and reducing errors or confusion caused by inconsistent data formats. Standardization standards may include uniform formats (for similar variables), uniform units (for similar variables), and value range standardization.
[0043] The data conversion process is: the cleaned data is converted to a uniform or pre-set format through built-in mapping tools, for example: using the mapping tool of the Pandas library to convert date data to a uniform format; Use the label encoding tool of the Scikit-learn library to convert category data to numerical encoding; Use the vectorization operation of the NumPy library to standardize the data value range to 0 to 1, etc.
[0044] The data quality control operation includes an abnormal value detection operation based on machine learning and an automatic data correction operation to improve data quality and ensure that the underlying data for analysis is accurate.
[0045] Specifically, the abnormal value detection uses preset rules and algorithms to analyze data and identify outliers, missing values, and inconsistent values in the data set that are not logical or do not match known patterns.
[0046] The detection process for outliers is as follows: use statistical methods (such as Z-Score, IQR method) or machine learning methods (such as cluster-based anomaly detection) to identify data points that exceed the preset range in numerical value; the specific implementation logic is: use statistical methods to identify outliers, and calculate the Z-Score absolute value of the data point, and if either method determines that it is an outlier, the data point is determined to be an outlier.
[0047] The judgment rule includes:
[0048] Rule 1: If a data point is lower than Q1-1.5IQR or higher than Q3+1.5IQR, it is an outlier, where IQR=Q3-Q1, and the upper quartile Q3 and the lower quartile Q1 are statistical quantities describing the distribution of data, used to divide the data set into four equal parts. Specifically, Q1 is the 25th percentile of the data, and Q3 is the 75th percentile of the data. When calculating the quartiles, first sort the data values in ascending order, then find the value at the preset percentile. That is, arrange all data from small to large, and the number at the lower 1 / 4 position is called the lower quartile (according to the percentage, that is, the number at the 25% position), also called the first quartile Q1; the number at the upper 1 / 4 position is called the upper quartile (according to the percentage, that is, the number at the 75% position), also called the third quartile Q3. If the size of the data set is even, Q1 and Q3 usually take the average of the two adjacent positions. The inter-quartile range (IQR) is a statistical quantity that describes the distribution range of the middle 50% of the data in the data set. It is the difference between the third quartile (Q3) and the first quartile (Q1). In statistics, it is a method for measuring the dispersion or dispersion of data. The calculation method of the inter-quartile range is as follows: IQR=Q3-Q1.
[0049] Rule 2: If the absolute value of the Z-Score of a data point is greater than 2, it is an outlier; if either rule 1 or rule 2 determines that it is an outlier, the data point is determined to be an outlier. Where Z-Score is a measurement unit that represents the distance of a data point from the mean, Z-Score=(x-μ) / σ.
[0050] The detection and processing procedure of missing values is: analyzing the mode of data missing, such as using completely random missing mode, random missing mode or non-random missing mode, wherein, if the missing probability of missing variable observation value is irrelevant to itself or other variables included in the study, the missing data is completely random missing; if the missing probability of missing variable observation value is related to other variables included in the study, and is irrelevant to itself under the control of variables included in the study, the missing data is random missing; if the missing data is neither completely random missing nor random missing, it is non-random missing; and corresponding strategies are used to process missing values, such as using average filling, hot card filling, prediction model interpolation and other calculation methods.
[0051] The detection procedure of inconsistent values is: checking logical errors and inconsistencies in the data by setting rules; for example, the user's treatment end date should not be earlier than the start date.
[0052] Further, the data quality control operation includes data correction operation based on machine learning, which automatically corrects or labels the detected errors or abnormalities for further review and processing.
[0053] The data correction procedure includes a correction step after error identification, and the correction procedure includes the following steps:
[0054] Step 21, data cleaning: cleaning the data identified as outliers, missing or inconsistent, and replacing, modifying or deleting by appropriate methods. The specific implementation is: when the number of abnormal values < 10% of the total data, replace the abnormal values with the mean or median of the entire data set; when the number of abnormal values ≥ 10% of the total data, remove the current feature column, and the remaining features are clustered by KNN, and the average value of K neighbor samples is used to replace the abnormal value.
[0055] Step 22, data verification: after cleaning, re-verify the data to ensure the effectiveness of the modification measures and the consistency of the data; specifically, the data verification procedure is still repeated to verify the data quality control steps described above. If the data verification result shows that the data quality cannot meet the requirements, redefining the data source, resetting the data flow or manual intervention can be used to adjust it to meet the requirements.
[0056] Step 23, data record: record the pre-update and post-update versions of the data to ensure the traceability of the data correction procedure.
[0057] Step 3, comprehensive analysis of the data processed by the quality control operation, and running according to the analysis result.
[0058] For example, Case 1: A diabetes treatment drug has shown certain effectiveness and safety in actual use, and the drug has shown good efficacy in clinical trials. To further understand its performance in a broader patient population, the statistical analysis process in a specific comprehensive analysis is as follows:
[0059] First step, descriptive statistical analysis, grouping of user population: according to patient characteristics, that is, key information such as age, gender, disease duration, comorbidities, etc., the user, that is, the patient, is divided into different subgroups, which helps to understand the key information of the data and determine the potential confounding factors that need to be controlled. Compare the key information of patients using the drug (i.e. treatment group) with patients using other treatment methods (i.e. control group) to ensure comparability. Among them, "comparability" means that the treatment group (patients using new drugs) and the control group (patients using other treatment methods) should be as similar as possible in the following key characteristics: such as in terms of age: the age distribution of patients in the two groups should be similar to avoid the potential impact of age on treatment effectiveness. In terms of gender: the gender ratio should be similar to prevent gender differences from affecting the research results. In terms of disease duration: the duration of diabetes (i.e. the length of time since diabetes diagnosis) should be similar, as the length of time may affect treatment response. In terms of comorbidities: the comorbidities (such as hypertension, heart disease, etc.) of patients in the two groups should be similar, as comorbidities may affect the effectiveness and safety of diabetes treatment. Other characteristics that may affect treatment effectiveness: such as body mass index (BMI), blood glucose control level (such as HbA1c value), lifestyle (such as diet and exercise habits), etc.
[0060] And the method to ensure "comparability" includes:
[0061] 1. Randomization: by randomly assigning patients to treatment and control groups, the impact of known and unknown confounding factors is reduced.
[0062] 2. Matching: match patients according to baseline characteristics so that the two groups are as similar as possible in these characteristics.
[0063] 3. Statistical control: use statistical methods (such as multivariate regression analysis, propensity score matching, etc.) to control for differences in baseline characteristics during the analysis phase. Other treatment methods can choose standard treatment options, which in clinical practice refer to a set of widely accepted and recognized treatment procedures and guidelines for specific diseases or symptoms; these options are usually based on the latest clinical research and evidence, and are published and updated by professional organizations or organizations; the development of standard treatment options aims to provide consistent treatment standards to ensure that patients receive the best medical care and help doctors make informed treatment decisions in clinical practice.
[0064] The content of the standard treatment plan usually includes drug treatment, surgical treatment, rehabilitation plan, nutrition guidance, and suggestions on monitoring and follow-up; these plans may also take into account the needs of specific patient groups, such as children, the elderly, or pregnant women, etc.; the standard treatment plan is designed to improve the consistency and quality of treatment and to help healthcare professionals better provide treatment services for patients.
[0065] Second step, propensity score matching: using propensity score matching method to control potential confounding factors. Under the same conditions, match patients using new drugs (treatment group) and patients using other treatment methods (control group) to make the two groups as similar as possible in key information. Specifically, using the propensity score matching (PSM) method can effectively control these confounding factors. By controlling confounding factors, the propensity score matching method can improve the internal validity of the study, making the evaluation of the effect of new drugs more accurate and reliable, and ensuring the balance of subgroups in baseline characteristics, which is the key information. Other methods such as multivariate adjustment method can also be used to control confounding factors.
[0066] And the steps of the PSM include:
[0067] 1. Calculate propensity score: first, select variables, then use the statistical method of logistic regression to calculate the probability of each patient receiving new drug treatment under known confounding factors, i.e. propensity score. The statistical method of logistic regression is the existing method, which will not be described in detail here.
[0068] 2. Matching process: match patients in the treatment group and the control group according to the propensity score to make the two groups as similar as possible in confounding factors. For example, one-to-one matching, caliper matching or other matching methods can be used. These matching methods are existing methods, which will not be described in detail here.
[0069] 3、Verify the matching effect: After matching is completed, check the balance of the two groups on the confounding factors to ensure that the matching is successful. That is, if the distribution of these confounding factors in the two groups is similar, it can be considered that the matching is successful, and the comparison and analysis of the treatment effect on the matched data set are carried out; among them, the specific matching steps and the judgment standard of whether the matching is successful are: when the current and posterior significantly reduces (usually less than 0.1 is considered acceptable), it means that the matching effect is good. Use the balance degree visualization comparison analysis (Love plot) to compare the distribution changes of each confounding factor before and after matching. If the SMD of most confounding factors after matching is close to zero, it indicates that the balance is improved. Use t-test for continuous variables and chi-square test for categorical variables to check whether the difference between the two groups before and after matching is significant. After successful matching, these tests should not be significant (p value is large). Compare the distribution of each confounding factor before and after matching by visualization method. The closer the distribution, the better the balance. If the verification result shows that the two groups are balanced on most confounding factors after matching, and the difference is small, it can be considered that the matching is successful.
[0070] Further, in clinical research, confounding factors (also known as confounding variables) refer to those factors that are related to both exposure (such as the use of new drugs) and outcome (such as treatment effect); confounding factors can affect the effectiveness and reliability of research results, so they need to be controlled in research design and analysis. For the case study of the diabetes treatment drug, possible confounding factors include: 1. Patient age: Patients of different ages may have different responses to treatment. 2. Gender: Men and women may differ in disease progression and treatment response. 3. Disease duration: The length of time since diabetes diagnosis may affect the severity of the disease and response to treatment. 4. Comorbidities: such as hypertension, heart disease, kidney disease, etc., which may affect treatment efficacy and safety. 5. Baseline glycemic control level: such as HbA1c value, patients with different glycemic control levels may have different responses to treatment. 6. Body mass index (BMI): obesity or underweight may affect drug metabolism and treatment effect. 7. Lifestyle: such as eating habits, exercise, smoking and drinking, etc., these factors may affect diabetes management and treatment effect. 8. Drug adherence: whether patients take medicine regularly according to doctor's advice, poor adherence may affect treatment effect. 9. Socioeconomic factors: such as income, education level, medical insurance, etc., which may affect patients' ability to access medical resources and health management. 10. Baseline health status: such as whether there are other chronic diseases or health problems, which may affect overall treatment response. 11. Drug type and dosage: the control group may use different treatment methods, so the type and dosage of drugs used need to be controlled.
[0071] Third step, survival analysis: Kaplan-Meier survival curve and Cox proportional hazards regression model method for survival analysis to assess the survival of different subgroups or event rates; in the study of drugs for diabetes, Kaplan-Meier survival curve and Cox regression analysis method can be used for survival analysis to evaluate the effect of new drugs on patient survival rate.
[0072] Specifically, Kaplan-Meier survival curve can be used to analyze the following variables: (1) treatment group: new drug group and control group. (2) Age: such as <50 years old, 50-65 years old, >65 years old. (3) Gender: male and female. (4) Disease duration: short duration (such as <5 years) and long duration (such as ≥5 years). (5) Comorbidities: with or without specific comorbidities (such as heart disease, hypertension). (6) Baseline glycemic control level: such as HbA1c low, medium, high. (7) Body mass index (BMI): such as normal weight, overweight, obesity.
[0073] And the results of Kaplan-Meier survival curve analysis usually include: (1) Survival curve: shows the survival probability of patients in each group at different time points during the study. (2) Median survival time: the median survival time of each group, that is, the time point at which 50% of patients are still alive. (3) Survival probability: survival probability at different time points, showing the difference between groups. (4) Log-rank test: used to compare whether there is a significant difference in survival curves between different groups. Further, Cox regression analysis is also used for analysis. Cox regression analysis (proportional hazards model) is used to consider the effects of multiple variables simultaneously to assess their independent effects on survival rate.
[0074] The variables that can be analyzed include: (1) treatment group (main variable): new drug group and control group. (2) Age: as a continuous variable or categorical variable (such as different age groups). (3) Gender: male and female. (4) Disease duration: as a continuous variable or categorical variable (such as short duration and long duration). (5) Comorbidities: no specific comorbidities, or the number of comorbidities. (6) Baseline glycemic control level: such as HbA1c value. (7) Body mass index (BMI): as a continuous variable or categorical variable (such as normal weight, overweight, obesity). (8) Other potential confounding factors: such as lifestyle (diet, exercise), socioeconomic factors (income, education level), etc.
[0075] The results of this Cox regression analysis typically include: (1) Hazard Ratio (HR): the effect of each variable on the risk of survival. HR > 1 indicates an increased risk, while HR < 1 indicates a decreased risk. (2) 95% Confidence Interval: the confidence interval for the hazard ratio, showing the accuracy of the estimate. (3) p-value: assessing the significance of the effect of each variable on survival, with p < 0.05 generally indicating significance. (4) Adjusted survival curve: the survival curve adjusted according to the regression model, reflecting the survival situation after multivariate adjustment.
[0076] A specific application example, assuming the research results are as follows: (1) Based on the Kaplan-Meier survival curve analysis, the median survival time of the treatment group is 5 years, and that of the control group is 4 years. The log-rank test shows that there is a significant difference between the survival curves of the two groups (p < 0.01). Analysis of different age groups shows that the treatment effect of the <50 age group is significantly better than that of the control group, while the >65 age group shows no significant difference. (2) Based on Cox regression analysis, the hazard ratio of the treatment group is 0.75 (95% CI: 0.60-0.90, p = 0.002), indicating that the use of the new drug significantly reduces the risk of death by 25%. The hazard ratio of age is 1.03 (95% CI: 1.01-1.05, p = 0.01), indicating that for every additional year of age, the risk of death increases by 3%. The hazard ratio of patients with comorbidities is 1.50 (95% CI: 1.20-1.80, p < 0.001), indicating that comorbidities significantly increase the risk of death by 50%. Through the above analysis, we can more comprehensively understand the efficacy and safety of the new drug in different patient populations, providing strong support for clinical decision-making.
[0077] Step 4: Safety evaluation based on integrated data: analyze the safety data collected during the study to evaluate the safety of different treatments or interventions. Compare the incidence of adverse events between the new drug and other treatment methods. The "integrated data" refers to data collected and aggregated from multiple sources. In drug safety evaluation, this may include information from clinical trials, epidemiological studies, drug monitoring databases, and other relevant materials. These data are integrated together to comprehensively evaluate the side effects and incidence of adverse events of the drug.
[0078] By integrating data, researchers can more comprehensively understand the safety characteristics of the drug, including its response in different populations, potential risk factors, and advantages and disadvantages compared to other treatment methods. This helps medical professionals and decision-makers better understand the safety of the drug and provide more comprehensive treatment recommendations for patients.
[0079] Step 5, Sensitivity Analysis: To test the robustness and reliability of the results, sensitivity analysis is performed by changing the matching criteria and / or adjusting the model parameters according to the actual situation. In this diabetes treatment drug study, some preliminary conclusions have been drawn through propensity score matching and survival analysis, that is, the new drug significantly reduces the risk of death in patients. However, in order to ensure the robustness of these results, sensitivity analysis is needed to test the impact of various assumptions on the research results. Through sensitivity analysis, the robustness of the research results can be ensured, and the results are still reliable and meaningful even under different assumptions and conditions.
[0080] Assumption Scenario One, Method of Dealing with Missing Values: Assuming that there is a certain proportion of missing values in the original data, the "mean imputation" method is used in the preliminary analysis; in order to test the robustness of the results, another common method of dealing with missing values, "multiple imputation", is used. Among them, multiple imputation: generate multiple data sets to fill in missing values, each data set is based on different assumptions, then analyze and integrate the results respectively. The corresponding result analysis is: if the analysis results (hazard ratio, p value, etc.) after multiple imputation are consistent with or change little (for example, the hazard ratio changes from 0.75 to 0.78, and is still significant) after mean imputation, the results are robust. If the results change significantly or lose significance (for example, the hazard ratio changes from 0.75 to 1.10, and is not significant), the robustness of the results is poor.
[0081] Assumption Scenario Two, Subgroup Analysis: Assuming that in the preliminary analysis, it is found that the new drug has different effects in different age groups. In order to test the robustness of the results, patients are divided into different subgroups (such as <50 years old, 50-65 years old, >65 years old), and analyzed respectively. In the subgroup analysis process, the hazard ratio and p value of each age group are calculated respectively. The corresponding result analysis is: if the analysis results of each subgroup are consistent with the overall analysis results (for example, the hazard ratio of each age group is between 0.70-0.80, and is significant), the results are robust. If the results of some subgroups are significantly different (for example, the hazard ratio of the <50 years old group is 0.50, but the hazard ratio of the >65 years old group is 1.20, and is not significant), the robustness of the results may be a problem, and further exploration of the heterogeneity between subgroups is needed.
[0082] Hypothetical Scenario 3, Different Criteria for Control Group Selection: In the initial analysis, patients who received a treatment were selected as the control group. To test the robustness of the results, the selection criteria for the control group were changed, for example, only patients who had received a specific standard treatment (e.g., metformin) were selected as the control group. The propensity score matching and survival analysis were repeated with the changed control group criteria. The corresponding result analysis is as follows: if the results of the analysis with the changed control group criteria are consistent with the original analysis (e.g., the hazard ratio is still between 0.70 and 0.80 and is significant), it indicates that the results are robust. If the results change significantly (e.g., the hazard ratio becomes 1.05 and is not significant), it indicates that the robustness of the results is poor, and the selection criteria for the control group may have a significant impact on the results.
[0083] Further comprehensive examples, assume that in the initial analysis, the missing values are filled with the mean, and the hazard ratio of the treatment group is 0.75 (95% CI: 0.60-0.90, p=0.002). The sensitivity analysis process is as follows: use multiple imputation to handle missing values, and repeat the propensity score matching and Cox regression analysis. The results are compared: multiple imputation results: hazard ratio is 0.78 (95% CI: 0.62-0.94, p=0.003). The results are explained: (1) good robustness: the hazard ratio changes from 0.75 to 0.78, and is still significant (p<0.01), indicating that the method for handling missing values has little effect on the research results, and the results have good robustness. (2) poor robustness: if the multiple imputation results show that the hazard ratio is 1.05 (95% CI: 0.85-1.25, p=0.25), it indicates that the method for handling missing values has a significant impact on the results, and the robustness of the results is poor, and the analysis method and data processing process need to be re-examined.
[0084] Comprehensive analysis also includes machine learning analysis. The main goal of using machine learning analysis in this case is to use machine learning models to predict the effects of different drugs in treating diabetes, thereby optimizing treatment strategies. The secondary goal is to identify key factors that affect the effectiveness of drug treatment and provide the basis for personalized treatment.
[0085] The specific analysis process is as follows:
[0086] Step 31, data acquisition, classification and preprocessing, the process is as follows:
[0087] Step 311, acquire the required data: acquire the required data from the corresponding data sources. For example, patient medical records, laboratory test results, imaging data, etc. can be obtained from electronic health records (EHR); past and ongoing clinical trial data on diabetes can be obtained from clinical trial data; information such as treatment dosage, side effects, and efficacy of different drugs can be obtained from drug databases.
[0088] Step 312, data type classification: for example, age, gender, race, etc. data are classified as demographic data; medical history, family history, lifestyle (diet, smoking, alcohol), comorbidities, etc. are classified as clinical data; pathological diagnosis results, imaging data, etc. are classified as diagnostic data; drug regimen, dosage, treatment duration, adverse reactions, etc. are classified as treatment data; follow-up information after treatment, including efficacy evaluation, recurrence, quality of life, etc. are classified as follow-up data.
[0089] Step 313, data preprocessing. The preprocessing process includes data cleaning, data standardization and normalization, and data set division steps. Among them, the data cleaning steps include: ① missing value processing: deleting records or features containing a large number of missing values, or using mean, median, interpolation method to fill in missing values, the specific process is the same as step 2, which will not be described here. ② Abnormal value processing: detect and process abnormal values (such as through IQR method, Z-score method), the specific process is the same as step 2, which will not be described here. ③ Data type conversion: convert categorical data to numerical data (such as one-hot encoding).
[0090] The steps of data standardization and normalization include: ① Standardization: standardizing numerical features (such as subtracting the mean and dividing by the standard deviation). ② Normalization: scaling data to a specific range (such as [0, 1]). The process of data set division is: usually according to the proportion of 70% (training set), 15% (validation set), 15% (test set) to divide the training set, validation set and test set.
[0091] Step 32, feature engineering processing, including feature selection processing and feature extraction processing; specifically:
[0092] Step 321, feature selection: correlation analysis is performed to calculate the correlation coefficient of the feature and the target variable, and low correlation features are removed. Chi-square test is performed, for categorical variables, significant features are selected using chi-square test. And recursive feature elimination (RFE) is used to recursively eliminate unimportant features and retain important features.
[0093] Step 322, feature extraction: through PCA dimension reduction, the number of features is reduced, the main information is retained, and principal component analysis (PCA) is performed. Through factor analysis, latent factors are extracted for factor analysis.
[0094] Step 33, model selection and training:
[0095] Step 331, in the model selection process, involves algorithm selection and benchmark model training. Specifically, the algorithm selection process is: selecting appropriate data features and research objectives, such as linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), neural network, XGBoost, etc. The benchmark model construction process is: constructing a simple benchmark model (such as linear regression, decision tree) for preliminary verification.
[0096] Step 332, in the model training process, involves cross-validation and hyperparameter optimization. Specifically, k-fold cross-validation is used to evaluate model performance, select the best parameters, and perform cross-validation. Grid search or random search is used to optimize model hyperparameters.
[0097] Step 34, model evaluation and optimization: during evaluation, draw a confusion matrix, calculate performance indicators and analyze; during optimization, optimize features based on feature selection and engineering, and optimize model parameters based on model tuning.
[0098] Specifically, step 341, in the model evaluation process, specifically, select appropriate performance indicators according to the task, such as accuracy, precision, recall, F1 value, AUROC for classification problems, and mean square error (MSE), root mean square error (RMSE), R 2 For classification problems, draw a confusion matrix and analyze the classification error types of the model.
[0099] Step 342, in the model optimization process, involves feature selection and engineering, and model tuning, where feature selection and engineering are used to further optimize features, such as combining features and interacting features. Model tuning is based on validation set results to further optimize model parameters.
[0100] Step 35, model interpretation: for example, use SHAP values to explain the contribution of each feature to the prediction result. Use LIME technology to explain the local model prediction result. Analyze the feature importance ranking in the model and identify key features.
[0101] Step 6, based on the analysis results of each angle from steps 1 to 5, obtain the comprehensive results of the analysis: based on this case, through the comprehensive analysis of each angle above, the overall effect and safety evaluation of the new diabetes drug in the real world can be obtained. For example, it may be found that the drug significantly reduces blood sugar levels in a specific patient population, while the incidence of side effects is lower than other treatment methods. The results can be used to guide clinical practice and decision-making.
[0102] For example, Case 2: To evaluate the effectiveness and safety of a drug (Drug W) for treating precancerous lesions of the stomach in a real clinical environment, a real-world study of a drug for treating precancerous lesions of the stomach is conducted. The target sample size selected is 2000 patients, of which 1000 patients are in the treatment group treated with the drug and 1000 patients are in the control group treated with the standard treatment regimen. The inclusion criteria for these patients are: age between 30-75 years old, diagnosed with precancerous lesions of the stomach (including intestinal metaplasia, mild or moderate dysplasia, confirmed by gastroscopy and biopsy), started treatment within the past 6 months, no confirmed gastric cancer, severe liver and kidney dysfunction, severe cardiovascular disease, and other major diseases, no allergy to the study drug or contraindications, and complete electronic health records. The primary health outcomes are: improvement or regression of precancerous lesions of the stomach (evaluated by gastroscopy and biopsy) and incidence of progression to gastric cancer. Secondary health outcomes are: symptom relief (such as stomach pain, indigestion, loss of appetite, etc.), overall quality of life (evaluated by patient-reported outcomes), and incidence of adverse reactions and side effects. Key indicators include: baseline and follow-up gastroscopy results, biopsy results (degree of intestinal metaplasia and dysplasia), symptom scores (based on standardized questionnaires such as stomach pain, indigestion, loss of appetite, etc.), overall quality of life scores, and incidence of adverse reactions (such as abdominal pain, diarrhea, allergic reactions, etc.). Clinical characteristics (key information) include: age, gender, race, weight, height, body mass index (BMI), baseline symptom scores, medical history (such as Helicobacter pylori infection, history of gastric ulcer), and concomitant medication use. Based on these data, the comprehensive analysis process is as follows:
[0103] Step 1: Descriptive statistical analysis, group the user population: describe the key information of the patients (such as age, gender, baseline symptom scores, etc.); use mean, standard deviation, median, quartile, etc. to describe continuous variables; use frequency and percentage to describe categorical variables. Specifically, continuous variables refer to those variables that can take any value, usually quantitative data, such as: age (in years), weight (in kilograms), height (in centimeters), body mass index (BMI) (weight divided by the square of height), baseline symptom scores (such as stomach pain, indigestion, loss of appetite scores), overall quality of life scores, etc. Categorical variables refer to those variables that can be divided into different categories, usually qualitative data; for example: gender (male, female), race (such as white, black, Asian, white or yellow, etc.), medical history (such as presence or absence of Helicobacter pylori infection, history of gastric ulcer), concomitant medication use (such as presence or absence of other drug use), incidence of adverse reactions (such as presence or absence of abdominal pain, diarrhea, allergic reactions, etc.). For continuous variables, commonly used statistical quantities include mean, standard deviation, median, and quartile.
[0104] Described as follows: the mean is the average of the variable, such as the average of the age, etc.; the standard deviation is the degree of dispersion of the data, such as the degree of dispersion of the age can be obtained by the standard deviation; the median is the middle value of the data, which can obtain the middle value of the continuous variable such as age, weight, etc. Quartile: the quartile value of dividing the data into four parts (25th and 75th percentile).
[0105] And for the classification variable, the commonly used statistical quantities include frequency and percentage. Described as follows: the frequency refers to the number of samples in each category. The percentage refers to the percentage of the number of samples in each category in the total samples. Example (assuming the following data): the mean of age is 55.2 years, the standard deviation is 10.5 years, the median is 56 years, and the quartile is 48 years and 63 years. The mean of weight is 70.3 kg, the standard deviation is 12.1 kg, the median is 69 kg, and the quartile is 61 kg and 78 kg. There are 1080 males (54%) and 920 females (46%) in gender. There are 1200 whites (60%), 400 blacks (20%), 300 blacks, whites or Asians (15%) in the Asian region, and 100 others (5%) in the race. There are 1500 (75%) with a history of previous disease (Helicobacter pylori infection) and 500 (25%) without. There are 800 (40%) with combined medication and 1200 (60%) without. The incidence of adverse reactions is 200 (10%) for abdominal pain, 150 (7.5%) for diarrhea, 50 (2.5%) for allergic reactions, and 1600 (80%) for no adverse reactions. Through these descriptive statistical analyses, the key information of the patients can be comprehensively understood, and the user groups can be grouped, providing basic data for subsequent analysis and comparison.
[0106] The second step is propensity score matching (PSM): using the propensity score matching method, matching the patients in the treatment group and the control group, controlling the potential confounding factors, which can include age, gender, race, weight, height, body mass index (BMI), baseline symptom score (such as stomach pain, indigestion, loss of appetite, etc.), history of previous disease (such as Helicobacter pylori infection, history of gastric ulcer), combined medication, baseline quality of life score, etc.; the propensity score is calculated by Logistic regression, and 1:1 matching is performed. Specifically, the steps of calculating the propensity score by Logistic regression are as follows: assuming there is a data set including 2000 patients, of which 1000 are treated with drug W (treatment group) and the other 1000 are treated with standard treatment (control group).
[0107] The specific matching step is: 1. Calculate the propensity score: First, select variables, select the mentioned confounding factors as independent variables, and the treatment group (receive drug W treatment) and the control group (receive standard treatment) as dependent variables (0 and 1). Construct a logistic regression model and use the logistic regression (Logistic Regression) method to calculate the propensity score of each patient receiving drug W treatment.
[0108] 2. Perform matching processing: Based on the propensity score of each patient calculated by the logistic regression model, sort the patients in the treatment group and the control group according to the propensity score. According to the propensity score, 1:1 matching is performed to make the distribution of these confounding factors in the treatment group and the control group as similar as possible (i.e. match each patient in the treatment group with the patient in the control group with the closest propensity score).
[0109] 3. Verify the matching result: After matching, the effectiveness of the matching result needs to be verified. The effect of matching can be evaluated by comparing the baseline characteristic distribution (such as age, gender, race, BMI, etc.) before and after matching. Standardized Mean Difference (SMD) can also be used to evaluate the balance of confounding factors before and after matching, and the smaller the SMD value, the better the effect of matching. For example, assume there is the following partial data, as shown in Table 1 below, where "1" in the group column of Table 1 represents the treatment group and "0" represents the control group:
[0110] Table 1
[0111]
[0112] Based on the data in Table 1, the propensity score of each patient is calculated by logistic regression: the propensity score of patient 1 (treatment group) is 0.65; the propensity score of patient 2 (control group) is 0.60; the propensity score of patient 3 (treatment group) is 0.70; the propensity score of patient 4 (control group) is 0.66. Then, compare and analyze the treatment effect of patients with similar matching propensity scores, and based on the above data in this case, the following matching can be performed: First, match the closest patient pair, i.e. patient 1 and patient 4 (difference 0.01), patient 1 (treatment group, propensity score 0.65) and patient 4 (control group, propensity score 0.66) are matched. Patient 3 (treatment group, propensity score 0.70) and patient 2 (control group, propensity score 0.60) are matched.
[0113] Step 3, survival analysis: use Kaplan-Meier survival curve to analyze the incidence of progression to gastric cancer; use Cox proportional hazards regression model to evaluate the risk of gastric cancer progression. It is used to evaluate the survival or event incidence between different subgroups. The specific analysis idea is the same as case 1, which will not be described in detail here.
[0114] Step 4: Safety assessment based on integrated data: Analyze the safety data collected in the study to assess the safety of different treatments or interventions. Specifically, use the integrated data to assess the difference in the incidence of adverse reactions and side effects between groups. The relevant statistical data can be used to study their safety in the real world; use the chi-square test or Fisher's exact test to compare the differences between categorical variables. Suppose we are interested in two groups of patients in the study: the treatment group (receiving drug W treatment) and the control group (receiving standard treatment), and we want to assess the difference in the incidence of adverse reactions (categorical variable) between the two groups of patients.
[0115] Specific steps are as follows: Step A, two categorical variables can be set: categorical variable 1, whether adverse reactions occur (yes / no): there are 50 cases of adverse reactions in the treatment group and 30 cases of adverse reactions in the control group; categorical variable 2, side effect type (mild / severe): there are 20 cases of severe side effects in the treatment group and 15 cases of severe side effects in the control group.
[0116] Step B, use the chi-square test to compare the difference in the incidence of adverse reactions: ① Set hypotheses: (1) null hypothesis (H0): the incidence of adverse reactions in the treatment group and the control group is the same; (2) alternative hypothesis (H1): the incidence of adverse reactions in the treatment group and the control group is different. ② Calculate the chi-square value: create a 2x2 contingency table corresponding to the occurrence of adverse reactions and the treatment group and control group, as shown in Table 2 below, and then use the chi-square test formula to calculate the chi-square value.
[0117] Step C, test the hypothesis: at the significance level (usually 0.05), compare the calculated chi-square value with the critical value corresponding to the degrees of freedom in the chi-square distribution table. If the chi-square value is greater than the critical value, reject the null hypothesis, indicating that there is a significant difference in the incidence of adverse reactions between the treatment group and the control group.
[0118] Specifically, the process for comparing the difference in the incidence of severe side effects using Fisher's exact test is as follows:
[0119] (1) Set hypotheses: null hypothesis (H0): the incidence of severe side effects in the treatment group and the control group is the same. Alternative hypothesis (H1): the incidence of severe side effects in the treatment group and the control group is different.
[0120] (2) Calculate the p-value of Fisher's exact test: use the Fisher's exact test formula to calculate the p-value.
[0121] (3) Test the hypothesis: at the significance level (usually 0.05), compare the calculated p-value with the set significance level.
[0122] If the p-value is less than the significance level, the null hypothesis is rejected, indicating that there is a significant difference in the incidence of serious adverse effects between the treatment group and the control group. Step D, result interpretation and safety relationship: if the results of the chi-square test and the Fisher's exact test show that there is a significant difference in the incidence of adverse reactions or side effects between the treatment group and the control group, it may mean that there is a certain problem with the safety of drug W, because a higher incidence of adverse reactions or serious side effects may reduce the safety of the drug. On the other hand, if there is no significant difference in the incidence of adverse reactions or side effects between the two groups, it can be considered that drug W is comparable to the treatment method in terms of safety, with higher safety. In summary, the chi-square test and the Fisher's exact test can help evaluate the difference in categorical variables between different groups, which is closely related to the safety of the drug, and a larger difference may mean lower safety, and a smaller difference means higher safety.
[0123] Table 2 contingency table
[0124] Category Adverse reaction occurrence (unit: group) Adverse reaction occurrence (unit: group) Treatment group 50 950 Control group 30 970
[0125] Step 5, sensitivity analysis: sensitivity analysis is performed by changing the matching criteria and / or adjusting the model parameters according to the actual situation, etc. to test the influence of various assumptions on the research results, in order to verify the robustness and reliability of the results. The specific case is: subgroup analysis is performed to evaluate the treatment effect of different subgroups (such as different age groups, different baseline symptom scores); test the influence of different assumptions on the research results to ensure the robustness and reliability of the results. Hypothesis 1: age has an impact on the treatment effect of drug W. Specifically, the influence of hypothesis 1 on the research results: if age has a significant impact on the treatment effect of drug W, then in patients of different ages, the treatment effect of drug W may be different, thereby affecting the robustness of the research results.
[0126] Perform result scenario analysis: (1) high robustness case: the assumed research results show that the treatment effect of drug W is significant in patients of different ages, with little difference. In this case, no matter the patient's age is young or old, the treatment effect of drug W is good, and the research results have high robustness.
[0127] (2) low robustness case: the assumed research results show that the treatment effect of drug W is significant in young patients, but poor in old patients. In this case, the age factor has a greater impact on the research results, and the robustness of the research results is low.
[0128] Specifically, the process of sensitivity analysis includes the following steps:
[0129] 1) Data Preparation: Divide the samples into different subgroups based on age, such as young group (30-50 years old), middle-aged group (51-65 years old), elderly group (66-75 years old), etc.
[0130] 2) Subgroup Analysis: Analyze each age subgroup and compare the treatment effects of drug W in different age groups. Survival analysis, symptom improvement analysis, and other methods can be used.
[0131] 3) Test Different Hypotheses: In subgroup analysis, test different hypotheses, such as whether age has a significant impact on the treatment effect of drug W. The impact of the hypothesis can be evaluated by comparing the treatment effect differences of different age subgroups.
[0132] 4) Robustness Evaluation: Based on the results of subgroup analysis and hypothesis testing, evaluate the robustness of the research results. If the treatment effect of drug W is not significantly different in different age subgroups, it means that the research results have high robustness.
[0133] Example: Suppose the research drug W has a significant difference in treatment effect between young and old patients. High robustness: Through sensitivity analysis, it is found that the treatment effect of drug W is significant in different age subgroups, and the difference is not large. In this case, no matter the patient's age is young or old, the treatment effect of drug W is good, and the research results have high robustness. Low robustness: Through sensitivity analysis, it is found that the treatment effect of drug W is significant in young patients, but poor in old patients. In this case, the age factor has a significant impact on the research results, and the robustness of the research results is low. Through the above analysis, the impact of age on the treatment effect of drug W can be evaluated, and the robustness of the research results can be judged.
[0134] Comprehensive analysis also includes machine learning analysis. The main goal of using machine learning analysis in this case is to use machine learning models to predict the effects of different drugs on precancerous lesions of gastric cancer, thereby optimizing treatment strategies. The secondary goal is to identify key factors affecting drug treatment effects and provide the basis for personalized treatment. The specific analysis process is as follows:
[0135] Step 31, data acquisition, classification and preprocessing, the process is as follows:
[0136] Step 311, acquire the required data: acquire the required data from the corresponding data sources. For example, patient medical records, laboratory test results, imaging data, etc. can be obtained from electronic health records (EHR); past and ongoing clinical trial data on precancerous lesions of gastric cancer can be obtained from clinical trial data; information such as treatment dose, side effects, efficacy of different drugs can be obtained from drug databases.
[0137] Step 312, data type classification: for example, age, gender, race, etc. data are classified as demographic data; medical history, family history, lifestyle (diet, smoking, alcohol), comorbidities, etc. are classified as clinical data; pathological diagnosis results, gastroscopy results, imaging data, etc. are classified as diagnostic data; drug regimen, dosage, treatment duration, adverse reactions, etc. are classified as treatment data; follow-up information after treatment, including efficacy evaluation, recurrence, quality of life, etc. are classified as follow-up data.
[0138] Step 313, data preprocessing. The preprocessing process includes data cleaning, data standardization and normalization, and data set division steps. Among them, the data cleaning step includes: ① Dealing with missing values: deleting records or features containing a large number of missing values, or using mean, median, interpolation method to fill in missing values, the specific process is the same as step 2, which will not be described here. ② Processing of abnormal values: detecting and processing abnormal values (such as through IQR method, Z-score method), the specific process is the same as step 2, which will not be described here. ③ Data type conversion: converting categorical data into numerical data (such as one-hot encoding). The data standardization and normalization step includes: ① Standardization: standardizing numerical features (such as subtracting the mean and dividing by the standard deviation). ② Normalization: scaling data to a specific range (such as [0, 1]). The data set division process is: usually according to the proportion of 70% (training set), 15% (validation set), 15% (test set) to divide the training set, validation set and test set.
[0139] Step 32, feature engineering processing, including feature selection processing and feature extraction processing; specifically:
[0140] Step 321, feature selection: correlation analysis is performed to calculate the correlation coefficient of the feature and the target variable, and low correlation features are removed. Chi-square test is performed, for categorical variables, significant features are selected using chi-square test. And using recursive feature elimination (RFE), recursively eliminating unimportant features and retaining important features. For example, key features (i.e. independent variables) include age (patient's age), gender (patient's gender), smoking history (whether there is a history of smoking), drinking history (whether there is a history of drinking), family history (whether there is a family history of gastric cancer), eating habits (such as high-salt diet), pathological indicators (such as tumor size, depth of invasion, etc.), biomarkers (such as CEA (carcinoembryonic antigen), CA19-9, etc.). The target variable (i.e. dependent variable) is usually the patient's disease state or prognosis, including: disease status (such as whether the patient has gastric cancer (0 = no, 1 = yes)), survival time (time from diagnosis to death), disease progression (such as tumor progression (progression / no progression)), treatment response (response to treatment (good / medium / poor)).
[0141] Example Explanation: Suppose there is a dataset containing the above-mentioned key features and target variable. Data cleaning and preprocessing can be performed, such as handling missing values, standardization, etc., and then calculate the correlation coefficient between each feature and the target variable to evaluate their relationship. Here, take "whether suffering from stomach cancer" as the target variable, and use Pearson correlation coefficient as an example. First, a DataFrame containing virtual data is created, and the gender variable is numerically processed, that is, the data is preprocessed, etc. Then use the corr() method of Pandas to calculate the correlation coefficient between each feature and the target variable, and calculate the correlation coefficient matrix. Then use the pearsonr function to calculate the Pearson correlation coefficient between each feature and the target variable, that is, calculate the Pearson correlation coefficient separately. Finally, through the correlation coefficient matrix and the Pearson correlation coefficient value, the features with strong correlation with the target variable (whether suffering from stomach cancer) can be identified, so as to be considered in further model construction and feature selection, that is, result interpretation.
[0142] Step 322, feature extraction: perform principal component analysis (PCA), dimensionality reduction, reduce the number of features, and retain the main information. Extract key latent factors through factor analysis. Factor analysis is a multivariate statistical method used to identify and extract latent factors or structures in data. These latent factors are called "factors", which explain the correlation patterns between the original observed variables.
[0143] The results of factor analysis mainly include the following aspects:
[0144] (1) Factor Loadings Matrix: The factor loadings matrix shows the load of each observed variable on each factor. Factor loadings reflect the correlation or contribution between variables and factors. Load values are usually between -1 and 1, and the larger the value (positive or negative), the stronger the explanatory power of the variable on the factor. That is, it shows the load of each variable on the two factors. Higher load values indicate that the variable contributes more to the factor.
[0145] (2) Factor Scores: Factor scores are the scores of each observation on the extracted factors. These scores represent the position of each observation in different factor dimensions. Factor scores can be used for further analysis, such as regression analysis, clustering analysis, etc. That is, it shows the scores of each observation on the two factors. Scores can be used for further analysis and visualization.
[0146] (3) Eigen values and Explained Variance: Eigen values reflect the explanatory power of each factor. Larger eigen values indicate that the factor can explain more total variance, i.e., showing the explanatory power of each factor, larger eigen values indicate that the factor explains more variance. Cumulative explained variance percentage is used to evaluate how much the extracted factors can explain the variance of the original data, i.e., showing the proportion of the extracted factors that explain the variance of the original data. Cumulative explained variance is used to evaluate the effectiveness of the factors.
[0147] (4) Factor Rotation: To improve the interpretability of factors, factor loadings are usually rotated. Common rotation methods include orthogonal rotation (such as Varimax) and oblique rotation (such as Promax). The rotated factor loading matrix is easier to interpret, and each factor usually has more obvious high loading variables.
[0148] (5) Factor Correlation Matrix: In oblique rotation, factors may be correlated, and the factor correlation matrix shows the correlation between factors. Through these results, researchers can understand the underlying structure of the data, identify key latent factors, and use these factors in further analysis. For example, in a gastric cancer study, key latent factors that affect disease progression can be identified through factor analysis, and risk assessment and prognosis prediction can be based on these factors.
[0149] Step 33, model selection and training:
[0150] Step 331, in the model selection process, algorithm selection and benchmark model training are involved. Specifically, the algorithm selection process is: selecting algorithms suitable for data characteristics and research objectives, such as linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), neural network, XGBoost, etc. The benchmark model construction process is: constructing a simple benchmark model (such as linear regression, decision tree) for preliminary verification.
[0151] For example, a simple machine learning model can be used to preliminarily verify the prediction ability of the above-mentioned features on the target variable. In the case of gastric cancer, a logistic regression model can be used to predict whether a patient has gastric cancer (target variable). The following are specific steps and preliminary verification result examples: Assume that there is a dataset containing some key features and target variables. These data will be used to build and verify the corresponding baseline model. First, preprocess the data, including handling missing values, standardization, converting categorical variables to numerical variables, etc.; use a logistic regression model as the baseline model. Divide the data into training set and test set, train the baseline model and perform preliminary verification, calculate the performance indicators of the model, such as accuracy, precision, recall and F1 score. Based on the corresponding code, the preliminary verification results of the baseline model can be obtained; assume that the output accuracy is 0.70, precision is 0.67, recall is 0.60, and F1 score is 0.63; among them, the accuracy represents the proportion of samples correctly predicted by the model in the total samples, and the accuracy of 0.70 means that the model correctly classifies 70% of the samples. The precision indicates the proportion of actual positive samples in the samples predicted by the model as positive, and the precision of 0.67 means that among the samples predicted by the model as positive, 67% are actually positive. The recall indicates the proportion of actual positive samples correctly predicted as positive, and the recall of 0.60 means that among the actual positive samples, 60% are correctly identified by the model. The F1 score is the harmonic mean of precision and recall, and the F1 score of 0.63 means that the model has achieved a relative balance between precision and recall. These preliminary verification results show that the baseline model has certain ability in predicting whether a patient has gastric cancer, but there is still room for improvement. Therefore, more complex models and feature selection methods can be used to further optimize the performance of the model.
[0152] Step 332, in the model training process, cross-validation and hyperparameter optimization are involved. Specifically, k-fold cross-validation is used to evaluate model performance, select the best parameters, and perform cross-validation. Grid search or random search is used to optimize model hyperparameters. Among them, selecting the best parameters is an important step in machine learning model optimization, common methods include grid search, random search and Bayesian optimization.
[0153] The following is an example of how to use these methods to select the best parameters in the case of gastric cancer: assuming that the logistic regression algorithm is used for modeling, the common parameters include regularization strength and regularization type: first define the parameter grid that needs to be adjusted, define the logistic regression model (the specific definition of the program code is not shown); then initialize the data, perform grid search to find the best parameter combination (the specific definition of the program code is not shown); finally, use the best parameters found to retrain the model and validate it on the test set, and calculate the performance indicators. Assuming the best parameters are C=1 and penalty='l2', the performance of the model is as follows: Accuracy=0.80, Precision=0.78, Recall=0.75, F1 Score=0.76. These results show that by using grid search, the best parameters can be found, and the established logistic regression model performs better in predicting the status of gastric cancer.
[0154] Step 34, model evaluation and optimization: during evaluation, draw the confusion matrix, calculate the performance indicators and analyze them; during optimization, optimize the features based on feature selection and engineering, and optimize the model parameters based on model tuning.
[0155] Specifically, step 341, during model evaluation, specifically, select appropriate performance indicators according to the task, such as accuracy, precision, recall, F1 value, AUROC for classification problems, and mean square error (MSE), root mean square error (RMSE), R 2 For classification problems, draw a confusion matrix to analyze the types of classification errors made by the model.
[0156] Step 342, during model optimization, involves the process of feature selection and engineering, model tuning, etc., where feature selection and engineering are used to further optimize features, such as combining features and interacting features. Model tuning is based on the validation set results to further optimize model parameters.
[0157] Step 35, model interpretation: for example, use SHAP values to explain the contribution of each feature to the prediction result. Use the LIME method to explain the local model prediction result. Analyze the feature importance ranking in the model to identify key features. For example, use SHAP values to explain the contribution of each feature to the prediction result: SHAP values are based on Shapley values in game theory, used to explain the contribution of each feature to the model prediction result. SHAP value summary shows the contribution of each feature to the model prediction result, including the high and low of feature value and the size of contribution. Assuming the output result shows the importance of the following features: Tumor_Size, CEA, Age, Smoking_History, Gender. The X-axis represents the SHAP value, indicating the influence of each feature on the model output. The Y-axis lists each feature, and the distribution of points represents the range of contribution of that feature to the prediction. Through this summary, it can be intuitively seen which features have the greatest impact on the model prediction, and how each feature affects the prediction result. A wider point distribution indicates that the contribution of that feature to the prediction result has a larger range of variation. Use the LIME method to explain the local model prediction result: the LIME method provides explanations for specific samples, making it easier to understand the model's decision-making on individual samples. A specific sample can be selected to explain its prediction result. Assuming the LIME explanation result shows the feature contribution of sample 1 as follows: Tumor_Size:+0.4, CEA:+0.3, Age:+0.2, Smoking_History:+0.1, Gender:-0.1. Feature importance ranking: based on the absolute value of the model coefficient, showing the importance of each feature, with the top-ranked features having a greater impact on the prediction result and being key features. Assuming the feature importance ranking (Feature Importance Ranking) is as follows: Tumor_Size:2.4567, CEA:1.9876, Age:1.5678, Smoking_History:1.2345, Gender:0.9876.
[0158] Step 6, obtain comprehensive results of analysis based on the analysis results of each angle in steps 1 to 5: in this case, through the comprehensive analysis of each angle above, the overall effectiveness and safety evaluation in the real world can be obtained.
[0159] As can be seen from the above case, through this comprehensive analysis process, which involves feature engineering and various analyses, the quality and efficiency of real-world medical research can be greatly improved, providing users with more accurate and personalized treatment plans.
[0160] Specifically, the idea of feature engineering is to introduce automated feature selection and optimization algorithms, such as model-based feature selection and genetic algorithms, to identify and extract useful features from raw data for subsequent analysis and model building. Deep learning techniques are used for feature construction to mine deep-level features and potential interaction effects from raw data, further enhancing the predictive power of the model.
[0161] The idea of statistical analysis is to use descriptive statistical analysis, hypothesis testing, regression analysis, and other methods to explore the relationship between data and evaluate the differences in treatment effects. Specifically, based on the explored data relationship, statistical analysis methods can be used to evaluate the differences in treatment effects. Using descriptive statistical analysis, the distribution, central tendency, and variability of the data can be summarized and described. By comparing the descriptive statistics (such as median) of different treatment groups or different time points, the differences in treatment effects can be preliminarily evaluated.
[0162] For example, comparing the mean or median of two treatment groups, if the mean of one group is significantly higher than the other, it may suggest that the treatment group has better treatment effects. Hypothesis testing can be used to verify whether the difference in treatment effects is statistically significant. Common hypothesis tests include t-tests, etc., for example, t-tests can be used to compare whether the means of two groups are significantly different, if the p-value of the t-test is less than the pre-set significance level (usually 0.05), the null hypothesis can be rejected, indicating that there is a significant difference between the two groups.
[0163] Regression analysis can be used to explore the relationship between independent variables and dependent variables, and further evaluate the differences in treatment effects, for example, when comparing the effects of two treatment options, a regression model can be established, with treatment options as independent variables and treatment effects as dependent variables, then the impact of independent variables on dependent variables is evaluated. Effect size in effect size analysis is an index used to describe the size of the difference between two groups, which is not affected by sample size. Common effect sizes include Cohen's d, Pearson correlation coefficient, etc., for example, by calculating Cohen's d, the effect size between two groups can be evaluated, so that the difference in treatment effects can be more intuitively understood. By analyzing the relationship between data and statistical significance, it can be determined whether there is a significant difference between different treatment groups, and further evaluate the size and clinical significance of the treatment effect.
[0164] The analysis approach of machine learning is to combine multiple machine learning modeling and prediction algorithms, including but not limited to random forest, gradient boosting machine and deep neural network, to identify patterns and trends in the data, predict treatment effects, and explore potential risks or benefits. Implement model training strategies such as adaptive learning rate adjustment and model ensemble techniques to improve the generalization ability and performance of the model. Among them, identifying patterns and trends in data is to understand the basic structure and regularity of data, predicting treatment effects is to predict the treatment response of patients according to data, and exploring potential risks or benefits is to evaluate the risks and benefits that different treatment options may produce. These tasks are closely related and together provide support and guidance for medical decision-making.
[0165] Among them, identifying patterns and trends in data usually refers to discovering the internal structure and regularity of data through machine learning; for example: using algorithms such as random forest, gradient boosting machine and deep neural network to learn from a large amount of data and discover patterns and relationships therein. Once the patterns and trends in the data are identified, the information can be used to predict treatment effects; for example, based on the clinical characteristics and medication of patients, machine learning algorithms are used to model and predict patient treatment responses, such as drug efficacy or side effects. Exploring potential risks or benefits is to analyze patient data using established models and evaluate the risks and benefits that different treatment options may produce; for example, through machine learning, the risk (risk score or likelihood) of side effects or complications caused by a certain treatment option can be evaluated, as well as the benefits that treatment may bring.
[0166] Model interpretation is to apply model interpretation and visualization tools such as SHAP values and LIME to deeply understand the predictive behavior of the model and the contribution of each feature to the model prediction, thereby improving transparency and interpretability. Based on the comprehensive analysis of the model, the potential value, benefits and risks of medical products in real-world applications, as well as the possible reaction differences of different user groups are proposed. Among them, SHAP (Shapley Additive exPlanations) value and LIME (Local Interpretable Model-agnostic Explanations) are two commonly used model interpretation and visualization tools.
[0167] SHAP value is a game theory-based method for explaining individual prediction model results. It analyzes the contribution of each feature to the prediction result and provides an importance score for each feature, thereby explaining the predictive behavior of the model.
[0168] LIME is a local explanation method for individual prediction instances, which generates some local approximation models around the model and explains the prediction results of these approximation models to explain the prediction behavior of the model. The correspondence between prediction behavior, features, contribution, transparency, and explainability: prediction behavior refers to the prediction or judgment made by the model under given input conditions. Explanation and visualization tools represented by SHAP and LIME can help understand the behavior of the model in the preset prediction, that is, explain why the model gives such a prediction result. Features are input variables for model prediction. Contribution represents the relative importance or impact of each feature on the model prediction result. Transparency and explainability refer to whether the internal working mechanism of the model can be understood and explained. By using tools such as SHAP and LIME, the contribution of each feature to the model prediction result can be analyzed, that is, the influence of each feature on the prediction behavior of the model can be understood. SHAP value and LIME can give the contribution score of each feature, helping to understand the dependence of the model on the prediction, thereby improving the transparency and explainability of the model, so that complex models can also be explained and understood, thereby enhancing the trust and application reliability of the model.
[0169] Based on the operation method of the software robot system for real world research, a software robot system for real world research can also be implemented, which includes a data input module, a data processing module and a data analysis module; and the data input module transmits the processed data to the data processing module for processing, and the processed data is transmitted to the data analysis module for analysis; the data input module is used for setting data flow, data extraction and data transmission operation; the data processing module is used for data preprocessing, data conversion and data quality control operation; the data analysis module is used for comprehensive analysis. The results after the experiment can be written into a detailed analysis report, including research methods, analysis process, result interpretation and its clinical significance. The results can also be output in the form of API or board, showing the real world research evidence of medical products.
[0170] Finally, it should be noted that the above content is only used to illustrate the technical solutions of the present application, and is not a limitation on the protection scope of the present application. Simple modifications or equivalent replacements of the technical solutions of the present application made by those skilled in the art do not deviate from the essence and scope of the present application.
Claims
1. A method of operation of a software robot system for real world research, characterized by: Comprising the following steps: Step 1, set up data flow, and perform data extraction operation: according to data type, configure data flow, and connect original data source; extract key information from original data source; and use the extracted information for data preprocessing; data types include structured data, semi-structured data and unstructured data; Data flow defines the path and flow mode of data from the original data source to the target system; The original data source includes database, file system, network service; Step 2, pre-process the data, and convert the pre-processed data, and perform data quality control operation on the converted data; The pre-processing process is: the initially received raw data is pre-processed, irrelevant data and duplicate records are removed in an automated manner, and standardized processing is performed; standardized processing refers to converting data into consistent format or arrangement form; The data conversion process is: the cleaned data is converted into a unified or preset format through the built-in mapping tool; The data quality control operation includes outlier detection operation based on machine learning and automatic data correction operation; The data correction process includes the following steps: Step 21, data cleaning: clean the data identified as outliers, missing or inconsistent; Step 22, data verification: after cleaning, re-verify the data to ensure the effectiveness of the modification measures and the consistency of the data; If the data verification result shows that the data quality cannot meet the requirements, adjust it to meet the requirements, and the adjustment methods include: redefining the data source, resetting the data flow or manual intervention; Step 23, data record: record the pre-update and post-update versions of the data to ensure the traceability of the data correction process; Step 3, perform comprehensive analysis on the data processed by the quality control operation, and run according to the analysis result to obtain the comprehensive analysis result.
2. The method of Claim 1, wherein: In step 1, the data extraction process is: the software robot system extracts key information from the original data source by using natural language processing and image recognition technology, wherein the key information at least includes user information, chief complaint, present illness history, past history, laboratory test result and imaging examination result.
3. The method of Claim 1, wherein: The outlier detection in the data quality control operation is to analyze the data using preset rules and algorithms to identify outliers, missing values and inconsistent values in the data set that are not logical or do not match the known patterns; wherein, The detection process of outliers is: using statistical methods or machine learning methods to identify data points that deviate from the preset range; The detection and processing process of missing values is: analyzing the pattern of data missing, and processing the missing values by using corresponding strategies; The detection process of inconsistent values is: checking the logical errors and inconsistencies in the data by setting rules.
4. The method of Claim 3, wherein: The preset rules include: using statistical methods to identify outliers, and determining the box line position of data points and calculating the absolute value of data point Z-Score, Rule 1: if a data point is lower than Q1-1.5IQR or higher than Q3+1.5IQR, it is an outlier, wherein IQR=Q3-Q1, Q3 is the upper quartile, Q1 is the lower quartile, and IQR is the interquartile distance.
5. The method of Claim 4, wherein: The preset rules include: identifying outliers by using statistical methods, determining the box line position of the data points, and calculating the absolute value of the Z-Score of the data points, Rule two: if the absolute value of the Z-Score of a data point is greater than 2, it is an outlier, wherein the Z-Score is a measurement unit; If the result of rule one or rule two is an outlier, the data point is determined to be an outlier.
6. The method of Claim 3, wherein: In step 21, the data identified as outliers, missing or inconsistent is cleaned up; specifically including: When the number of outliers is less than 10% of the total data, the mean or median of the entire data set is used to replace the outliers; when the number of outliers is greater than or equal to 10% of the total data, the current feature column is removed, and the remaining features are clustered by KNN, and the average value of K neighbor samples is used to replace the outliers.
7. The method of Claim 1, wherein: The comprehensive analysis includes a statistical analysis process, which includes the following steps: First, descriptive statistical analysis is performed, and the user group is grouped: according to the key information, the users are divided into different subgroups; Second, use propensity score matching method to control potential confounding factors and ensure balance between subgroups in key information; Third, survival analysis is performed: survival analysis is performed using the Kaplan-Meier survival curve and Cox proportional hazards regression model method to evaluate the survival or event rate of different subgroups; Fourth, safety evaluation combined with integrated data: analyze the safety data collected in the study to evaluate the safety of different treatments or interventions; Fifth, sensitivity analysis by changing matching criteria and / or adjusting model parameters to verify the robustness and reliability of the results; Sixth, based on the analysis results of each angle from the first to the fifth steps, the comprehensive results of the analysis are obtained.
8. The method of Claim 7, wherein: In the fifth step, the hazard ratio is used to represent the robustness and reliability, which specifically includes: using multiple imputation method to handle missing values, re-performing propensity score matching and survival analysis, and comparing the hazard ratios before and after re-analysis using multiple imputation method. If the difference between the hazard ratios before and after is greater than a threshold value, the analysis method and data processing flow are re-considered.
9. The method of Claim 7, wherein: In step 3, the comprehensive analysis also includes machine learning analysis, and the machine learning analysis process includes the following steps: Step 31, data acquisition, classification and preprocessing; wherein the data preprocessing process includes data cleaning, data standardization and normalization, and data set division; Step 32, feature engineering processing: including feature selection processing and feature extraction processing; Step 33, model selection and model training, wherein k-fold cross-validation is used to evaluate model performance and perform cross-validation; Step 34, model evaluation and optimization: when evaluating, draw a confusion matrix, calculate performance indicators and analyze; when optimizing, perform feature optimization based on feature selection and engineering, and perform model parameter optimization based on model tuning; Step 35, model interpretation according to the evaluation and optimization results.
10. A software robot system for real world research, characterized by: The method for operating a software robot system for real-world research according to any one of claims 1-9, the software robot system comprising a data input module, a data processing module, and a data analysis module; and the data input module transmits the extracted data to the data processing module for processing, and the processed data is transmitted to the data analysis module for analysis; The data input module is used for setting data flow, data extraction, and data transmission operations; The data processing module is used for data preprocessing, data conversion, and data quality control operations; The data analysis module is used for comprehensive analysis.
Citation Information
Patent Citations
Software robot system for real world research and operation method
CN118412142A