A data processing system based on multi-source data
By integrating a multi-source data processing system that includes a user interface, feature filtering, machine learning models, and model training modules, this system addresses the high barriers to entry, operational complexity, and fragmented processes of existing medical data analysis tools, enabling efficient and accurate data analysis without the need for programming.
Patent Information
- Application Number
- CN202411836697.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2044-12-13
AI Technical Summary
Existing medical data analysis tools are complex to operate, have high barriers to entry, lack intelligent support, and have fragmented data processing and analysis workflows, making it difficult to meet the needs of medical researchers without a programming background.
This paper provides a data processing system based on multi-source data, which integrates a user interface module, a feature selection module, a machine learning model module, and a model training module. It supports automated processing and intelligent analysis of multi-dimensional data, including data preprocessing, feature selection, model training, and result display.
It reduces medical researchers' reliance on programming, improves the efficiency and accuracy of data analysis, provides personalized processing capabilities, supports one-stop intelligent analysis, and avoids complex operations and human errors.
Smart Images

Figure CN119862403B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of data analysis and artificial intelligence technology, and more specifically, to a data processing system based on multi-source data. Background Technology
[0002] With the rapid development of big data technology and artificial intelligence (AI), data analysis has been widely applied in various industries. In the medical field, with the proliferation of electronic medical records, genomic data, imaging data, and other health monitoring data, researchers face enormous challenges in data processing and analysis. While traditional statistical methods and data analysis tools have some effectiveness, they often struggle to adapt to the dimensionality, complexity, and diversity of predictive tasks.
[0003] Especially in the medical field, clinical data is often multidimensional, heterogeneous, and frequently contains missing values and noise. To effectively extract useful knowledge from this complex data, researchers need to leverage machine learning and artificial intelligence technologies for data modeling, feature extraction, risk prediction, and personalized intervention recommendations. Therefore, building an intelligent analysis platform that integrates data preprocessing, feature selection, machine learning modeling, and predictive evaluation has become a key technological requirement for advancing medical and clinical research.
[0004] With the increasing volume and complexity of medical research data, data analysis has become a fundamental and crucial task in medical research. Medical research data is typically multi-dimensional, highly complex, and heterogeneous, involving various types of data such as medical records, genetic data, and imaging data. The analysis and modeling of this data requires a high level of technical expertise and professional background. However, current medical data analysis tools generally suffer from the following major drawbacks:
[0005] High technical barriers: Most existing analytical tools require users to have a certain level of programming ability, especially the application of machine learning algorithms. Users typically need to understand the principles of the algorithms, parameter tuning, data preprocessing, and other technical details. Medical researchers often lack sufficient programming background, leading to difficulties when using these tools and hindering their ability to fully utilize machine learning and artificial intelligence technologies for data analysis.
[0006] Operational complexity: While some analytical tools possess certain data processing and modeling capabilities, their user interfaces are typically complex, with cumbersome workflows involving extensive manual configuration and parameter settings. For users without a programming background, this high level of complexity significantly increases the barrier to entry and impacts efficiency.
[0007] Lack of intelligent support: Current tools often lack intelligent functions, requiring users to manually select models, adjust parameters, and evaluate results, often without automated optimization or recommendation features. This makes it easy for medical researchers to produce inaccurate or unstable results when faced with large amounts of data due to inappropriate selection or incomplete analysis.
[0008] Fragmented data processing and analysis workflows: Existing tools typically separate data preprocessing, feature selection, model training, and prediction evaluation processes, requiring researchers to perform each step individually. This makes the entire analysis process cumbersome and prone to errors. This fragmented work model increases the workload, especially for users lacking a data science background. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide a data processing system based on multi-source data, which aims to solve at least one of the above-mentioned technical problems.
[0010] In a first aspect, the technical solution of the present invention to solve the above-mentioned technical problems is as follows: a data processing system based on multi-source data, the system comprising:
[0011] The module includes a user interface module, a feature selection module, a machine learning model module, and a model training module.
[0012] The user interface module is used to acquire multidimensional data uploaded by the user, acquire the target analysis task of the multidimensional data selected by the user, and acquire the target machine learning algorithm model selected by the user.
[0013] The feature filtering module is used to select relevant features of the target variable from the multidimensional data based on the target variable;
[0014] The machine learning model module is used to provide the user with machine learning algorithm models corresponding to different analysis tasks, the different analysis tasks including the target analysis task, and the machine learning algorithm models corresponding to the different analysis tasks including the target machine learning algorithm model.
[0015] The model training module is used to train the target machine learning algorithm model based on the relevant features of the target variable to obtain a target model, and then perform the processing corresponding to the target analysis task based on the relevant features of the target variable using the target model.
[0016] The beneficial effects of this invention are as follows: The system of this invention enables users to efficiently process and analyze multidimensional data, automatically perform machine learning modeling, and supports personalized processing. This system not only reduces the reliance of medical researchers on programming and improves the efficiency and accuracy of data analysis, but also provides more precise references for clinical decision-making. Furthermore, this system integrates multiple analysis modules, including a user interface module, a feature selection module, a machine learning model module, and a model training module. Through one-stop intelligent analysis, it can automate the entire data processing workflow, avoiding user switching and complex operations between multiple steps, making the entire analysis process more efficient and seamless.
[0017] Based on the above technical solution, the present invention can be further improved as follows.
[0018] Furthermore, the system also includes a data preprocessing module, which is used to preprocess the multidimensional data. The preprocessing includes at least one of missing value imputation, data standardization, outlier detection, and data normalization.
[0019] Furthermore, the user interface module is specifically used for:
[0020] Get the first trigger operation of the user on the data upload identifier displayed on the user operation interface, and get the multidimensional data uploaded by the user based on the first trigger operation;
[0021] Get the user's first selection operation on the target analysis task identifier among the various analysis task identifiers displayed on the user operation interface, and based on the first selection operation, get the target analysis task corresponding to the target analysis task identifier selected by the user;
[0022] Obtain the user's second selection operation on the target machine learning identifier among the various machine learning identifiers displayed on the user operation interface, and based on the second selection operation, obtain the target machine learning algorithm model corresponding to the target machine learning identifier selected by the user.
[0023] Furthermore, the system also includes a result display and analysis module, which is used to visualize the processing results output by the model training module.
[0024] Furthermore, the visualization display types include data distribution maps, model prediction charts, and feature importance bar charts.
[0025] Furthermore, the system also includes a report generation module, which is used to generate a report based on at least one of the output results of the data preprocessing module, the feature filtering module, the user interface module, and the model training module.
[0026] Furthermore, each analysis task corresponds to a processing result, and the user interface module is also used for:
[0027] Obtain the second trigger operation of the user in response to the target processing result identifier among the various processing result identifiers displayed on the user operation interface;
[0028] Based on the second triggering operation, the target processing result corresponding to the target processing result identifier is obtained and displayed.
[0029] Furthermore, the user interface module is also used for:
[0030] The user's settings for the target parameters among the various analysis parameters displayed on the user interface are obtained.
[0031] Based on the settings, the set target parameters are obtained and displayed.
[0032] Furthermore, if the number of each analysis parameter is greater than a set number, all analysis parameters are hidden under the menu corresponding to the parameter setting identifier displayed on the user interface. When the user interface module obtains the user's setting operation for the target parameter among the analysis parameters displayed on the user interface, and based on the setting operation, obtains and displays the set target parameter, it is specifically used for:
[0033] Obtain the user's third selection operation for the parameter setting identifier, and display the menu corresponding to each of the analysis parameters based on the third selection operation;
[0034] Obtain the user's settings for the target parameters among the various analysis parameters; based on the settings, obtain and display the set target parameters.
[0035] Furthermore, the feature selection module is specifically used to: select relevant features of the target variable from the multidimensional data according to the target variable through a target algorithm, wherein the target algorithm is any one of the correlation analysis algorithm, LASSO regression algorithm, and recursive feature elimination algorithm.
[0036] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below.
[0038] Figure 1 This is a schematic diagram of the structure of a data processing system based on multi-source data, provided as an embodiment of the present invention. Detailed Implementation
[0039] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0040] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0041] This invention provides a possible implementation, such as... Figure 1 As shown, a schematic diagram of a data processing system based on multi-source data is provided, such as... Figure 1 As shown, the system may include:
[0042] The module includes a user interface module, a feature selection module, a machine learning model module, and a model training module.
[0043] The user interface module is used to acquire multidimensional data uploaded by the user, acquire the target analysis task of the multidimensional data selected by the user, and acquire the target machine learning algorithm model selected by the user.
[0044] The feature filtering module is used to select relevant features of the target variable from the multidimensional data based on the target variable;
[0045] The machine learning model module is used to provide the user with machine learning algorithm models corresponding to different analysis tasks, the different analysis tasks including the target analysis task, and the machine learning algorithm models corresponding to the different analysis tasks including the target machine learning algorithm model.
[0046] The model training module is used to train the target machine learning algorithm model based on the relevant features of the target variable to obtain a target model, and then perform the processing corresponding to the target analysis task based on the relevant features of the target variable using the target model.
[0047] The system of this invention enables users to efficiently process and analyze multidimensional data, automatically perform machine learning modeling, and supports personalized processing. This system not only reduces the reliance of medical researchers on programming and improves the efficiency and accuracy of data analysis, but also provides more precise references for clinical decision-making. Furthermore, this system integrates multiple analysis modules, including a user interface module, feature selection module, machine learning model module, and model training module. Through one-stop intelligent analysis, it automates the entire data processing workflow, avoiding user switching and complex operations between multiple steps, making the entire analysis process more efficient and seamless.
[0048] The following specific embodiments further illustrate the solution of the present invention. The present invention provides a data processing system based on multi-source data, aiming to overcome the aforementioned shortcomings of the prior art, particularly improving the efficiency and accuracy of data analysis for medical researchers without programming backgrounds. Specifically, the present invention addresses the deficiencies of the prior art through the following technical means:
[0049] 1) Simplified User Interface and Intelligent Functions: This invention provides a user-friendly interface through its user interface module, enabling researchers to quickly perform data analysis and machine learning predictions without programming. The system design lowers the operational threshold; users can complete tasks such as data uploading, preprocessing, model training, and evaluation simply through clicks and selections, greatly improving the efficiency of data analysis.
[0050] 2) One-stop data analysis and modeling process: This invention integrates multiple analysis modules such as data preprocessing, feature selection, machine learning modeling, and model evaluation. Through one-stop intelligent analysis function, the system can automatically complete the entire data analysis process, avoiding users switching between multiple steps and complex operations, making the entire analysis process more efficient and coherent.
[0051] 3) Intelligent Recommendation and Automated Optimization: The system's machine learning model module embeds multiple machine learning algorithms and features intelligent model selection and automatic optimization. Through automated model selection and hyperparameter tuning, the system can intelligently recommend suitable analysis models based on different types of tasks (such as classification or regression) and automatically optimize hyperparameters. This avoids the tedious work of manually selecting algorithms and adjusting parameters, reduces human error in the analysis process, and improves the accuracy of the analysis results.
[0052] 4) Machine learning prediction without programming: The system of this invention provides a code-free operation method, enabling researchers without programming experience to easily apply machine learning for data analysis. The system automatically completes tasks such as data preprocessing, feature selection, model training, and prediction, allowing researchers to focus on interpreting the analysis results and making decisions without needing to delve into the details of programming and machine learning algorithms.
[0053] 5) Automatic Generation of Results and Analysis Reports: Through the report generation module, this invention can automatically generate analysis reports. The reports include data processing, feature selection, model evaluation, and prediction results, helping researchers clearly understand the data analysis process and results. This automated report generation not only saves a significant amount of time but also avoids potential biases that may arise from manual interpretation of analysis results.
[0054] In summary, this invention addresses the problems of high technical barriers, operational complexity, insufficient intelligent support, and fragmented data analysis workflows found in existing medical data analysis tools through a series of technical means. This system, by simplifying operations, providing intelligent recommendations, automatically optimizing, and eliminating the need for coding, enables medical researchers with no programming experience to efficiently perform data preprocessing, modeling, prediction, and result evaluation, significantly improving the efficiency and accuracy of data analysis and promoting the development of medical research towards a more intelligent and efficient direction.
[0055] Based on the technical means provided by this system, this embodiment provides a detailed description of a data processing system based on multi-source data provided in this application. The system includes:
[0056] The module includes a user interface module, a feature selection module, a machine learning model module, and a model training module.
[0057] The user interface module is used to acquire multidimensional data uploaded by the user, acquire the target analysis task of the multidimensional data selected by the user, and acquire the target machine learning algorithm model selected by the user.
[0058] The multidimensional data includes medical records, genetic data, and imaging data. The target analysis task refers to the desired type of result after processing the multidimensional data, i.e., the type of processing result corresponding to the multidimensional processing. The target machine learning model refers to which machine learning algorithm model is used to process the multidimensional data to obtain the processing result corresponding to the target analysis task.
[0059] As an example, a target analysis task might be to extract data that characterizes the features of a target from multidimensional data.
[0060] The feature filtering module is used to select relevant features of the target variable from the multidimensional data based on the target variable;
[0061] The machine learning model module is used to provide the user with machine learning algorithm models corresponding to different analysis tasks, the different analysis tasks including the target analysis task, and the machine learning algorithm models corresponding to the different analysis tasks including the target machine learning algorithm model.
[0062] The machine learning algorithm models corresponding to different analysis tasks include regression algorithm models (such as linear regression, ridge regression, Lasso regression, etc.), classification algorithm models (such as decision trees, support vector machines, random forests, K-nearest neighbors, etc.), and ensemble learning method models (such as XGBoost, AdaBoost, etc.). These are used to implement different analysis tasks.
[0063] The model training module is used to train the target machine learning algorithm model based on the relevant features of the target variable to obtain a target model, and then to process the relevant features of the target variable based on the target model to obtain the processing result.
[0064] Optionally, the main application areas of this system include medical research, clinical data analysis, precision medicine, disease prediction, and personalized treatment.
[0065] Optionally, the feature selection module is specifically used to: select relevant features of the target variable from the multidimensional data according to the target variable using a target algorithm, wherein the target algorithm is any one of a correlation analysis algorithm, a LASSO regression algorithm, and a recursive feature elimination algorithm.
[0066] The target variable represents the parameter corresponding to the processing result of the target analysis task. For example, if the processing result of the target analysis task identifies the type of parameter A in multidimensional data, then the corresponding target variable is parameter A.
[0067] Based on the characteristics of multidimensional data, this module automatically selects features related to the target variable and uses statistical or machine learning methods (such as correlation analysis, LASSO regression, recursive feature elimination, etc.) to screen features, helping researchers select the most representative features and improve the model's predictive performance.
[0068] Optionally, the user interface module is specifically used for:
[0069] Get the first trigger operation of the user on the data upload identifier (e.g., a button) displayed on the user operation interface, and get the multidimensional data uploaded by the user based on the first trigger operation;
[0070] Get the user's first selection operation on the target analysis task identifier among the various analysis task identifiers displayed on the user operation interface, and based on the first selection operation, get the target analysis task corresponding to the target analysis task identifier selected by the user;
[0071] Obtain the user's second selection operation on the target machine learning identifier among the various machine learning identifiers displayed on the user operation interface, and based on the second selection operation, obtain the target machine learning algorithm model corresponding to the target machine learning identifier selected by the user.
[0072] Specifically, a target machine learning algorithm model can be selected based on the target analysis task and the relevant characteristics of the target variable.
[0073] This user interface simplifies the complex analysis process, providing users with operable function buttons and options, greatly reducing the barrier to entry.
[0074] Optionally, the system further includes a data preprocessing module, which is used to preprocess the multidimensional data, the preprocessing including at least one of missing value imputation, data standardization, outlier detection and data normalization.
[0075] This module supports users in preprocessing input multidimensional data, including data cleaning processes such as missing value imputation, data standardization, and outlier detection, to ensure that the data meets the input requirements of machine learning models.
[0076] Optionally, each analysis task corresponds to a processing result, and the user interface module is further used for:
[0077] Obtain the second trigger operation of the user in response to the target processing result identifier among the various processing result identifiers displayed on the user operation interface;
[0078] Based on the second triggering operation, the target processing result corresponding to the target processing result identifier is obtained and displayed.
[0079] Optionally, the user interface module is further used for:
[0080] The user's settings for the target parameters among the various analysis parameters displayed on the user interface are obtained.
[0081] Based on the settings, the set target parameters are obtained and displayed.
[0082] Here, the analysis parameters refer to the parameters that need to be analyzed. The analysis parameters can be used as target variables. In this application, users can add new analysis parameters and modify the added analysis parameters to provide personalized services to users.
[0083] Furthermore, if the number of each analysis parameter is greater than a set number, meaning that all analysis parameters cannot be displayed at once on the user interface, then all analysis parameters are hidden under the menu corresponding to the parameter setting identifier displayed on the user interface. When the user interface module obtains the user's setting operation for the target parameter among the various analysis parameters displayed on the user interface; and based on the setting operation, obtains and displays the set target parameter, it is specifically used for:
[0084] Obtain the user's third selection operation for the parameter setting identifier, and display the menu corresponding to each of the analysis parameters based on the third selection operation;
[0085] Obtain the user's settings for the target parameters among the various analysis parameters; based on the settings, obtain and display the set target parameters.
[0086] Optionally, the system further includes a result display and analysis module, which is used to visualize the processing results output by the model training module.
[0087] Furthermore, the visualization display types include data distribution maps, model prediction charts, and feature importance bar charts.
[0088] The system can visualize the processing results, presenting them in the form of charts or reports for easy understanding and use by researchers. This module supports multi-dimensional result display, including data distribution maps, model prediction charts, feature importance bar charts, etc., helping users intuitively understand the analysis process and results.
[0089] Optionally, the system further includes a report generation module, which is used to generate a report based on at least one of the output results of the data preprocessing module, the feature filtering module, the user interface module, and the model training module.
[0090] Based on the data analysis and model evaluation results, the system automatically generates a concise and clear analysis report. The report includes data cleaning steps, feature selection results, model selection, training process, evaluation indicators, etc., which facilitates decision-making by researchers.
[0091] Optionally, in the model training module, during the model training process, it is also used to evaluate the model's performance through cross-validation and / or model evaluation (such as accuracy, precision, recall, F1 score, etc.) and provide optimization suggestions to the user.
[0092] Optionally, since the data used in this application is multidimensional data, and the data characteristics of different dimensions are different, the feature filtering module, when selecting relevant features of the target variable from the multidimensional data based on the target variable, is specifically used for:
[0093] Based on the target variable, determine the relevant features corresponding to the target variable from the medical record data;
[0094] Based on the target variable, determine the relevant features corresponding to the target variable from the gene data;
[0095] Based on the target variable, the target image corresponding to the target variable is determined from the image data, and the target image is used as the relevant feature corresponding to the target variable.
[0096] Furthermore, based on the target variable, relevant features corresponding to the target variable can be determined from medical record data and / or genetic data using natural language processing techniques. Specifically, semantically related keywords to the target variable are identified from the medical record data and / or genetic data, and these semantically related keywords are determined as relevant features corresponding to the target variable.
[0097] Furthermore, based on the target variable, the target image corresponding to the target variable can be determined from the image data using image processing techniques. Specifically, the target variable can be first converted into a target feature (e.g., a feature vector), and then a target region similar to the target feature can be determined from the image data based on the target feature. The target region is then extracted from the image data as the relevant feature corresponding to the target variable.
[0098] Furthermore, if the data types (text and images) of the relevant features determined from data of different dimensions are different, then the relevant features of different data types can be unified into features of the same data type, for example, all of them can be unified into feature vectors.
[0099] Furthermore, since the multidimensional data includes medical record data, genetic data, and image data, although relevant features corresponding to the target variable can be determined from data of different dimensions, the relevant features determined from data of different dimensions are not necessarily related. Based on this, in the scheme of this application, after determining the relevant features corresponding to data of different dimensions, similarity calculation is performed on each relevant feature, and some relevant features with low similarity (the similarity calculation result is less than the set similarity threshold) are excluded based on the similarity calculation result.
[0100] Furthermore, although the proposed solution can perform subsequent processing based on the relevant features corresponding to each of the multidimensional data, there may still be some missing information in the original multidimensional data. In such cases, feature expansion processing can be performed based on the relevant features corresponding to each of the multidimensional data. For example, features similar to the relevant features or target variables can be obtained from other data and added to the relevant features corresponding to each of the determined multidimensional data.
[0101] Furthermore, image data can also include different types of images, such as X-ray images, CT images, and MRI images.
[0102] For the region corresponding to the target variable, the corresponding region in different types of images is first aligned. This involves using the center point of the region as the base point and aligning the base points in different types of images. Then, using one image from each type of image as the base image, the other images in each type are rotated according to the shape of the region in the base image, ensuring that the shape of the region in the rotated image follows the same direction as the shape of the region in the base image. Finally, features can be extracted from the rotated image and the base image based on the target variable to obtain the relevant features for each image.
[0103] Furthermore, the relevant features corresponding to all images can be fused, and the fused features can be used as the relevant features corresponding to the image data.
[0104] Compared with the prior art, the solution of the present invention has the following advantages:
[0105] (1) Simplified user interface, lowering the technical threshold
[0106] Existing technical problems: Current machine learning analysis tools typically require users to have some programming knowledge or a deep understanding of machine learning algorithms. Many researchers (especially medical researchers) do not have a data science background, and they often have to rely on statistical software and manually adjust parameters for analysis, which makes data analysis cumbersome and error-prone.
[0107] The advantages of this invention: This invention presents complex data analysis and machine learning processes through a graphical interface by designing a simple and intuitive user interface module. Users only need to select buttons and drop-down menus to complete operations such as data uploading, task selection, and parameter setting, without writing any code. This design enables researchers without programming experience to easily perform efficient data analysis.
[0108] Performance Analysis: By eliminating programming requirements, the barrier to entry is significantly lowered, greatly improving user convenience. This is especially beneficial for researchers without technical backgrounds in fields like medicine, allowing them to focus on the business aspects of data analysis without needing to understand complex algorithm implementations. This not only accelerates the data analysis process but also avoids resource waste caused by technical barriers.
[0109] (2) Automated data preprocessing reduces human intervention
[0110] Existing technical problems: In traditional machine learning analysis tools, data preprocessing usually needs to be performed manually by the user, and different datasets require different preprocessing methods. This requires users to have high levels of expertise in steps such as data cleaning, missing value handling, and outlier detection when performing data analysis. This not only increases the difficulty of use but also makes it easy for improper operation to affect the model's performance.
[0111] Advantages of this invention: The data preprocessing module of this invention, through embedded intelligent algorithms, can automatically perform operations such as missing value imputation, outlier detection, standardization, and normalization of data. The system will automatically select an appropriate preprocessing method based on the type of input data and task requirements.
[0112] Results Analysis: The automation of this module eliminates the need for manual intervention in the data cleaning process, reducing human error and improving data quality. In large volumes of medical data, data preprocessing is a crucial step affecting analytical accuracy. This invention ensures the consistency and standardization of this process, thereby providing more reliable data support for subsequent machine learning model training.
[0113] (3) Intelligent feature selection to optimize model performance
[0114] Existing technical problems: In traditional data analysis, feature selection is usually done by users based on experience, involving a large amount of trial and error. The selection of different features directly affects the predictive ability of the model; incorrect feature selection may lead to poor model performance, or even overfitting or underfitting.
[0115] Advantages of this invention: The feature selection module of this invention automatically evaluates and selects the features most relevant to the target variable through intelligent algorithms (such as correlation analysis, LASSO regression, recursive feature elimination, etc.), helping users to select the most representative feature set and maximize the predictive ability of the model.
[0116] Performance Analysis: Through automated feature selection, this invention avoids feature selection errors caused by users' lack of data processing experience or human interference. This intelligent screening effectively reduces the impact of redundant features and noisy data, improving the accuracy and stability of the model. Especially when processing complex medical data, optimized feature selection can significantly improve the model's predictive performance, providing a more reliable basis for scientific research.
[0117] (4) Integrate multiple machine learning algorithms and flexibly select models.
[0118] Existing technical problems: Current data analysis tools often have fixed algorithmic frameworks or offer a limited variety of algorithms. If users encounter complex analysis problems, they may not be able to find a suitable algorithm to solve them. Moreover, model selection in existing tools usually requires users to manually choose based on experience, lacking intelligent support.
[0119] Advantages of this invention: The machine learning model module of this invention embeds various regression, classification, and ensemble learning algorithms (such as linear regression, support vector machine, decision tree, random forest, XGBoost, etc.). Users can select the most suitable algorithm for modeling based on the task type and data characteristics.
[0120] Results Analysis: The flexibility and versatility of this invention allow users to choose the appropriate algorithm for different problems. For complex multi-dimensional data in fields such as medicine, this invention provides more algorithm options, effectively improving the accuracy and flexibility of the analysis.
[0121] (5) Automated report generation improves the standardization of analysis reports.
[0122] Existing technical problems: Traditional analysis tools typically output results in the form of charts or tables, requiring users to spend a significant amount of time compiling these results into reports. This is not only time-consuming but may also lead to omissions or inconsistencies in the report content due to human factors.
[0123] Advantages of this invention: The results display and analysis module of this invention supports the automatic generation of analysis reports, which cover data preprocessing, feature selection, model training and evaluation, and results analysis. The reports include not only numerical indicators but also visual charts and images, facilitating intuitive understanding by researchers.
[0124] Results Analysis: Automated report generation avoids omissions, errors, and non-standardization issues that may occur when manually writing reports. Through a systematic report generation process, not only is the standardization and consistency of reports improved, but researchers also save significant time on report writing, allowing them to focus on higher-level research. This plays a crucial role in accelerating the research process and facilitating data sharing.
[0125] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.
Claims
1. A data processing system based on multi-source data, characterized in that, include: The module includes a user interface module, a feature selection module, a machine learning model module, and a model training module. The user interface module is used to acquire multidimensional data uploaded by the user, acquire the target analysis task of the multidimensional data selected by the user, and acquire the target machine learning algorithm model selected by the user. The feature filtering module is used to select relevant features of the target variable from the multidimensional data based on the target variable; The machine learning model module is used to provide the user with machine learning algorithm models corresponding to different analysis tasks, the different analysis tasks including the target analysis task, and the machine learning algorithm models corresponding to the different analysis tasks including the target machine learning algorithm model. The model training module is used to train the target machine learning algorithm model based on the relevant features of the target variable to obtain a target model, so as to perform the processing corresponding to the target analysis task based on the relevant features of the target variable using the target model; The multidimensional data includes medical record data, genetic data, and image data. The feature filtering module is specifically used for: From the medical record data and / or genetic data, identify keywords that are semantically related to the target variable, and determine the semantically related keywords as the relevant features corresponding to the target variable; The target variable is converted into a target feature. Based on the target feature, a target region similar to the target feature is determined from the image data. The target region is extracted from the image data and used as the relevant feature corresponding to the target variable. For image data, which contains different types of images, the feature filtering module is further used for: For the part corresponding to the target variable, the corresponding parts in different types of images are first aligned. One image in different types of images is used as the base image. Other images in different types of images are rotated according to the shape of the part in the base image so that the shape of the part in the rotated image is the same as the shape of the part in the base image. Based on the target variable, features are extracted from the rotated image and the base image to obtain the relevant features corresponding to each image.
2. The system according to claim 1, characterized in that, The system also includes a data preprocessing module, which is used to preprocess the multidimensional data. The preprocessing includes at least one of missing value imputation, data standardization, outlier detection, and data normalization.
3. The system according to claim 1, characterized in that, The user interface module is specifically used for: Get the first trigger operation of the user on the data upload identifier displayed on the user operation interface, and get the multidimensional data uploaded by the user based on the first trigger operation; Get the user's first selection operation on the target analysis task identifier among the various analysis task identifiers displayed on the user operation interface, and based on the first selection operation, get the target analysis task corresponding to the target analysis task identifier selected by the user; Obtain the user's second selection operation on the target machine learning identifier among the various machine learning identifiers displayed on the user operation interface, and based on the second selection operation, obtain the target machine learning algorithm model corresponding to the target machine learning identifier selected by the user.
4. The system according to claim 1, characterized in that, The system also includes a result display and analysis module, which is used to visualize the processing results output by the model training module.
5. The system according to claim 4, characterized in that, The visualization display types include data distribution maps, model prediction charts, and feature importance bar charts.
6. The system according to claim 2, characterized in that, The system also includes a report generation module, which is used to generate a report based on at least one of the output results of the data preprocessing module, the feature filtering module, the user interface module, and the model training module.
7. The system according to claim 1, characterized in that, Each analysis task corresponds to a processing result. The user interface module is also used for: Obtain the second trigger operation of the user for the target processing result identifier among the various processing result identifiers displayed on the user operation interface; Based on the second triggering operation, the target processing result corresponding to the target processing result identifier is obtained and displayed.
8. The system according to claim 1, characterized in that, The user interface module is also used for: The user's settings for the target parameters among the various analysis parameters displayed on the user interface are obtained. Based on the settings, the set target parameters are obtained and displayed.
9. The system according to claim 8, characterized in that, If the number of each analysis parameter exceeds a set number, all analysis parameters are hidden under the menu corresponding to the parameter setting identifier displayed on the user interface. When the user interface module obtains the user's setting operation for the target parameter among the analysis parameters displayed on the user interface, and based on the setting operation, obtains and displays the set target parameter, it is specifically used for: Obtain the user's third selection operation for the parameter setting identifier, and display the menu corresponding to each of the analysis parameters based on the third selection operation; Obtain the user's settings for the target parameters among the various analysis parameters; based on the settings, obtain and display the set target parameters.
10. The system according to claim 1, characterized in that, The feature selection module is specifically used to: select relevant features of the target variable from the multidimensional data according to the target variable through a target algorithm, wherein the target algorithm is any one of the correlation analysis algorithm, LASSO regression algorithm, and recursive feature elimination algorithm.
Citation Information
Patent Citations
Learning training platform of automatic valuation model
CN117235524A