A data requirement modeling system based on machine learning
By introducing a machine learning-based data demand modeling system into machine learning systems, the difficulties in data modeling and training data set description are solved, and clear description of data requirements and automatic extraction of key features are realized, improving the efficiency and adaptability of data modeling.
Patent Information
- Application Number
- CN202510112613.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The prior art has difficulties in data modeling and training dataset description of machine learning systems, and has failed to effectively consider the requirements of data characteristics in the requirements analysis stage.
A machine learning-based data requirements modeling system is proposed, including interactive interfaces and knowledge bases. The knowledge base consists of a context learning layer and an attribute specification layer. The context learning layer stores a pre-generated feature tree model, and the attribute specification layer defines nodes in the feature tree model. The interactive interface collects user data characteristics and models and guides information through the knowledge base feedback.
It realizes that the data requirements of the machine learning system are clearly described in the requirements analysis stage, automatically identify and extract key features, reduce manual intervention, adapt to data changes, and maintain the effectiveness and scalability of the model.
Smart Images

Figure CN119557289B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning, and in particular to a data demand modeling system based on machine learning. Background Art
[0002] With the development of technology, the field of artificial intelligence (AI) has spawned the rapid development of a new generation of intelligent systems represented by machine learning. However, behind the rapid development, the software development process and software quality are difficult to guarantee, which has become a new problem that needs to be solved in the field of machine learning.
[0003] As one of the main factors affecting software quality, the data model and its construction process reflect whether the developer has a correct analysis and understanding of the target system, and whether the subsequent programming and verification process is based on a reliable requirements analysis. At present, the data model description mechanism and modeling method based on machine learning systems are still in the initial stage, and there are relatively few data modeling methods for machine learning systems.
[0004] In traditional solutions, modeling methods are mainly used to abstract and describe system functions, and the requirements analysis process focuses on functional module design. For example, data modeling work mainly revolves around database design (including data structure and database operation) and big data management, or only discusses data collection of machine learning systems from the perspective of data management. Modeling tools are also only used for database design and visualization.
[0005] However, the new generation of machine learning systems is constantly evolving in the data-driven learning process, making it difficult to apply traditional modeling methods in demand description. The above solutions are not fully applicable to data modeling and training data set description in machine learning systems, and do not consider the data characteristics required in the demand analysis phase of machine learning systems. Summary of the invention
[0006] In order to solve the above problems, this application proposes a data demand modeling system based on machine learning, including an interactive interface and a knowledge base:
[0007] The knowledge base includes a bottom context learning layer and an upper attribute specification layer;
[0008] The context learning layer stores a plurality of pre-generated feature tree models; wherein the feature tree models are selected in advance based on different scenario requirements, corresponding data features are selected from basic elements, and the corresponding feature tree models are obtained by outputting one or more selected machine learning models;
[0009] The attribute specification layer stores a plurality of attribute specifications, wherein the attribute specifications define each node in the feature tree model by means of context attributes;
[0010] The interactive interface collects data features of the data set required by the user, transmits the data features to the knowledge base, and receives modeling guidance information fed back by the knowledge base based on the stored feature tree model, so as to display the modeling guidance information to the user; wherein the modeling guidance information at least includes the form of a feature tree model.
[0011] In one example, the context learning layer includes a metadata model module, a data set module, and an organizational structure module;
[0012] The metadata model module includes a plurality of metadata models, each metadata model corresponds to a single basic element; the data type of the basic element includes at least: an intrinsic element and a learning-related element;
[0013] By using a feature-oriented domain analysis method, data features corresponding to the basic elements are identified, so as to construct the metadata model according to the data features;
[0014] The data set module obtains the data value of the data feature for the metadata model and generates a corresponding data set;
[0015] The organizational structure module stores a feature tree model composed of condition combinations corresponding to the data features; the feature tree model is obtained by selecting a machine learning model according to scenario requirements and data set output.
[0016] In one example, the metadata model for the inherent element is a unified setting, and its corresponding mandatory features include: at least one of: label, date, and presentation mode, and its corresponding optional features include at least: authenticity;
[0017] The metadata model for the learning-related elements is a personalized setting, and its corresponding mandatory features include at least: semantics, and its corresponding optional features include at least: metrics;
[0018] The semantics is used to describe the data background, and the metrics are used to describe the data attributes.
[0019] In one example, the process of constructing the context learning layer includes:
[0020] Acquire multiple dimensional data corresponding to the basic elements, and preprocess the dimensional data;
[0021] Calculating the statistical indicators corresponding to the dimension data and the correlation between the statistical indicators so as to display them through the interactive interface;
[0022] Perform feature selection based on the statistical indicators and the correlation between the statistical indicators to obtain the required data features;
[0023] Select one or more machine learning models from among the pre-stored machine learning models to obtain different machine learning model applications;
[0024] By applying different machine learning models, a corresponding feature tree model is output for the data set, and performance indicators of the feature tree model are counted;
[0025] sorting the feature tree models according to the performance indicators so as to be displayed through the interactive interface;
[0026] The feature tree model selected by the user is added to the context learning layer.
[0027] In one example, feature selection is performed based on the statistical indicators and the correlation between the statistical indicators to obtain the required data features, specifically including:
[0028] Based on the scenario requirements, several most important data features are selected as designated features, and the designated features are used as the required data features;
[0029] According to the correlation between the statistical indicators, data features whose correlation with the specified feature is higher than a first preset degree are selected as the required data features, and data features whose correlation with the specified feature is lower than a second preset degree are deleted;
[0030] Among the remaining data features, among the multiple data features whose correlation with each other is higher than a first preset degree, only one of the data features is retained as the required data feature, and the remaining data features of the multiple data features are deleted;
[0031] The remaining data features are selected as the required data features.
[0032] In one example, in the attribute specification layer, for the attribute specifications corresponding to the non-leaf nodes in the feature tree model, the definition of the non-leaf nodes is dynamically obtained by fusion based on the definition in the context attributes corresponding to the non-leaf nodes.
[0033] In one example, the interactive interface includes a feature selection module, a selected feature display module, and an attribute-based specification module;
[0034] The feature selection module supports the user to select the data features of the required data set and displays the modeling guidance information to the user;
[0035] The selected feature display module displays the data features selected by the user and their corresponding hierarchical structures;
[0036] The attribute-based specification module summarizes the data features selected by the user and their corresponding attribute specifications.
[0037] In one example, the selected feature display module locates the position of the data feature in the modeling guidance information based on the user's operation on the selected data feature;
[0038] The attribute-based specification module automatically generates corresponding attribute specifications through the knowledge base based on user operations.
[0039] In one example, the process of generating the modeling guidance information includes:
[0040] For the feature tree model stored in the knowledge base, starting from the root node, a depth-first search is performed, and during the search process, the feature type and dependency relationship of each node are identified;
[0041] During the search process, if a non-leaf node is found and its feature type is a mandatory feature, the non-leaf node is recorded and the search continues;
[0042] If a non-leaf node is found and its feature type is an optional feature, determine whether it forms a dependency relationship;
[0043] If a dependency relationship is formed and the dependency is judged to be valid, the non-leaf node is recorded;
[0044] If a dependency relationship is formed and the dependency is judged to be invalid, the non-leaf node and its subtree are skipped;
[0045] If no dependency relationship is formed, the non-leaf node is recorded and the guidance sequence composed of the recorded nodes is output as modeling guidance information.
[0046] In one example, the process of generating the modeling guidance information further includes:
[0047] During the search process, if the feature type of the searched non-leaf node is an or type, or an alternative type, all nodes at the same level are searched, and among all nodes at the same level, nodes whose feature types do not constitute a dependency relationship are recorded and output, and based on the user's selection, it is determined whether to output nodes whose feature types constitute a dependency relationship;
[0048] The search traverses to a leaf node. If the leaf node has no dependency relationship with other nodes or the dependency relationship is valid, the leaf node is recorded.
[0049] Otherwise, the leaf node is skipped and the search for the feature tree model continues.
[0050] The data demand modeling system based on machine learning proposed in this application can bring the following beneficial effects:
[0051] A two-layer data modeling approach is proposed, where the bottom layer model serves as the basis for the upper layer model. The bottom layer consists of metadata models to describe the learning environment, while the upper layer contains attribute-based specifications. This enables a clear description of the data requirements of the machine learning system during the requirements analysis phase.
[0052] Based on this approach, a support tool was developed to guide users through the data modeling process. The tool consists of two parts: a knowledge base that stores the pre-prepared knowledge obtained by using the two-layer data modeling approach; and an interactive interface that guides users through data modeling in an interactive way, presenting the hierarchy of selected features and generating specifications for data requirements.
[0053] Based on this, key features can be automatically identified and extracted, reducing the need for human intervention and manual feature selection. The model can be retrained and updated as new data is continuously input to adapt to data changes and remain effective. It is more adaptable to large amounts of data and the model is more scalable. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0055] Figure 1 A schematic diagram of the architecture of data demand modeling based on machine learning in an embodiment of the present application;
[0056] Figure 2 A schematic diagram of the architecture of inherent elements in a metadata model template in one scenario of an embodiment of the present application;
[0057] Figure 3 A schematic diagram of the architecture of learning related elements in a metadata model template in one case of an embodiment of the present application
[0058] Figure 4 This is a schematic diagram of the architecture of a feature tree model in one scenario in an embodiment of the present application;
[0059] Figure 5 This is a flow chart of a method for constructing a context learning layer in one scenario in an embodiment of the present application;
[0060] Figure 6 This is a schematic diagram of an interactive interface in one scenario in an embodiment of the present application. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0062] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0063] like Figure 1 As shown, an embodiment of the present application provides a data demand modeling system based on machine learning, including: an interactive interface and a knowledge base.
[0064] In a machine learning system, a knowledge base is used to store knowledge, models, algorithms, and data sets related to machine learning, so that users can efficiently access, retrieve, and share the required information. The knowledge in the knowledge base in the embodiment of the present application is formed based on a two-layer data requirement modeling method for a machine learning system. The bottom layer uses a metadata model for contextual learning to describe the elements in the machine learning system and its environment and the relationships between them. The feature-oriented analysis (FODA) method is used to describe the context through a feature model, and the definition, relationship, and constraints of the features are given. The upper layer is a set of specifications based on attributes. The data requirements of the machine learning system are described and standardized by deriving the required attributes from good contextual learning.
[0065] Specifically, the knowledge base mainly includes the underlying context learning layer and the upper attribute specification layer.
[0066] like Figure 1 As shown, the context learning layer is mainly used to describe the elements of the machine learning system, the environment and the relationship between the two. It stores multiple pre-generated feature tree models. Among them, the feature tree model is a corresponding feature tree model obtained by selecting corresponding data features from the basic elements based on different scenario requirements and outputting one or more selected machine learning models. The machine learning model can be stored in the context learning layer or obtained through external interface docking.
[0067] The property-based specification layer contains multiple property specifications. The property specifications define each node in the feature tree model through context attributes.
[0068] The interactive interface is mainly used to interact with users. It is responsible for presenting modeling guidance information, collecting dataset features and attributes required by users, and displaying the currently selected features and their hierarchical structure.
[0069] The interactive interface collects the data features of the data set required by the user through the user's selection and transmits the data features to the knowledge base. The knowledge base stores the feature tree model, which can feedback the modeling guidance information based on the feature tree model stored in it to display the modeling guidance information to the user. Among them, the modeling guidance information at least includes the feature tree model, and of course can also include other information, such as the architecture and parameters of the machine learning model corresponding to the feature tree model.
[0070] In one embodiment, Figure 1 As shown, the context learning layer includes a metadata model module, a dataset module, and an organizational structure module.
[0071] The metadata model module includes multiple metadata models. The metadata model is the basic element of a specific data model, which can describe the structure and characteristic information of the data. Each metadata model corresponds to a single basic element.
[0072] Specifically, basic elements include two types: intrinsic elements and learning-related elements. Intrinsic elements are independent of specific machine learning applications and are more general in various machine learning applications, while learning-related elements depend on specific machine learning applications and are particularly suitable for a certain machine learning application.
[0073] For example, in the task of data modeling, taking the product recommendation system of an online store as an example, the metadata model can be shown in Table 1:
[0074] Table 1 Metadata model diagram
[0075]
[0076] Through the feature-oriented domain analysis method (FODA), the data features corresponding to the basic elements are identified to build a metadata model based on the data features. The feature-oriented domain analysis method is mainly used to identify, analyze and organize the features of the software system from the perspective of the domain in order to support the variability and reusability of the system. The process can be roughly divided into: modeling according to requirements, feature extraction based on modeling, organizing and classifying the extracted features, and representing the features and their relationships as feature models.
[0077] In the example of Table 1, for "user", it can be analyzed that the feature "recently purchased products" corresponds to the product IDs that the user recently purchased; for "product", it can be analyzed that the feature "product category" corresponds to a specific category label, such as "electronic products" or "clothing".
[0078] In the dataset module, for the metadata model, the data values of the data features are obtained and the corresponding datasets are generated. The dataset module contains multiple datasets. The dataset is the element value or attribute corresponding to the specified metadata model and can be regarded as a set of feature values. In the context of multimodal data forms, the forms of datasets are also diverse.
[0079] Still taking the example of Table 1 as an example, the data set of user A can be:
[0080] User ID: 123; Registration date: 2021-01-01; Most recently purchased products: Product ID: 456; Purchase frequency: 5 times.
[0081] The data set for product B can be:
[0082] Product ID: 456; Name: Smart watch; Price: 299; Product category: Electronic products; Sales volume: 1500.
[0083] Furthermore, in order to support different machine learning tasks, a unified model can be used to describe the characteristics of intrinsic elements. At the same time, the feature model of learning related elements needs to be adjusted according to different machine learning systems.
[0084] like Figure 2 and Figure 3 As shown, a corresponding metadata model template is pre-set.
[0085] In the template, the metadata model for the inherent element is uniformly set (that is, in different inherent elements, the architecture of the corresponding metadata model is consistent), and its corresponding mandatory features include: at least one of: label, date, and presentation method, and its corresponding optional features include at least: authenticity.
[0086] The metadata model for learning-related elements is a personalized setting (that is, the architecture of the corresponding metadata model may be inconsistent in different learning-related elements), and its corresponding mandatory features include at least: semantics, and its corresponding optional features include at least: measurement.
[0087] The feature model of learning related elements consists of two parts: semantics and metrics of each data set. Semantics is a mandatory feature that describes the data background and how the system and structure work, including the source, type, purpose, relationship and organization of the data. Metrics are optional features that describe data attributes and include attributes that can be used to measure data sets, providing more dimensional information for data analysis.
[0088] Among them, features can be divided into mandatory features and optional features. Mandatory features refer to features that must be included. Each instance needs to include this feature when using the feature model. Optional features refer to features that can be selected or not. Users decide whether to add this feature based on their needs. For example, in Figure 2 and Figure 3 In the , "label", "yes, no", "date", "presentation", and "semantics" are mandatory features, and "authenticity", "coverage", and "measurement" are optional features.
[0089] The relationship between features can also include or groups (for a single feature, its feature type is also called or type), alternative groups (for a single feature, its feature type is also called alternative type), and dependency relationships. Or groups refer to the fact that features are selectable in this group, and users can select one or more features. Alternative groups refer to the fact that features are mutually exclusive in this group, and users can only select one feature. Dependency relationships refer to the dependency relationship between features, and the existence or value of a feature value depends on another feature.
[0090] Still taking the example of Table 1, taking the data set of user A as an example, an example of its inherent elements can be: the label is "user ID123", the date is "2021-01-01", the presentation method is text, and the authenticity can be judged based on the data source of this data to give a score or result. Coverage is an attribute used to describe the scope or applicable scope of the data set. The scope here can refer to the dimensions such as fields, time periods, geographical areas, age ranges, etc. covered by the data set information. For example, it can be "sales data of Company A from 2023 to 2024", "sales of product ID456 in various countries", "sales data of product ID456 among people aged 18-35", etc.
[0091] As for learning related elements, because "recently purchased products" and "purchase frequency" are closely related, they describe the purchase behavior of the same user (user ID: 123) for the same product (product ID: 456), and the purchase frequency is the number of times a specific product is purchased. The two together express the purchase behavior of a certain product and are features of the same dimension, so the two are regarded as different contents in the same semantic branch, that is, they belong to the same semantics.
[0092] As for metrics, they can take many forms. Here are some common examples to explain:
[0093] 1. Quantity-related metrics, such as average, sum, standard deviation, growth rate, etc., all belong to this category. For example, under the semantics of "recently purchased product ID: 456", it may have metrics such as "total sales of product ID 456", "average sales amount per product", and "sales growth rate of product 456".
[0094] 2. Frequency metrics, such as frequency of occurrence, frequency distribution, click-through rate, etc. For example, if you want to advertise the product "Product ID: 456", the frequency metrics in this activity may include "ad click-through rate", "ad display frequency", "average ad click frequency of user ID: 123", etc.
[0095] 3. Association metrics. For example, Pearson correlation coefficient, Spearman rank correlation coefficient, covariance, mutual information, etc. For example, to analyze the association between advertising and sales of "product ID: 456", the metrics include "Pearson correlation coefficient between advertising expenditure and sales of product ID: 456" and "mutual information between users' purchasing behavior and their age groups".
[0096] The organizational structure module helps to clearly display the relationship between feature nodes of different data features. It stores the feature tree model composed of condition combinations corresponding to data features. Among them, the feature tree model is obtained by selecting a machine learning model based on scenario requirements and data set output.
[0097] Commonly used feature structures include OR / AND tree structures (herein referred to as feature tree models, and can also be referred to as tree structures, feature trees, etc.). AND relationships mean that all conditions must be met at the same time, and OR relationships mean that at least one condition must be met. By setting up a tree structure, data set information can be classified according to hierarchical relationships, making complex logical relationships clear at a glance and helping to locate target data more quickly.
[0098] Here we assume that the current scenario requirement is to recommend electronic products to users. The recommendation conditions can be based on the user's purchasing behavior and product characteristics. The feature tree structure (that is, the feature tree model) can be designed as follows Figure 4 As shown, the node description is as follows:
[0099] Scenario requirement: recommend electronic products.
[0100] Under the AND node, inclusion condition 1 (AND) can be: A: User's purchase frequency > 3 times, B: User's most recently purchased product category = electronic products. This means that this condition is only met if the user's purchase frequency is greater than 3 times and has recently purchased electronic products.
[0101] Condition 2 (AND) can be: C: the average rating of the user > 4.0, D: the number of electronic products the user has browsed > 5. This means that this condition can be met when the user has a high rating and has browsed multiple electronic products.
[0102] Under the OR node, condition 3 (OR) can be: E: the user's rating of a certain type of electronic product is > 4.5, F: the user has mentioned related products on social media. Users only need to meet one of the conditions, such as a rating higher than 4.5 or a related product has been mentioned on social media, to be recommended to them.
[0103] In the contextual learning layer, knowledge (embodied by basic elements, data features, etc.) is organized through a feature tree model with the same hierarchical structure, stored in JavaScript Object Notation (JSON) format to form a knowledge base, and continuously expanded or updated.
[0104] JSON has simple syntax, clear hierarchical structure, and is easy for computers to parse. Each non-leaf node contains two keywords: Type (recording the node type, including mandatory, optional, alternative, and or) and Relation (recording the relationship between nodes, including None and dependency). Each leaf node contains four keywords: Type, DefinitionType, DefinitionConstraint, and Relation. DefinitionType records the value type used by the node definition, and DefinitionConstraint is used to specifically define the constraints of the node.
[0105] In one embodiment, the above describes the architecture of the data requirement modeling system based on machine learning, and the architecture of each layer of the knowledge base contained therein. Figure 5 As shown in Figure 1, the construction process of the context learning layer includes:
[0106] S1: Acquire multiple dimensional data corresponding to basic elements, and preprocess the dimensional data.
[0107] The data demand modeling system based on machine learning (referred to as the modeling system) receives and stores the basic elements uploaded by users, presenting them in the form of table data, with the first row containing non-repeated column names. The data in the form of table ensures structure and facilitates subsequent processing. The character encoding is UTF-8 or GB2312 to prevent garbled characters when processing data.
[0108] The modeling system cleans and transforms the collected data to prepare for the subsequent feature value extraction. The transformation processing methods include missing value processing, outlier processing, standard normalization, maximum and minimum value normalization, and norm normalization.
[0109] S2: Calculate the statistical indicators corresponding to the dimension data and the correlation between the statistical indicators so as to display them through the interactive interface.
[0110] By selecting and evaluating data features, the most valuable features are extracted. The feature data is calculated to obtain multiple statistical indicators (herein referred to as statistical indicators), including mean, mode, median, quantile, variance and standard deviation, etc., and the correlation between data is calculated, including mutual information, F test, P value, chi-square test value, Pearson correlation coefficient, Spearman correlation coefficient and other correlation coefficients.
[0111] Visualization tools can be used to visualize statistical indicators in the interactive interface. For example, indicator results can be plotted as histograms, pie charts, and violin plots, and indicator results of correlation calculations can be plotted as scatter plots.
[0112] S3: Perform feature selection based on the statistical indicators and the correlation between the statistical indicators to obtain the required data features.
[0113] After obtaining the visualized statistical indicators, users can select the most valuable features based on some criteria, such as high correlation, important category features, removal of low correlation and avoidance of multicollinearity, so as to perform further processing and obtain the required data features, thereby reducing the data dimension and enhancing the validity of the data.
[0114] Specifically, first, based on the scenario requirements, select several of the most important data features as designated features, and use the designated features as the required data features. For example, if the scenario requirement is to recommend electronic products to users, then price and brand are often among the most important data features, and they are used as the initially selected required data features as designated features. The selection process can be obtained through machine learning models such as large language models as the first batch of required data features.
[0115] According to the correlation between the statistical indicators, the data features with a correlation with the specified features higher than the first preset degree are selected as the required data features, as the second batch of required data features. The data features with a high correlation with the specified features are selected as the required data features. For example, memory, storage and price usually have a strong positive correlation, so these features are the most valuable.
[0116] Data features with a correlation with the specified feature lower than a second preset level are deleted. Features with a low correlation lower than a second preset level (which is lower than the first preset level) are deleted. Some features (such as battery capacity) have a low correlation with price, so the feature is removed.
[0117] Among the remaining data features, among the multiple data features whose correlation is higher than the first preset degree, only one of the data features is retained as the data features required for the third batch, which is important in different scenario requirements. If two features (for example, memory and storage) are highly correlated, it can be considered to retain only one of the features to avoid multicollinearity. At this time, the remaining data features of the multiple data features are deleted.
[0118] The remaining data features are selected as the required data features as the fourth batch of required data features.
[0119] In step S2 and step S3, it is assumed that there is a data set about a smartphone for example for explanation.
[0120] It can contain the following feature data: screen size, processor type, memory, storage, camera pixels, battery capacity, screen resolution, brand, release year, etc. The current scenario requirement is to predict the price of smartphones through these data features.
[0121] First, calculate the statistical indicators. Taking screen size as an example, you can calculate the average screen size of all mobile phones (used to determine the concentrated range of screen sizes of most mobile phones), the most common screen size (for example, whether a large number of mobile phones have a screen size of 6.1 inches), the median screen size (to understand whether the distribution of data is biased towards large or small screens), and check whether 75% of mobile phone screen sizes are smaller than a certain value (to help understand the distribution of screen sizes).
[0122] Then calculate the correlation between statistical indicators. Taking the Pearson correlation coefficient as an example, we can calculate the linear correlation between memory and price, the linear correlation between storage and price, and the linear correlation between camera pixels and price. For some features, such as the nonlinear correlation between processor type and price, we can calculate the corresponding indicator of the Spearman coefficient. If the price is classified into three categories: high, medium and low, the chi-square test can be used to deal with the relationship between price and brand or processor type, etc.
[0123] Then visualize the results. For example, draw a histogram to show numerical features like memory, storage, battery capacity, etc., so that users can observe the distribution. Draw a violin plot to show correlations, such as the relationship between brand or processor type and price. Draw a scatter plot to reveal linear / non-linear relationships, such as the relationship between memory and storage. If there is a positive correlation between memory and price, then memory is an important feature.
[0124] Finally, feature selection is performed. After analysis, the most valuable features are selected according to the corresponding standards to obtain the required data features for the next step of processing. The required data features may include: Highly correlated features: For example, there is usually a strong positive correlation between memory, storage and price, so these features are the most valuable; Important category features: For example, the brand usually has a significant impact on the price, especially some high-end brands will significantly increase the price. At the same time, low-correlation features can also be removed. Some features (such as battery capacity) have a low correlation with price, so the feature is removed. Multicollinearity can also be avoided. If two features (such as memory and storage) are highly correlated, you can consider retaining only one of them.
[0125] S4: Select one or more machine learning models from the pre-stored machine learning models to obtain different machine learning model applications.
[0126] The modeling system uses a single machine learning model or a combination of multiple machine learning models to train and identify the data using the extracted feature values. The machine learning model can be selected by the user or automatically selected.
[0127] The machine learning model includes but is not limited to the logistic regression algorithm, the naive Bayes classification algorithm, the decision tree classification algorithm, the linear regression algorithm, the ridge regression algorithm, the K-means clustering algorithm, the hierarchical clustering algorithm, the linear discriminant analysis algorithm, the t-SNE algorithm, etc. When one or more of them are selected, they are applied as machine learning models. Among them, if multiple machine learning models are selected, they can be identified by weighted summing of the results of multiple models, fusing the hierarchies of multiple models, etc.
[0128] The above machine learning models cover different learning methods (supervised learning, unsupervised learning, etc.), which improves the diversity of model selection. You can select algorithms based on data features, giving priority to classification algorithms. If it is a regression problem, consider linear regression, decision tree regression, etc. For large-scale data sets, you can choose efficient algorithms such as random forest and XGBoost. Considering the complementarity between algorithms, you can choose multiple algorithm combinations. For example, the combination of random forest and XGBoost under the ensemble method has both the robustness of random forest and the powerful model that can capture complex patterns using XGBoost.
[0129] S5: By applying different machine learning models, a corresponding feature tree model is output for the data set, and performance indicators of the feature tree model are counted.
[0130] Specifically, k-fold cross validation can be used to evaluate the stability and generalization ability of each machine learning model application. Each model will be evaluated in k trainings to avoid deviations caused by single-division dataset problems. Performance indicators used for ranking evaluation include accuracy, precision, recall, F1 value, AUC-ROC, etc. For multiple algorithm combinations, ranking can be performed based on the combined performance (for example, average accuracy, weighted F1 value, etc.), so as to obtain the state of each machine learning model application when outputting the feature tree model, and measure the state through performance indicators.
[0131] S6: Sort the feature tree models according to the performance indicators so as to be displayed through the interactive interface.
[0132] The ranking of the selected single machine learning algorithm or multiple machine learning algorithm combinations is displayed. According to the results of the performance evaluation, the algorithms or combinations are sorted in descending order, and the ranking list is displayed to the user. The single machine learning algorithm or multiple machine learning algorithm combinations ranked first are regarded as the optimal processing situation, and the result is recommended to the user.
[0133] S7: Add the feature tree model selected by the user to the context learning layer.
[0134] If the user selects by ranking, the model will be trained after the user selects one or more machine learning algorithms, and the results will be stored in the knowledge base in the form of a feature tree model. If the user does not make a selection, the feature tree model output by the top-ranked machine learning model will be added to the context learning layer and thus added to the knowledge base.
[0135] In one embodiment, for the attribute specification layer, the attribute specifications are used to define leaf nodes and non-leaf nodes in the feature tree model. Among them, the leaf node refers to the bottom node in the tree structure, which is the terminal node of the tree, and there is no child node below. It generally represents the final feature classification or the specific value of the feature. The non-leaf node is the intermediate node of the tree structure. It has child nodes and can pass information or data to the child nodes. The non-leaf node generally refers to the intermediate step of a decision, or the description of a certain attribute, or the judgment of a certain feature to guide subsequent branches.
[0136] For non-leaf nodes, they can be defined as follows:
[0137] 1. Represent the intermediate state of the decision, helping the tree to deduce the flow or classification path of data from the root to the leaves.
[0138] 2. Represents a conditional judgment that determines how data is further segmented or classified and controls the logic of data branching.
[0139] 3. Represent the features of the middle layer, passing the feature information of the previous layer to the next layer until the final result is obtained.
[0140] Unlike traditional data modeling methods, data modeling methods for machine learning systems are more dynamic and rely on a data-driven self-evolution process. Therefore, in addition to the fixed scenarios mentioned above, which need to be defined in a fixed way, in some specific scenarios, it is difficult to describe their behavior or data form in a fixed way. Instead, they need to be expressed with the help of attributes based on the learning context. The definition should be made based on specific analysis of specific reasons.
[0141] At this time, for the attribute specifications corresponding to the non-leaf nodes in the feature tree model, based on the definitions in the context attributes corresponding to the non-leaf nodes, the definitions of the non-leaf nodes are dynamically obtained through fusion.
[0142] For example, if we specify ω(m, <e>) represents the value of element e in the feature tree model m, and parent.child represents the value of the child element child of element parent. Since leaf nodes have a relatively fixed definition in the feature tree model, if e is a non-node, ω(m, <e>)=ω(m,<e.child> )=v to indicate that the value of the element e of the non-leaf node in the feature tree model m is v.
[0143] In one embodiment, Figure 6 As shown, the interactive interface includes a feature selection module, a selected feature display module, and an attribute-based specification module.
[0144] The feature selection module (located in the middle pane) provides modeling guidance information for users, supports users to select the required data features in the dataset, defines these features according to the actual needs of the project, and displays modeling guidance information to users.
[0145] The selected feature display module (located in the left pane) displays the data features selected by the user and their corresponding hierarchical structure, which is convenient for users to intuitively understand and manage. It can also locate the position of the data feature in the modeling guidance information based on the user's operation (for example, click operation) in the selected data feature, so that users can make necessary modifications.
[0146] The attribute-based specification module (located in the lower pane) summarizes the data features selected by the user and their corresponding attribute specifications. Based on the user's operation, you can also click the button in the upper left corner to automatically generate the corresponding attribute specifications through the knowledge base.
[0147] In one embodiment, Figure 6 On the basis of the model, an important function of the interactive interface is to guide users to perform data modeling. The generation process of modeling guidance information includes:
[0148] For the feature tree models stored in the knowledge base (for example, according to the scenario requirements selected by the user, the closest feature tree model is selected), starting from the root node, a depth-first search is performed, and during the search process, the feature type and dependency relationship of each node are identified.
[0149] During the search process, if a non-leaf node is found and its feature type is a mandatory feature, it means that the non-leaf node is a necessary node, and the non-leaf node is recorded and the search continues. If a non-leaf node is found and its feature type is an optional feature, it is determined whether it constitutes a dependency relationship.
[0150] If a dependency relationship is established and the dependency is judged to be valid, the dependency relationship of the node is considered to be established and the non-leaf node can be recorded. If a dependency relationship is established and the dependency is judged to be invalid, the non-leaf node and its subtree are skipped and no longer searched, and other non-leaf nodes are searched.
[0151] If no dependency relationship is formed, it is also considered that the non-leaf node can be recorded, and the guidance sequence composed of the recorded nodes is output as the modeling guidance information.
[0152] Furthermore, during the search process, it can be determined whether to continue searching the subtree in depth according to the user's selection.
[0153] If the feature type of the searched non-leaf node is "or type" or "alternative type", all nodes at the same level are searched, and among all nodes at the same level, nodes whose feature types do not constitute a dependency relationship are recorded and output, and based on the user's selection, it is determined whether to output nodes whose feature types constitute a dependency relationship. At this time, the user can choose whether to record and output nodes whose feature types constitute a dependency relationship according to their own needs.
[0154] Continue searching until the leaf node is reached. If the leaf node has no dependency relationship with other nodes, or the dependency relationship is valid, the leaf node is recorded and the corresponding guidance sequence is output.
[0155] Otherwise, skip the leaf node and continue searching the feature tree model until all nodes of the feature tree model are searched.
[0156] A two-layer data modeling approach is proposed, where the bottom layer model serves as the basis for the upper layer model. The bottom layer consists of metadata models to describe the learning environment, while the upper layer contains attribute-based specifications. This enables a clear description of the data requirements of the machine learning system during the requirements analysis phase.
[0157] Based on this approach, a support tool was developed to guide users through the data modeling process. The tool consists of two parts: a knowledge base that stores the pre-prepared knowledge obtained by using the two-layer data modeling approach; and an interactive interface that guides users through data modeling in an interactive way, presenting the hierarchy of selected features and generating specifications for data requirements.
[0158] Based on this, key features can be automatically identified and extracted, reducing the need for human intervention and manual feature selection. The model can be retrained and updated as new data is continuously input to adapt to data changes and remain effective. It is more adaptable to large amounts of data and the model is more scalable.
[0159] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.< / e> < / e>
Claims
1. A data requirement modeling system based on machine learning, characterized in that: Includes interactive interface and knowledge base; The knowledge base includes a bottom context learning layer and an upper attribute specification layer; The context learning layer stores a plurality of pre-generated feature tree models; wherein the feature tree models are selected in advance based on different scenario requirements, corresponding data features are selected from basic elements, and the corresponding feature tree models are obtained by outputting one or more selected machine learning models; The attribute specification layer stores a plurality of attribute specifications, wherein the attribute specifications define each node in the feature tree model by means of context attributes; The interactive interface collects data features of a data set required by a user, transmits the data features to the knowledge base, and receives modeling guidance information fed back by the knowledge base based on a stored feature tree model, so as to display the modeling guidance information to the user; wherein the modeling guidance information at least includes the form of a feature tree model; The construction process of the context learning layer includes: Acquire multiple dimensional data corresponding to the basic elements, and preprocess the dimensional data; Calculating the statistical indicators corresponding to the dimension data and the correlation between the statistical indicators so as to display them through the interactive interface; According to the statistical indicators and the correlation between the statistical indicators, feature selection is performed to obtain the required data features, specifically including: based on the scenario requirements, selecting several most important data features as designated features, and using the designated features as the required data features; according to the correlation between the statistical indicators, selecting data features whose correlation with the designated features is higher than a first preset degree as the required data features, and deleting data features whose correlation with the designated features is lower than a second preset degree; among the remaining data features, among the multiple data features whose correlation with each other is higher than the first preset degree, only one of the data features is retained as the required data feature, and the remaining data features of the multiple data features are deleted; and the remaining data features are selected as the required data features; Select one or more machine learning models from among the pre-stored machine learning models to obtain different machine learning model applications; By applying different machine learning models, a corresponding feature tree model is output for the data set, and performance indicators of the feature tree model are counted; sorting the feature tree models according to the performance indicators so as to be displayed through the interactive interface; Add the feature tree model selected by the user to the context learning layer; The generation process of the modeling guidance information includes: For the feature tree model stored in the knowledge base, starting from the root node, a depth-first search is performed, and during the search process, the feature type and dependency relationship of each node are identified; During the search process, if a non-leaf node is found and its feature type is a mandatory feature, the non-leaf node is recorded and the search continues; If a non-leaf node is found and its feature type is an optional feature, determine whether it forms a dependency relationship; If a dependency relationship is formed and the dependency is judged to be valid, the non-leaf node is recorded; If a dependency relationship is formed and the dependency is judged to be invalid, the non-leaf node and its subtree are skipped; If no dependency relationship is formed, the non-leaf node is recorded and the guidance sequence composed of the recorded nodes is output as modeling guidance information; The generation process of the modeling guidance information further includes: During the search process, if the feature type of the searched non-leaf node is an or type, or an alternative type, all nodes at the same level are searched, and among all nodes at the same level, nodes whose feature types do not constitute a dependency relationship are recorded and output, and based on the user's selection, it is determined whether to output nodes whose feature types constitute a dependency relationship; The search traverses to a leaf node. If the leaf node has no dependency relationship with other nodes or the dependency relationship is valid, the leaf node is recorded. Otherwise, skip the leaf node and continue searching the feature tree model; In the attribute specification layer, for the attribute specifications corresponding to the non-leaf nodes in the feature tree model, based on the definitions in the context attributes corresponding to the non-leaf nodes, the definitions of the non-leaf nodes are dynamically obtained by fusion; The interactive interface includes a feature selection module, a selected feature display module, and an attribute-based specification module; The feature selection module supports the user to select the data features of the required data set and displays the modeling guidance information to the user; The selected feature display module displays the data features selected by the user and their corresponding hierarchical structures; The attribute-based specification module summarizes the data features selected by the user and their corresponding attribute specifications; The selected feature display module locates the position of the data feature in the modeling guidance information based on the user's operation on the selected data feature; The attribute-based specification module automatically generates corresponding attribute specifications through the knowledge base based on user operations.
2. The system according to claim 1, characterized in that The context learning layer includes a metadata model module, a data set module, and an organizational structure module; The metadata model module includes a plurality of metadata models, each metadata model corresponding to a single basic element; The data types of the basic elements include at least: intrinsic elements and learning-related elements; By using a feature-oriented domain analysis method, data features corresponding to the basic elements are identified, so as to construct the metadata model according to the data features; The data set module obtains the data value of the data feature for the metadata model and generates a corresponding data set; The organizational structure module stores a feature tree model composed of condition combinations corresponding to the data features; the feature tree model is obtained by selecting a machine learning model according to scenario requirements and data set output.
3. The system according to claim 2, characterized in that The metadata model for the inherent element is a unified setting, and its corresponding mandatory features include: at least one of: label, date, and presentation mode, and its corresponding optional features include at least: authenticity; The metadata model for the learning-related elements is a personalized setting, and its corresponding mandatory features include at least: semantics, and its corresponding optional features include at least: metrics; The semantics is used to describe the data background, and the metrics are used to describe the data attributes.