Machine learning model construction method and device, storage medium and program product
By receiving user input modeling requirements and training data sets, using pre-trained language models to generate detailed modeling configuration information, and automatically controlling automated machine learning tools to perform model construction tasks, solving the problems of cumbersome construction process and high professional knowledge requirements in traditional machine learning models, and achieving an efficient and automated model construction process.
Patent Information
- Application Number
- CN202510374713.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
The construction process of traditional machine learning models is cumbersome, time-consuming, and requires high professional knowledge of operators, which limits non-professional personnel to quickly and efficiently complete model construction.
By receiving the modeling requirements and training data sets input by users, modeling configuration prompts are generated, and detailed modeling configuration information is generated using pre-trained language models, and automated machine learning tools are automatically controlled to perform machine learning model construction tasks, realizing the full process automation from data processing to model training.
It lowers the professional threshold and significantly improves modeling efficiency, allowing users to obtain high-quality machine learning models without deep understanding of machine learning theory, reducing the errors that may be caused by manual settings.
Smart Images

Figure CN120218285A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the technical field of model construction, and particularly to a method for constructing a machine learning model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Currently, with the continuous development of big data and artificial intelligence technologies, machine learning has been widely applied in various fields. However, traditional modeling processes usually rely on manual operations by data or algorithm experts, including steps such as feature selection, algorithm decision-making, and parameter adjustment. This approach not only makes the modeling process cumbersome and time-consuming, but also requires high professional knowledge of the operators, making it difficult for non-professionals to quickly and efficiently complete model construction, thus limiting the popularization and promotion of the technology. Summary of the Invention
[0003] In view of this, one or more embodiments of this specification provide a method for constructing a machine learning model, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:
[0005] According to a first aspect of one or more embodiments of this specification, a method for constructing a machine learning model is proposed, including:
[0006] Receiving a modeling requirement input by a user and a training data set for generating training samples, where the modeling requirement is at least used to describe the task and performance requirements of the machine learning model to be constructed;
[0007] Generating a modeling configuration prompt based on the modeling requirement and the training data set, and inputting the modeling configuration prompt into a pre-trained language model to generate modeling configuration information for describing the configuration content of each link in the automated modeling process by using the pre-trained language model;
[0008] Based on the training data set and the modeling configuration information, controlling an automated machine learning tool to execute a machine learning model construction task, where the machine learning model construction task includes: constructing a machine learning model to be trained and training samples, and training the machine learning model to be trained by using the training samples to obtain a trained machine learning model.
[0009] According to a second aspect of the embodiments of this specification, an electronic device is provided, including:
[0010] A processor;
[0011] A memory for storing instructions executable by the processor;
[0012] Wherein, when the processor executes the executable instructions, it is used to implement the method described in the first aspect.
[0013] According to a third aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in the first aspect.
[0014] According to a fourth aspect of the embodiments of the present specification, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.
[0015] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects:
[0016] In the embodiments of the present specification, the user can input the modeling requirements and the training data set for generating training samples according to actual needs, without the need to perform other operations, reducing the professional threshold and improving the user experience. The device automatically generates modeling configuration prompts by combining the modeling requirements and the training data set input by the user, and uses a pre-trained language model to generate detailed modeling configuration information, optimizing the construction process of the machine learning model, greatly shortening the time required for traditional manual configuration, reducing the errors that may be brought by manual settings, and finally automatically controlling the machine learning tool to execute the machine learning model construction task according to the training data set and the modeling configuration information, realizing the full-process automation from data processing to model training, significantly improving the modeling efficiency, and enabling the user to obtain a high-quality machine learning model without in-depth knowledge of machine learning theory.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a schematic diagram of the architecture of a machine learning model construction service system provided by an exemplary embodiment.
[0019] Figure 2 is a flowchart of a method for constructing a machine learning model provided by an exemplary embodiment.
[0020] Figure 3 is a schematic diagram of a machine learning model construction template provided by an exemplary embodiment.
[0021] Figure 4 is a schematic diagram of constructing a machine learning model provided by an exemplary embodiment.
[0022] Figure 5 is a schematic diagram of evaluating a machine learning model provided by an exemplary embodiment.
[0023] Figure 6 It is a schematic diagram of an optimized training report provided by an exemplary embodiment.
[0024] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment. Detailed implementation manners
[0025] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0026] It should be noted that: in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0027] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.
[0028] Based on the problems in the related art, the embodiments of this specification provide a method for constructing a machine learning model, an electronic device, a computer-readable storage medium, and a computer program product.
[0029] In a possible application scenario, based on the user's privacy protection requirements, the method for constructing a machine learning model provided by the embodiments of this specification can be directly deployed in the electronic device to which the user belongs. The electronic device includes but is not limited to physical servers, server clusters, cloud servers, smart phones / mobile phones, tablet computers, personal digital assistants (PDAs), laptop computers, and desktop computers, etc.
[0030] In another possible application scenario, as Figure 1 shown Figure 1It is a schematic diagram of the architecture of a machine learning model construction service system provided by an exemplary embodiment. The system may include a server 11, a network 12, and several user terminals, such as a PC (Personal Computer) 13, a mobile phone 14, etc.
[0031] The server 11 may be a physical server including an independent host, or the server 11 may be a virtual server hosted by a host cluster. During operation, the server 11 may run the server-side program of the machine learning model construction application to implement the construction service platform for the corresponding machine learning model.
[0032] The PC 13 and the mobile phone 14 are only some types of user terminals that users can use. In fact, users can obviously also use user terminals of the following types: tablet devices, laptop computers, personal digital assistants (PDAs), wearable devices (such as smart glasses, smart watches, etc.). One or more embodiments of this specification do not limit this. During operation, the user terminal may run the client-side program of the machine learning model construction application and can be implemented as the client of the machine learning model construction service. Among them, the application program of the client of the above-mentioned machine learning model construction service can be started and run on the user terminal. The client-side program may be a native application installed on the user terminal, or the client-side program may be a small program, a fast application, or other similar forms. Of course, when using web technologies such as HTML5 or similar, relevant functions can be implemented through the page displayed by the browser. Here, the browser may be an independent browser application or a browser module embedded in some applications.
[0033] For the network 12 for interaction between user terminals such as the PC 13 and the mobile phone 14 and the server 11, it can be specifically selected to use a wired or wireless network to achieve communication based on the communication methods supported by the corresponding user terminals. This specification does not limit this. For example, the PC 13 can support both wired and wireless communication, so wired or wireless networks can be used to achieve communication according to needs, while the mobile phone 14 usually only supports wireless communication, so a wireless network can be used to achieve communication.
[0034] The machine learning model construction method provided by the embodiments of this specification can be executed by the server 11, or by the user terminal, or a part of the machine learning model construction method can be executed by the user terminal and another part by the server 11. This embodiment does not make any restrictions on this.
[0035] In some embodiments, please refer to Figure 2, showing a flowchart of a method for constructing a machine learning model, applied to an electronic device (such as the above-mentioned user terminal, server, or other electronic devices), the method includes:
[0036] In S201, receive the modeling requirements input by the user and the training data set for generating training samples, where the modeling requirements are at least used to describe the tasks and performance requirements of the machine learning model to be constructed.
[0037] In this step, the electronic device receives the modeling requirements and the training data set input by the user, where the modeling requirements are at least used to describe the tasks and performance requirements of the model to be constructed, enabling the electronic device to accurately capture the user's modeling intention, ensuring that subsequent configurations closely follow the actual requirements. The training data set is used to construct multiple training samples, and each training sample contains the input features and labels of the model. Users do not need to master complex technical details, thereby lowering the professional threshold and providing basic input for subsequent automated processing, enabling non-professional users to accurately express their requirements and improving the overall user experience.
[0038] For example, suppose a user wants a credit risk assessment model, whose task is to predict whether a borrower is likely to default based on information such as the borrower's historical credit behavior, salary level, borrowing purpose, etc. The performance requirement is: accuracy rate ≥ 85%.
[0039] Another example is that a user wants a credit card fraud warning model, whose task is to analyze the transaction patterns and behavioral characteristics of customers and detect and warn of potential credit card fraud in real time. The performance requirement is: recall rate ≥ 95%, especially in the case of a large volume of transactions, it is necessary to ensure that every potential fraud transaction can be detected.
[0040] Still another example is that a user wants a recommendation system, whose task is to recommend personalized products or content based on the user's historical behavior, interest preferences, and other relevant data, improving the user's purchase conversion rate and customer stickiness. The performance requirement is: accuracy rate ≥ 85%, and the output of the recommendation system should have a certain degree of interpretability.
[0041] In a possible implementation, the electronic device can provide input controls, and the user can input the modeling requirements through the input controls according to actual needs and upload the relevant training data set. It can be understood that the input method of the modeling requirements can be keyboard input (including but not limited to physical keyboards, virtual keyboards, etc.), voice input (voice-to-text), handwriting input (such as writing with a finger or a stylus on a touch screen), or scanning input (such as taking a picture of a paper document with a scanner, a smartphone camera, etc., and then using optical character recognition technology to recognize and convert the text in the picture into editable text), etc., but not limited to this.
[0042] In another possible implementation, the electronic device can receive the text content input by the user and perform intent recognition on the text content; if the machine learning model construction intent is recognized, a machine learning model construction template is output. The machine learning model construction template includes at least a first control for inputting modeling requirements and a second control for inputting a training data set, so that the user can input modeling requirements and a training data set. Furthermore, the electronic device can receive the modeling requirements and the training data set input by the user through the first control and the second control respectively. By providing the machine learning model construction template in this embodiment, the user can be guided to input according to the preset format and structure, reducing the error rate of data input, facilitating the standardization of subsequent data processing, reducing the user's understanding of complex operations, being suitable for non-professional users, further lowering the operation threshold, and improving the overall usage efficiency. For example, please refer to Figure 3 , which shows a schematic diagram of the machine learning model construction template.
[0043] Exemplarily, the electronic device can also output modeling prompt information, which is used to prompt the user for relevant content that must be input, such as modeling requirements, the content included in the training data set, etc. For example, in the case of supervised learning, the training data set needs to include labels for implementing supervision.
[0044] In S202, a modeling configuration prompt is generated based on the modeling requirements and the training data set, and the modeling configuration prompt is input into the pre-trained language model to generate modeling configuration information for describing the configuration content of each link in the automated modeling process by using the pre-trained language model.
[0045] Among them, the pre-trained language model, such as the large language model (LLM), refers to an artificial intelligence model based on deep learning technology, especially trained using a large corpus, aiming to understand and generate text similar to human language, and has powerful natural language understanding and generation capabilities. The goal of the LLM is to achieve various applications through natural language processing capabilities, such as text generation, translation, summarization, question answering, and dialogue systems, so as to help improve the efficiency of human-computer interaction and the degree of automation.
[0046] This step uses the pre-trained language model to achieve automated generation of modeling configuration, greatly shortening the time required for traditional manual configuration, reducing the errors that may be brought by manual settings, and ensuring the rationality and consistency of the configuration.
[0047] Exemplarily, the modeling configuration information includes but is not limited to at least one of the following:
[0048] (1)Selected model algorithm: Refers to the type of algorithm used for training and constructing the model. In machine learning, there are many types of model algorithms, including but not limited to regression models such as linear regression and logistic regression; classification models such as support vector machines, decision trees, and K-nearest neighbors; neural network models such as convolutional neural networks and recurrent neural networks; clustering algorithms such as K-means and hierarchical clustering; reinforcement learning algorithms such as Q-learning and Deep Q Network (DQN).
[0049] (2)Training sample construction strategy: Involves how to select and construct the data used for training the model. The quality and quantity of the samples directly affect the performance of the model. The training sample construction strategies include but not limited to: a. Dataset division, usually including the division of the training set, validation set, and test set. b. Data augmentation: Under limited data, methods are adopted to increase the diversity of the data (such as image rotation, flipping, etc.). c. Balanced dataset: If the data categories are unbalanced, the number of samples in different categories can be balanced by undersampling or oversampling, etc. d. Data preprocessing: Such as data cleaning, normalization, missing value handling, etc.; e. Feature selection and extraction: Select the most useful input features from the training dataset, or extract important input features through certain methods (such as principal component analysis).
[0050] (3)Selected hyperparameters: Refer to the parameters set before model training. These parameters usually cannot be learned from the data and need to be set or tuned manually. They include but not limited to: a. Learning rate, which refers to the step size controlling the gradient descent process. b. Regularization parameters such as L1 regularization and L2 regularization, which are used to control the model complexity and prevent overfitting. c. Batch size, which refers to the number of training samples used in each parameter update. d. Tree depth (in the decision tree model), which controls the maximum depth of the tree to prevent overfitting.
[0051] (4)Hyperparameter tuning methods, including but not limited to: a. Grid search, which searches for different combinations of hyperparameters through an exhaustive method. b. Random search, which randomly selects hyperparameters for training, avoiding exhausting all combinations and saving computational resources. c. Bayesian optimization, which predicts which hyperparameters may give better results through a probability model, thereby reducing the search space. d. Genetic algorithms, which simulate the process of natural selection to optimize hyperparameters.
[0052] (5) Training parameter settings refer to other parameters that need to be set during model training and affect the model learning process. These include, but are not limited to: a. Optimization algorithms, such as gradient descent, Adam (Adaptive Moment Estimation), Adagrad (Adaptive Gradient Algorithm), etc. b. Loss functions, such as mean squared error, cross-entropy loss, etc., which are used to measure the gap between the model's predictions and the true values. c. Number of training epochs, which refers to the number of times the entire dataset is passed through the model for training. d. Early stopping strategy. If the performance on the validation set does not improve over several consecutive training epochs, training can be stopped early to prevent overfitting.
[0053] (6) Evaluation metrics for model performance are used to evaluate the effectiveness of the model after training. These include, but are not limited to: a. Accuracy, which is the proportion of correctly predicted samples to the total number of samples in a classification problem. b. Precision, which is the proportion of truly positive samples among all samples predicted as positive. c. Recall, which is the proportion of samples correctly predicted as positive among all actually positive samples. d. F1-score, which is the harmonic mean of precision and recall and is used when dealing with imbalanced data. e. AUC-ROC curve, which is used to evaluate the classification ability of a binary classification model.
[0054] (7) Model validation methods are used to evaluate and validate the generalization ability of the model. Validation methods include, but are not limited to: a. Cross-validation, where the dataset is divided into K subsets, and K training and tests are performed. Each time, a different subset is used as the test set, and the rest are used as the training set. Finally, the results are averaged. b. Leave-One-Out Cross-Validation (LOOCV), where each time only one sample is used as the test set, and the others are used as the training set, which is suitable for cases with a small dataset. c. Division of the training set and test set: The dataset is simply divided into a training set and a test set, usually in a ratio of 70% for training and 30% for testing.
[0055] In a possible implementation, the modeling configuration prompt can be directly generated based on the modeling requirements and the training dataset, so that the pre-trained language model generates the modeling configuration information according to the modeling configuration prompt.
[0056] For example, the modeling configuration prompt can be expressed in the following form:
[0057]
[0058] In another possible implementation, to further improve the accuracy of the modeling configuration information, a knowledge base containing several modeling examples can be pre-set in the electronic device. Different modeling examples in the knowledge base are used to describe the construction process of machine learning models for different tasks and / or different data scales. Each modeling example includes information such as detailed task descriptions, data characteristics, model selection, parameter settings, training processes, etc., and is classified according to task types and data scales. Task types include, but are not limited to: classification tasks (such as binary classification, multi-class classification), regression tasks (such as linear regression, non-linear regression), clustering tasks (such as K-means, hierarchical clustering), recommendation systems, time series prediction (such as LSTM networks), and reinforcement learning, etc. Data scales include, but are not limited to: (1) small data sets (such as less than 10,000 samples), which are suitable for data sets with fewer samples, and may use simple models or adopt strategies such as data augmentation. (2) medium-scale data sets (such as between 10,000 and 100,000 samples), which are suitable for most real-world problems, and more complex models and moderate hyperparameter tuning can be considered. (3) large data sets (such as more than 100,000 samples, even reaching the million or above level), for problems with extremely large amounts of data, efficient algorithms (such as deep learning) and distributed training may be required.
[0059] Then, during the actual application process, please refer to Figure 4 , the electronic device can recall at least one target modeling example from the pre-set knowledge base containing several modeling examples based on the modeling requirements and the training data set, and then generate a modeling configuration prompt based on the modeling requirements, the training data set, and at least one target modeling example, so that the pre-trained language model generates modeling configuration information according to the modeling configuration prompt; wherein, at least one target modeling example is used to provide a reference for generating the modeling configuration information by the pre-trained language model. In this embodiment, by recalling the pre-defined target modeling example and adding the target modeling example to the modeling configuration prompt, the pre-trained language model can fully draw on past successful modeling experiences to determine more accurate modeling configuration information, improve the intelligent level of the modeling process, accelerate the execution efficiency of automated modeling, and also reduce the subsequent adjustment costs caused by configuration mismatches.
[0060] Exemplarily, the similarity between the embedding vector obtained by converting the modeling requirements and the training data set and the embedding vector obtained by converting the target modeling experience satisfies a preset condition. For example, the preset condition is that the similarity is greater than or equal to 90%, but it is not limited to this, and specific settings can be made according to the actual application scenario. It can ensure that the target modeling experience recalled from the knowledge base highly matches the actual modeling requirements and the training data set, thereby improving the relevance and accuracy of the subsequent generated configuration.
[0061] Exemplarily, the task of the machine learning model in the target modeling example is the same as the task described in the modeling requirements, ensuring that the task of the selected modeling example is consistent with the actual modeling requirements, making the subsequent generated configuration targeted and reducing the risk brought by task mismatch.
[0062] Exemplarily, the difference between the data scale of the target modeling example and the data volume of the training data set is less than a preset difference, and the preset difference can be specifically set according to the actual application scenario, and this embodiment does not impose any restrictions on this. By making the data scale in the target modeling example similar to the data scale of the actual training data set, the practicability of the generated configuration is improved.
[0063] In S203, based on the training data set and the modeling configuration information, control the automated machine learning tool to execute the machine learning model construction task. The machine learning model construction task includes: constructing a machine learning model to be trained and training samples, and using the training samples to train the machine learning model to be trained to obtain a trained machine learning model.
[0064] In this step, the electronic device automatically controls the machine learning tool to execute the machine learning model construction task according to the training data set and the modeling configuration information generated in S202, realizing the full-process automation from data processing to model training, and significantly improving the modeling efficiency.
[0065] Automated machine learning (AutoML) tool refers to constructing, training, and optimizing a machine learning model in an automated way. Exemplarily, please refer to Figure 4 , in the process of the automated machine learning (AutoML) tool executing the machine learning model construction task, it makes full use of the provided modeling configuration information and training data set to execute the following processes:
[0066] (1) Data preprocessing and training sample construction: The automated machine learning tool first performs preprocessing operations such as data cleaning, normalization, missing value filling, and data augmentation on the training data set according to the training sample construction strategy in the modeling configuration information to generate qualified training samples. In addition, according to the preset sample construction strategy, the data will also be reasonably divided (such as training set, validation set, and test set) to ensure the effectiveness of subsequent training and evaluation.
[0067] (2) Model architecture construction: The automated machine learning tool automatically builds the corresponding model structure according to the selected model algorithm in the modeling configuration information, which not only involves the initialization of the network structure or algorithm process, but also sets the initial hyperparameters and training parameters (such as learning rate, batch size, number of iterations, etc.) according to the requirements to construct the machine learning model to be trained.
[0068] (3) Hyperparameter Tuning and Training Process: During the model training process, the automated machine learning tool automatically adjusts the model parameters using the hyperparameters and tuning methods (such as grid search, random search, or Bayesian optimization) selected in the modeling configuration information. During continuous training and validation, the automated machine learning tool detects the model's performance in real-time based on model performance evaluation metrics (such as accuracy, recall, F1-score, etc.), and uses a predetermined model validation method (such as cross-validation) to ensure the generalization ability of the model, thereby iteratively optimizing the model parameter combination.
[0069] (4) Model Output and Report Generation: When the model training and validation reach the expected results, the automated machine learning tool outputs the trained machine learning model. At the same time, all the data, parameter settings, and performance evaluation results of the entire training process are integrated into a detailed training report, providing a basis for subsequent model deployment and further optimization.
[0070] Through this series of automated processes, the automated machine learning tool realizes the full-process automation from data preprocessing, sample construction, model architecture design, hyperparameter tuning to model validation, which not only significantly improves the modeling efficiency and model quality, but also greatly reduces the dependence on manual intervention and professional knowledge, enhancing the overall user experience.
[0071] In some embodiments, to further accelerate the model training efficiency, the automated machine learning tool can pre-store a pre-trained base model, which is trained by the provider of the machine learning model construction service using the massive data it has collected, and the base model has learned rich basic knowledge. Then, during the process of the automated machine learning tool performing the machine learning model construction task, this base model can be fully utilized to accelerate the model training efficiency.
[0072] Exemplarily, during the process of the automated machine learning tool performing the machine learning model construction task, it can obtain the input features of the base model from the training dataset, input the input features into the base model. After being trained with a large amount of data, the base model can automatically extract high-level feature representations, thereby generating an output result. Based on the training dataset and the output result, training samples for the machine learning model to be trained are generated; among them, the training samples include the input features and labels of the machine learning model to be trained, and the output result is used as one of the input features of the machine learning model to be trained. That is to say, the generated training samples not only contain the input features and corresponding labels determined based on the training dataset, but also add the output result of the base model as an additional input feature to enhance the expressive ability of the sample data. The process of extracting training samples from the training dataset and other processes can refer to the above description and will not be elaborated here.
[0073] In this embodiment, the base model has learned rich feature representations from a large amount of data, and can quickly extract high-level information from the data, reducing the number of iterations required for subsequent model training, thereby significantly improving the training efficiency. Using the output of the base model as supplementary features can provide more comprehensive and abstract feature information for the model to be trained, which helps to improve the generalization ability and prediction accuracy of the model.
[0074] It can be understood that if the base model is used to accelerate the construction efficiency of the machine learning model, after the machine learning model is trained, the trained machine learning model is used in conjunction with the base model. That is to say, in the actual application process, for a prediction requirement, the input features of the base model need to be obtained from the input data corresponding to the prediction requirement, and the input features are input into the base model to obtain the output result. Then, the input features of the machine learning model are obtained from the input data corresponding to the prediction requirement, and the input features of the machine learning model and the output result of the base model are input into the machine learning model for prediction.
[0075] In some embodiments, during the process of the automated machine learning tool performing the machine learning model construction task, the electronic device can query the execution progress of the automated machine learning tool and feedback the execution progress to the user. For example, the execution progress information can be displayed on the display interface.
[0076] In some embodiments, the electronic device can obtain the trained machine learning model and its training report returned by the automated machine learning tool after completing the machine learning model construction task. The training report is at least used to record the configuration parameters and performance evaluation results of the trained machine learning model. Exemplarily, the training report includes but is not limited to: (1) Configuration parameters: Record the adopted model architecture, hyperparameter settings, training sample construction strategy, and other training parameters. (2) Training process data: Show the changes of key indicators during the training process, such as loss value, accuracy, recall rate, F1 score, etc., as well as dynamic data such as learning curves. (3) Performance evaluation results: Provide the performance evaluation data of the model on the validation set and the test set to help judge the generalization ability of the model. (4) Abnormal information: Point out the abnormal situations and performance bottlenecks that occur during the training process, providing a reference for subsequent model tuning.
[0077] Please refer to Figure 5, the electronic device can generate model evaluation prompts based on the modeling requirements and training reports, and input the model evaluation prompts into the pre-trained language model to use the pre-trained language model to evaluate whether the trained machine learning model meets the modeling requirements. For example, the performance evaluation results in the training report can be compared with the performance requirements in the modeling requirements to determine whether the trained model meets the preset modeling requirements. If the pre-trained language model reaches a non-compliance conclusion, the pre-trained language model can be further used to modify the modeling configuration information corresponding to the trained machine learning model. This embodiment realizes the closed-loop feedback of model construction, can timely discover and correct the problem of insufficient model performance, automatically generates modification suggestions through the pre-trained language model, greatly reduces the dependence on professionals, and improves the overall modeling efficiency.
[0078] Exemplarily, the model evaluation prompt further includes at least one modification example, and different modification examples are used to describe the modification strategies taken for various unexpected situations of the machine learning model, such as performance metrics, generalization ability, data bias, and abnormal behavior, so as to provide a reference for the pre-trained language model to modify the modeling configuration information corresponding to the trained machine learning model, so that the pre-trained language model can accurately adjust the modeling configuration information.
[0079] For example, the modification examples include but are not limited to: (1) When the accuracy of the model on the validation set is lower than, it is recommended to adjust the hyperparameters (such as increasing the learning rate or optimizing the regularization parameter) and improve the feature engineering strategy, such as introducing more discriminative features, to improve the accuracy. (2) When there is an overfitting phenomenon in the model (the performance difference between the training set and the validation set is large), it is recommended to introduce or strengthen the regularization measure (such as introducing L2 regularization) to enhance the generalization ability of the model. (3) If there is a data bias problem, for example, when the prediction effect of the model on certain specific data subsets is significantly lower than that of other subsets, it is recommended to resample or balance the data, or add a targeted data augmentation strategy to the modeling configuration to eliminate the data bias. These modification examples provide adjustment strategies for different unexpected situations for the pre-trained language model, so that it can more accurately locate problems and optimize when modifying the modeling configuration information.
[0080] In a possible case, please refer to Figure 5, if the output of the pre-trained language model does not conform to the conclusion and the modified modeling configuration information, the electronic device can use the training data set and the modified modeling configuration information to control the automated machine learning tool to execute the machine learning model construction task again. In this embodiment, by reusing the modified configuration information and the training data set through the automated machine learning tool to execute the machine learning model construction task, the model can be continuously corrected and improved in subsequent iterations, improving the performance and stability of the final model. Moreover, this process automatically re-executes the machine learning model construction task, reducing the dependence on manual debugging and intervention, which can shorten the development cycle and reduce the risk of human errors.
[0081] Exemplarily, in order to avoid getting into an infinite retry state, a preset retry count can be set. That is, if the output of the pre-trained language model does not conform to the conclusion and the modified modeling configuration information, and the execution count of the machine learning model construction task for the modeling requirement has not reached the preset retry count, the electronic device will use the training data set and the modified modeling configuration information to control the automated machine learning tool to execute the machine learning model construction task again. By setting the preset retry count, the fault tolerance can be enhanced when the model does not meet the expectation, and at the same time, it can also prevent the electronic device from entering an infinite retry state when facing an unsolvable error.
[0082] If the output of the pre-trained language model does not conform to the conclusion and the modified modeling configuration information, and the execution count of the machine learning model construction task for the modeling requirement has reached the preset retry count, then the electronic device can directly output the trained machine learning model and the training report.
[0083] In another possible scenario, if the output of the pre-trained language model conforms to the conclusion, the electronic device can output the trained machine learning model and the training report.
[0084] Exemplarily, please refer to Figure 6 , considering that the training report output by the automated machine learning tool usually contains a large amount of technical data and information in non-natural language formats, such as logs, matrices, numerical tables, etc., which are not easy for users to directly read and understand. Then the training report output by the automated machine learning tool can be input into the pre-trained language model first, so as to use the pre-trained language model to perform data conversion on the training report in a way that is easy for users to understand, obtaining the converted training report, and then outputting the trained machine learning model and the converted training report to the user. In this embodiment, through data conversion by the pre-trained language model, the training report can be transformed into a more intuitive and easy-to-understand natural language description, enabling non-professional users to easily obtain key information. Without the need for users to have a deep machine learning background, they can understand the training process, key performance indicators, and potential problems of the model through the converted training report, which helps non-technical users such as relevant process personnel and product managers to participate in the modeling decision-making.
[0085] The various technical features in the above embodiments can be combined arbitrarily as long as there is no conflict or contradiction between the features. However, due to space limitations, not all combinations are described one by one. Therefore, any combination of the various technical features in the above embodiments also falls within the scope disclosed in this specification.
[0086] In some embodiments, the embodiments of this specification also provide an electronic device, including: a processor; a memory for storing executable instructions that can be executed by the processor; wherein, the processor realizes the method described in any one of the above by running the executable instructions.
[0087] Figure 7 is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 7 , at the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710. Of course, there may also be other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 702 reads the corresponding computer program from the non-volatile memory 710 into the memory 708 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of this specification do not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0088] In some embodiments, the device for constructing a machine learning model can be applied to a device as shown in Figure 7 to implement the technical solutions of this specification. Among them, the device for constructing a machine learning model can include:
[0089] An information receiving module, configured to receive the modeling requirements input by the user and a training data set for generating training samples, where the modeling requirements are at least used to describe the tasks and performance requirements of the machine learning model to be constructed;
[0090] A modeling configuration information obtaining module, configured to generate a modeling configuration prompt based on the modeling requirements and the training data set, and input the modeling configuration prompt into a pre-trained language model to use the pre-trained language model to generate modeling configuration information for describing the configuration content of each link in the automated modeling process;
[0091] A machine learning model construction module, configured to control an automated machine learning tool to perform a machine learning model construction task based on the training data set and the modeling configuration information. The machine learning model construction task includes: constructing a machine learning model to be trained and training samples, and training the machine learning model to be trained with the training samples to obtain a trained machine learning model.
[0092] Exemplarily, the modeling configuration information acquisition module is specifically configured to recall at least one target modeling example from a knowledge base pre-set with a number of modeling examples based on the modeling requirements and the training data set. Different modeling examples are used to describe the construction processes of machine learning models for different tasks and / or different data scales; generate a modeling configuration prompt based on the modeling requirements, the training data set, and the at least one target modeling example; wherein the at least one target modeling example is used as a reference for generating the modeling configuration information for the pre-trained language model.
[0093] Exemplarily, the similarity between the embedding vector obtained by converting the modeling requirements and the training data set and the embedding vector obtained by converting the target modeling experience satisfies a preset condition.
[0094] Exemplarily, the task of the machine learning model in the target modeling example is the same as the task described in the modeling requirements.
[0095] Exemplarily, the difference between the data scale in the target modeling example and the data volume of the training data set is less than a preset difference.
[0096] Exemplarily, the modeling configuration information includes at least one of the following: a selected model algorithm, a training sample construction strategy, selected hyperparameters, a hyperparameter tuning method, training parameter settings, an evaluation metric for model performance, and a model verification method.
[0097] Exemplarily, the apparatus further includes an evaluation module, a retry module, and an output module.
[0098] The evaluation module is configured to obtain the trained machine learning model and its training report returned by the automated machine learning tool after completing the machine learning model construction task. The training report is at least used to record the configuration parameters and performance evaluation results of the trained machine learning model; generate a model evaluation prompt based on the modeling requirements and the training report, and input the model evaluation prompt into the input pre-trained language model to use the pre-trained language model to evaluate whether the trained machine learning model meets the modeling requirements and modify the corresponding modeling configuration information of the trained machine learning model when a non-conformance conclusion is drawn.
[0099] The retry module is used to, if the output of the pre-trained language model does not conform to the conclusion and the modified modeling configuration information, use the training dataset and the modified modeling configuration information to control the automated machine learning tool to execute the machine learning model construction task again.
[0100] The output module is used to, if the output of the pre-trained language model conforms to the conclusion, output the trained machine learning model and the training report.
[0101] Exemplarily, the model evaluation prompt further includes at least one modification example, and different modification examples are used to describe the modification strategies adopted for the machine learning model in various situations that do not meet expectations, providing a reference for the pre-trained language model to modify the modeling configuration information corresponding to the trained machine learning model.
[0102] Exemplarily, the retry module is specifically used to, if the output of the pre-trained language model does not conform to the conclusion and the modified modeling configuration information, and the number of executions of the machine learning model construction task for the modeling requirement has not reached the preset retry times, use the training dataset and the modified modeling configuration information to control the automated machine learning tool to execute the machine learning model construction task again. The output module is further used to, otherwise, output the trained machine learning model and the training report.
[0103] Exemplarily, the output module is specifically used to input the training report into the pre-trained language model, so as to use the pre-trained language model to perform data conversion on the training report in a manner easy for users to understand, and obtain the converted training report; output the trained machine learning model and the converted training report.
[0104] Exemplarily, the information receiving module is specifically used to receive the text content input by the user, and perform intent recognition on the text content; if the machine learning model construction intent is recognized, output a machine learning model construction template, and the machine learning model construction template at least includes a first control for inputting modeling requirements and a second control for inputting a training dataset; receive the modeling requirements and the training dataset input by the user through the first control and the second control respectively.
[0105] Exemplarily, the automated machine learning tool prestores a pre-trained base model; the machine learning model construction task further includes: obtaining the input features of the base model from the training dataset, inputting the input features into the base model to obtain an output result, and generating training samples for the machine learning model to be trained based on the training dataset and the output result; wherein, the training samples include the input features and labels of the machine learning model to be trained, and the output result is used as one of the input features of the machine learning model to be trained.
[0106] For the implementation processes of the functions and roles of each module in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0107] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor realizes the steps of the method as described in any of the above embodiments by running the executable instructions.
[0108] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any of the above embodiments are realized.
[0109] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0110] Based on the same concept as the above method, this specification also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method as described in any of the above embodiments are realized.
[0111] The above are only the preferred embodiments of one or more embodiments of this specification, and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.
Claims
1. A method for constructing a machine learning model, comprising: Receive a modeling requirement input by a user and a training data set for generating training samples, wherein the modeling requirement is at least used to describe the task and performance requirements of the machine learning model to be constructed; Generate a modeling configuration prompt based on the modeling requirement and the training data set, and input the modeling configuration prompt into a pre-trained language model to generate modeling configuration information for describing the configuration content of each link in the automated modeling process using the pre-trained language model; Based on the training data set and the modeling configuration information, the automated machine learning tool is controlled to perform a machine learning model building task, wherein the machine learning model building task includes: building a machine learning model to be trained and training samples, and using the training samples to train the machine learning model to be trained to obtain a trained machine learning model.
2. The method according to claim 1, wherein generating a modeling configuration prompt based on the modeling requirement and the training data set comprises: Based on the modeling requirement and the training data set, recall at least one target modeling example from a preset knowledge base containing a plurality of modeling examples, where different modeling examples are used to describe the construction process of machine learning models for different tasks and / or different data scales; A modeling configuration prompt is generated based on the modeling requirement, the training data set, and the at least one target modeling example; wherein the at least one target modeling example is used to provide a reference for generating the modeling configuration information for the pre-trained language model.
3. According to the method of claim 2, the similarity between the embedding vector converted from the modeling requirement and the training data set and the embedding vector converted from the target modeling experience meets a preset condition; The task of the machine learning model in the target modeling example is the same as the task described in the modeling requirement; And / or, the difference between the data scale in the target modeling example and the data volume of the training data set is less than a preset difference.
4. According to the method of claim 1, the modeling configuration information includes at least one of the following: a selected model algorithm, a training sample construction strategy, a selected hyperparameter, a hyperparameter tuning method, a training parameter setting, an evaluation index of model performance, and a model verification method.
5. The method according to claim 1, further comprising: Obtaining the trained machine learning model and its training report returned by the automated machine learning tool after completing the machine learning model building task, wherein the training report is at least used to record configuration parameters and performance evaluation results of the trained machine learning model; Generate a model evaluation prompt based on the modeling requirement and the training report, input the model evaluation prompt into the input pre-trained language model, so as to use the pre-trained language model to evaluate whether the trained machine learning model meets the modeling requirement, and modify the modeling configuration information corresponding to the trained machine learning model when a conclusion is drawn that it does not meet the modeling requirement; If the output of the pre-trained language model does not conform to the conclusion and the modified modeling configuration information, using the training data set and the modified modeling configuration information to control the automated machine learning tool to perform the machine learning model building task again; If the output of the pre-trained language model meets the conclusion, the trained machine learning model and the training report are output.
6. According to the method described in claim 5, the model evaluation prompt also includes at least one modification example, and different modification examples are used to describe the modification strategies adopted for the machine learning model under various unexpected situations, so as to provide a reference for the pre-trained language model to modify the modeling configuration information corresponding to the trained machine learning model.
7. The method according to claim 5, wherein if the output of the pre-trained language model does not conform to the conclusion and the modified modeling configuration information, using the training data set and the modified modeling configuration information to control the automated machine learning tool to perform the machine learning model building task again, comprises: If the output of the pre-trained language model does not conform to the conclusion and the modified modeling configuration information, and the number of executions of the machine learning model building task for the modeling requirement does not reach the preset number of retries, the training data set and the modified modeling configuration information are used to control the automated machine learning tool to execute the machine learning model building task again; Otherwise, the trained machine learning model and the training report are output.
8. According to the method of claim 5 or 7, the outputting of the trained machine learning model and the training report comprises: Inputting the training report into the pre-trained language model, so as to use the pre-trained language model to perform data conversion on the training report in a manner that is easy for a user to understand, thereby obtaining a converted training report; Output the trained machine learning model and the converted training report.
9. The method according to claim 1, wherein receiving the modeling requirements input by the user and the training data set for generating training samples comprises: Receive text content input by a user and perform intent recognition on the text content; If the machine learning model building intention is identified, output a machine learning model building template, wherein the machine learning model building template includes at least a first control for inputting modeling requirements and a second control for inputting a training data set; The modeling requirement and the training data set are received, respectively input by the user through the first control and the second control.
10. The method according to claim 1, wherein the automated machine learning tool has a pre-trained basic model pre-stored; The machine learning model building task also includes: Acquire input features of the basic model from the training data set, input the input features into the basic model to obtain output results, and generate training samples of the machine learning model to be trained based on the training data set and the output results; wherein the training samples include input features and labels of the machine learning model to be trained, and the output results serve as one of the input features of the machine learning model to be trained.
11. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 10 by executing the executable instructions.
12. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Cited By
LLM-based machine learning model training method, system and device, medium and program
CN121413811A
Model training method and device, computer equipment, storage medium and program product
CN121459092A