Method and device for model training based on large model, and intelligent agent

Generate model training components by receiving natural language inputs through large models, solving the problem of high user technical capabilities in the existing technology, and achieving the ease of use and efficiency of model training.

CN120338101APending Publication Date: 2025-07-18BEIJING BAIDU NETCOM SCI & TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400027.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the model training process has high requirements for users' technical capabilities, and users need to understand and provide complex technical details in detail, resulting in a high training threshold.

Method used

Through the large model, we receive user input in natural language form, understand user requirements and generate corresponding model training components, including data set loading, preprocessing, model architecture and training parameters, etc., simplifying the user interaction process.

Benefits of technology

This reduces the technical capability requirements for users to train models. Users only need to complete model training and prediction through short natural language interactions, which improves the ease of use and efficiency of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338101A_ABST
    Figure CN120338101A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for model training based on a large model, and an intelligent agent, and relates to the technical field of artificial intelligence, in particular to a large model technology and a model training technology. According to the implementation scheme, user input in a natural language form for model training is received; processing the user input by using the large model to obtain a user requirement for model training included in the user input; and generating a model training component corresponding to the user requirement for model training by using the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, and in particular, to large model technologies and model training technologies. Specifically, the present disclosure relates to a method, apparatus, electronic device, computer-readable storage medium, computer program product, and agent for model training based on a large model. Background Art

[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for model training based on a large model.

[0005] According to one aspect of the present disclosure, there is provided a method for model training based on a large model, including: receiving a user input in natural language form for model training; processing the user input by using the large model to obtain a user requirement for model training included in the user input; and generating a model training component corresponding to the user requirement for model training by using the large model.

[0006] According to another aspect of the present disclosure, there is provided an apparatus for model training based on a large model, including: a receiving unit configured to receive a user input in natural language form for model training; a user requirement obtaining unit configured to process the user input by using the large model to obtain a user requirement for model training included in the user input; and a generating unit configured to generate a model training component corresponding to the user requirement for model training by using the large model.

[0007] According to another aspect of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to the embodiments of the present disclosure.

[0008] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to the embodiments of the present disclosure.

[0009] According to another aspect of the present disclosure, there is also provided a computer program product including a computer program, wherein when the computer program is executed by a processor, it implements the method according to the embodiments of the present disclosure.

[0010] According to another aspect of the present disclosure, there is also provided an agent, including: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a generation model based on the target task, and executing the method according to the embodiments of the present disclosure by invoking the generation model to obtain a model training component for model training generated for the input information; and an output module for outputting the model training component generated by the processing module.

[0011] According to one or more embodiments of the present disclosure, the requirements for the technical capabilities of users can be reduced when training a model and using the model for prediction.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings exemplarily illustrate embodiments and constitute a part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The illustrated embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein can be implemented according to an embodiment of the present disclosure is shown;

[0015] Figure 2 An exemplary flowchart of a method for model training according to an embodiment of the present disclosure is shown;

[0016] Figure 3 Shows an exemplary block diagram of an artificial intelligence development visualization platform according to an embodiment of the present disclosure;

[0017] Figure 4 Shows an exemplary block diagram of a large model-based device for model training according to an embodiment of the present disclosure;

[0018] Figure 5 Shows a structural block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure. Detailed Description of the Embodiments

[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0020] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0021] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0022] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0023] Figure 1 Shows a schematic diagram of an exemplary system 100 in which various methods and devices described herein can be implemented according to an embodiment of the present disclosure. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.

[0024] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable the execution of a method for model training according to an embodiment of the present disclosure.

[0025] In some embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software-as-a-service (SaaS) model.

[0026] In Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0027] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to interact with the large model-based agent related to the present disclosure. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.

[0028] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.

[0029] Network 110 may be any type of network known to those skilled in the art, which may use any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0030] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.

[0031] The computing units in server 120 can run one or more operating systems including any of the above - mentioned operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0032] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 can also include one or more applications to display data feeds and / or real - time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.

[0033] In some embodiments, server 120 can be a server of a distributed system or a server incorporating a blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, which addresses the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0034] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network - based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be relational databases, for example. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.

[0035] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key - value repositories, object repositories, or conventional repositories supported by a file system.

[0036] Figure 1The system 100 can be configured and operated in various ways to enable the application of various methods and apparatuses described in this disclosure.

[0037] Figure 2 FIG. shows an exemplary flowchart of a method for model training according to an embodiment of the present disclosure.

[0038] In step S202, a user input in natural language form for model training is received.

[0039] In step S204, the large model is used to process the user input to obtain the user requirements for model training included in the user input.

[0040] In step S206, the large model is used to generate model training components corresponding to the user requirements for model training.

[0041] Using the method for model training provided by the embodiments of the present disclosure can reduce the requirements for the technical capabilities of users when training models and using models for prediction. By using the trained large model to understand the natural language information of the user input and generate model training components corresponding to the user requirements included in the user input information, the user can provide the requirements for model training to the large model in the form of natural language, and the large model can directly generate each component required for model training for the user through its own semantic understanding and generation capabilities.

[0042] The principle of the present disclosure will be described in detail below.

[0043] In step S202, a user input in natural language form for model training can be received.

[0044] The user can input the requirements for model training to the large model in a natural language interaction manner. In some embodiments, the user can interact with an intelligent agent, and the intelligent agent can call the large model for generating model training components to implement method 200. In some examples, the intelligent agent interacting with the user can be implemented in the form of a model training assistant. Such an intelligent agent can be embedded in software programs such as a computer operating system and a model training system for the user to call. During this process, the user can explain the required information through one or more inputs.

[0045] In step S204, the large model can be used to process the user input to obtain the user requirements for model training included in the user input.

[0046] The large model used in the embodiments of the present disclosure may be a large language model or any form of large model that supports multi-modal content (such as text, image, video, audio) input / output. The pre-trained large model has the ability to solve general problems. Based on the pre-trained large model, the large model can be further fine-tuned by instructions, enabling the large model to learn the ability to process user requests at various stages of model training. In some embodiments, the large model can be processed by prompt engineering, that is, relevant prompt tuning is performed according to different stages of model training, and then the large model learns how to process the needs proposed by users at various stages of model training. Exemplary user inputs at different stages can be provided to the large model, and the large model can be guided by instructions on how to understand user inputs for different training stages and how to generate appropriate response results. Without departing from the embodiments of the present disclosure, various tuning methods can be used to implement the prompt tuning involved in the embodiments of the present disclosure.

[0047] In the training set loading stage, the large model can be trained to understand what dataset the user is using, where the dataset is obtained, where to download the publicly available dataset if needed, where the downloaded dataset should be saved, and how to load the dataset, etc. In the example, since the pre-trained large model has the ability to solve general problems, it is not necessary to require the user to provide all information related to dataset loading in the user input.

[0048] Exemplary user inputs can be: "I need to train an image classification model, and the dataset is CIFAR-10"; or "I need to train an image classification model, and the dataset includes data in the medical field". Using the above inputs, the user can describe the dataset for model training in a natural language manner without giving specific information about the dataset type and acquisition method. The large model can use its own semantic understanding ability and general knowledge to determine the specific name and / or acquisition method of the dataset required by the user.

[0049] In the training dataset preprocessing stage, the large model can be trained to understand what algorithm the user uses to preprocess the dataset, how to configure the parameters between algorithms, and recommend settings for the user.

[0050] Exemplary user inputs can be: "Please preprocess the CIFAR-10 dataset, including image normalization (scaling pixel values to the range [0,1]) and data augmentation (random horizontal flipping and random cropping)". In some examples, the description "scaling pixel values to the range [0,1]" or "random horizontal flipping and random cropping" can also be omitted, and the large model can use its own semantic understanding ability and general knowledge to determine the specific preprocessing means of the dataset.

[0051] In the model training stage, a large model can be trained to enable it to understand what kind of model the user is using, the corresponding loss function, etc., as well as the use and recommendation of relevant configuration parameters. If the user does not use a publicly available and mature model architecture, the large model can generate the code corresponding to the model after understanding the model architecture described by the user.

[0052] Exemplary user input can be: Please use the ResNet-18 model for training and initialize appropriate training parameters (such as setting the learning rate to 0.001 and using the Adam optimizer). After training is completed, save the model to the / models / resnet18 directory. An alternative user input can be: Based on ResNet-18, streamline the convolutional neural network structure and initialize appropriate training parameters (such as setting the learning rate to 0.001 and using the Adam optimizer). After training is completed, save the model to the / models / cifar10_resnet18 directory.

[0053] In step S206, a large model can be used to generate a model training component corresponding to the user requirements for model training.

[0054] In some embodiments, the model training component can be a visualization component. The large model can generate the code for generating the visualization component. The user can generate the corresponding visualization component by running the code in the software dedicated to model training. The large model can also generate any other information required for generating the visualization component. In the embodiments of the present disclosure, by using the language generation ability of the large model to generate the information for implementing the model training component according to the identified user requirements, the effect of directly generating the model training component by the large model in response to user input can be achieved, without having to call other auxiliary programs (such as search) to generate the model training component. Thus, the model training assistant based on the large model can be simplified and the resources required by the model training assistant can be reduced.

[0055] In some embodiments, the model training component may include a dataset loading component, a dataset augmentation component, a model training component, and a model output component. When the model training component is a visualization component, the user can drag and configure the visualization component on the visualization interface to complete the construction of the overall training link. In some implementation manners, when the large model knows that the task the user wants to achieve is model training, the large model can guide the user to give user requirements one by one according to each stage of model training by means of natural language interaction with the user. The large model will generate corresponding visualization components for each stage according to the user requirements input in the form of natural language and connect the visualization components of each stage together according to the training program. In the example, if the content input by the user misses the necessary information for generating the model training component, the large model can output guiding information to the user by means of natural language interaction to prompt the user to input the corresponding necessary information.

[0056] Step S206 may include determining the type and acquisition method of the dataset for model training, acquiring the dataset according to the acquisition method, and generating a corresponding dataset loading component based on the loading method of the acquired dataset. In some embodiments, the user can describe to the large model in natural language what the dataset to be trained is. If the required dataset needs to be obtained by downloading, the user can also indicate to the large model where to store the dataset after downloading. After understanding the user's intention, the large model will create a corresponding dataset loading component.

[0057] Step S206 may include determining the preprocessing parameters of the dataset for model training and generating a corresponding dataset augmentation component based on the preprocessing parameters. In some embodiments, the user can describe to the large model the algorithms for dataset preprocessing and data augmentation. If the required preprocessing algorithm is not pre-configured in the model training program, the large model can automatically generate the corresponding algorithm according to the user requirements and create a corresponding dataset augmentation component for dataset preprocessing.

[0058] Step S206 may include determining the model architecture and model parameters for model training and generating a corresponding model training component based on the model architecture and model parameters. In some embodiments, the user can describe to the large model the model structure or name to be used for training. If the required model structure is an existing structure already configured in the model training software, it can be directly called. If the required model structure is not included in the existing structures, the large model can automatically create the corresponding model structure according to the understood user intention. Further, after determining the model structure, the large model can initialize the training parameters of the model to be trained according to the training experience knowledge it has.

[0059] Step S206 may include determining the model output information of the trained model and generating corresponding model output components based on the model output information. In some embodiments, the model output information may include a storage location. The large model may create model output components according to the storage location specified by the user or determined according to a predetermined rule, and store the trained model in the above storage location.

[0060] In some embodiments, method 200 may further include: receiving a user query in natural language form for model prediction, processing the user query using the large model to obtain the user requirements for model prediction included in the user query, and generating, using the large model, a prediction component corresponding to the user requirements for model prediction.

[0061] After the link for model training is set up, when the user needs to use the trained model, the user may input a user query in natural language form to the large model. In the example, the content of the user query may include information specifying the data for prediction, information about the model for prediction, and information about the storage location of the prediction result. The large model may generate a prediction component according to the user intention identified from the user query for prediction using the trained model. In the example, the prediction component is also a visualization component.

[0062] Using the embodiments of the present disclosure can lower the threshold for user training and prediction, and only a short language interaction with the large model is required to complete the model training and use the trained model for data prediction.

[0063] Figure 3 An exemplary block diagram of an artificial intelligence development visualization platform according to an embodiment of the present disclosure is shown.

[0064] As Figure 3 shown, platform 300 may include a dataset loading component 310, a dataset augmentation component 320, a model training component 330, a model output component 340, and a prediction component 350.

[0065] The user may perform natural language interaction with a model training assistant based on a large model to configure the above model training components on the visualization platform 300.

[0066] The following details an exemplary process of the user using the model training assistant for model training assistance.

[0067] Step 1: The user may input to the model training assistant "I need to use the voc12 dataset for training, and you help me download it to this address: [specified address]", where the specified address within the square brackets may be specified by the user according to the actual situation.

[0068] In response to Step 1, the model training assistant can perform the following actions: download the voc12 dataset, create a training visualization job A, and create and configure a dataset loading component 310 in the created visualization job.

[0069] Further in response to Step 1, the model training assistant can output a reply to the user: "I have downloaded the voc12 dataset to the specified location for you, created a training visualization job A for you, and the dataset loading plugin has been set up."

[0070] Step 2: The user can input to the model training assistant: "I need to crop my voc dataset to a size of 512*512*3 and perform translation and rotation operations."

[0071] In response to Step 2, the model training assistant can perform the following actions: create a dataset augmentation component 320 in the visualization job according to the requirements.

[0072] Further in response to Step 2, the model training assistant can output a reply to the user: "I have created a dataset augmentation component in the visualization job according to your requirements."

[0073] Step 3: The user can input to the model training assistant: "I want to train my data using Resnet50 and store the trained model parameters at [specified location]", where the specified location within the square brackets can be specified by the user according to the actual situation.

[0074] In response to Step 3, the model training assistant can perform the following actions: create a model training component 330 and set the optimal training parameters and a model output component 340.

[0075] Further in response to Step 3, the model training assistant can output a reply to the user: "I have created a model training component for you and set the optimal training parameters and a model output component."

[0076] After the model training assistant replies to the user for Step 3, a model training link will be set up in the visualization platform 300, and the user can call and execute the already set model training link according to the actual situation.

[0077] Step 4: The user can input to the model training assistant: "I want to use the trained model to predict the data I stored at [specified location] and store the prediction results at [specified location]."

[0078] In response to Step 4, the model training assistant can perform the following actions: create a model prediction component, and the generated results will be generated at the location specified by the user after the job is successful.

[0079] Further in response to step 4, the model training assistant may output a reply to the user: "A model prediction component has been created for you, and the generated results will be generated at the location you specified after the job is successful."

[0080] After the model training assistant replies to the user for step 4 performed by the user, the user may call and execute the prediction component generated on the visualization platform and perform a prediction on the specified data.

[0081] Figure 4 An exemplary block diagram of an apparatus for model training based on a large model according to an embodiment of the present disclosure is shown.

[0082] As Figure 4 shown, the apparatus 400 includes a receiving unit 410, a user requirement acquisition unit 420, and a generating unit 430.

[0083] The receiving unit 410 may be configured to receive a user input in natural language form for model training.

[0084] The user requirement acquisition unit 420 may be configured to process the user input using a large model to obtain the user requirements for model training included in the user input.

[0085] The generating unit 430 may be configured to generate a model training component corresponding to the user requirements for model training using a large model.

[0086] In some embodiments, the model training component includes a dataset loading component, a dataset augmentation component, a model training component, and a model output component.

[0087] In some embodiments, generating a model training component corresponding to the user requirements using a large model may include: determining the type and acquisition method of the dataset for model training; acquiring the dataset according to the acquisition method; and generating a corresponding dataset loading component based on the loading method of the acquired dataset.

[0088] In some embodiments, generating a model training component corresponding to the user requirements using a large model includes: determining the preprocessing parameters of the dataset for model training; and generating a corresponding dataset augmentation component based on the preprocessing parameters.

[0089] In some embodiments, generating a model training component corresponding to the user requirements using a large model includes: determining the model architecture and model parameters for model training; and generating a corresponding model training component based on the model architecture and model parameters.

[0090] In some embodiments, generating a model training component corresponding to the user requirements using a large model includes: determining the model output information of the model to be trained; and generating a corresponding model output component based on the model output information.

[0091] In some embodiments, the apparatus 400 may further include a prediction component generation unit configured to: receive a user query in natural language form for model prediction; process the user query using a large model to obtain user requirements for model prediction included in the user query; and generate a prediction component corresponding to the user requirements for model prediction using the large model.

[0092] In some embodiments, the model training component is a visualization component.

[0093] In some embodiments, by performing instruction tuning on the large model, the large model learns the ability to process user requirements at various stages of model training.

[0094] It should be understood that Figure 4 each module or unit of the apparatus 400 shown in Figure 2 may correspond to each step in the method 200 described with reference to

[0095] Accordingly, the operations, features, and advantages described above for the method 200 also apply to the apparatus 400 and its included modules and units. For the sake of brevity, certain operations, features, and advantages are not described herein again.

[0096] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure, etc., of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0097] According to an embodiment of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to the embodiment of the present disclosure.

[0098] According to an embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to the embodiment of the present disclosure.

[0099] According to an embodiment of the present disclosure, there is also provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the method according to the embodiment of the present disclosure.

[0100] According to an embodiment of the present disclosure, an agent is further provided, including: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a generation model based on the target task, and obtaining a model training component for model training generated for the input information by calling the generation model to execute the method according to the embodiment of the present disclosure; and an output module for outputting the model training component generated by the processing module.

[0101] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0102] According to an embodiment of the present disclosure, an electronic device, a readable storage medium, and a computer program product are further provided.

[0103] Referring to Figure 5 , a block diagram of an electronic device 500 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0104] As Figure 5 shown, the electronic device 500 includes a computing unit 501, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0105] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device capable of inputting information into the electronic device 500. The input unit 506 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 507 can be any type of device capable of presenting information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 can include, but is not limited to, magnetic disks and optical discs. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0106] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of method 200 described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute method 200 in any other suitable manner (e.g., by means of firmware).

[0107] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0108] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0109] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0111] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain network.

[0112] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.

[0113] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0114] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. A method for model training based on a large model, comprising: Receiving a user input in natural language form for model training; Processing the user input using the large model to obtain user requirements for model training included in the user input; And Generating, using the large model, model training components corresponding to the user requirements for model training.

2. The method according to claim 1, wherein The model training components include a dataset loading component, a dataset augmentation component, a model training component, and a model output component.

3. The method according to claim 2, wherein, Generating, using the large model, model training components corresponding to the user requirements includes: Determining the type and acquisition method of the dataset for model training; Acquiring the dataset according to the acquisition method; and Generating a corresponding dataset loading component based on the loading method of the acquired dataset.

4. The method according to claim 2, wherein, Generating, using the large model, model training components corresponding to the user requirements includes: Determining the preprocessing parameters of the dataset for model training; and Generating a corresponding dataset augmentation component based on the preprocessing parameters.

5. The method according to claim 2, wherein Generating, using the large model, model training components corresponding to the user requirements includes: Determining the model architecture and model parameters for model training; and Generating a corresponding model training component based on the model architecture and model parameters.

6. The method according to claim 2, wherein, Generating, using the large model, model training components corresponding to the user requirements includes: Determining the model output information of the model to be trained; and Generating a corresponding model output component based on the model output information.

7. The method according to claim 1, further comprising: Receiving a user query in natural language form for model prediction; Processing the user query using the large model to obtain user requirements for model prediction included in the user query; And Generating, using the large model, a prediction component corresponding to the user requirements for model prediction.

8. The method according to any one of claims 1-7, wherein, The model training component is a visualization component.

9. The method according to any one of claims 1-7, wherein, By performing instruction tuning on the large model, the large model learns the ability to process user requirements at each stage of model training.

10. An apparatus for model training based on a large model, comprising: A receiving unit configured to receive a user input in natural language form for model training; A user requirement acquisition unit configured to process the user input using the large model to obtain user requirements for model training included in the user input; And A generating unit configured to generate, using the large model, model training components corresponding to the user requirements for model training.

11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.

13. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-9.

14. An intelligent agent, comprising: An input module for receiving input information; A processing module for determining a target task based on the input information received by the input module, determining a generation model based on the target task, and executing the method according to any one of claims 1-9 by calling the generation model to obtain a model training component for model training generated for the input information; An output module for outputting the model training component generated by the processing module.