Processing method and device, electronic equipment and storage medium
By creating a basic model interface in the terminal device, loading and switching different model parameters, and using LoRA technology for fine-tuning, the problems of low loading speed and update efficiency of large models are solved, and efficient model management and user experience optimization are achieved.
Patent Information
- Application Number
- CN202510901231.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-03
AI Technical Summary
When users load and update large local models through terminals, they have high demands on network bandwidth, resulting in low model loading speed and update efficiency.
By creating an interface for the basic model, loading the parameters of different models, and using LoRA technology for fine-tuning, flexible switching and updating of model parameters can be achieved, and model functions can be dynamically adjusted to adapt to different application scenarios.
It improves the loading speed and usage efficiency of local models, optimizes the user experience, reduces network bandwidth consumption, and ensures model performance and data security.
Smart Images

Figure CN120743377A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of large model technology, and in particular to a processing method, device, electronic device and storage medium. Background Art
[0002] With the rapid increase in the number of parameters in large language models, users have high demands on network bandwidth when loading and updating local large models through terminals, resulting in low model loading speed and update efficiency. Summary of the Invention
[0003] In view of this, the present disclosure provides a processing method, apparatus, electronic device, storage medium, and computer program.
[0004] One aspect of the present disclosure provides a processing method, including: loading model parameters of a base model; creating a first interface of the base model; loading model parameters of the first model based on the first interface so that the base model has the model parameters of the base model and the model parameters of the first model.
[0005] According to an embodiment of the present disclosure, it also includes: creating a second interface of the basic model; loading the model parameters of the second model based on the second interface so that the basic model also has the model parameters of the second model; the model parameters of the first model are different from the model parameters of the second model.
[0006] According to an embodiment of the present disclosure, it also includes: loading the model parameters of the third model, the third model has the same name as the first model, and the model parameters of the third model are different from the model parameters of the first model; using the model parameters of the third model to overwrite the model parameters of the first model; loading the model parameters of the third model based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the replaced third model.
[0007] According to an embodiment of the present disclosure, it also includes: unloading the model parameters of the first model; loading the model parameters of the second model based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the second model.
[0008] According to an embodiment of the present disclosure, it also includes: acquiring input text; determining identification information of the input text based on the input text, the identification information including at least one of a domain identification and a function identification; in response to the identification information indicating the first model and / or the second model, running a basic model including model parameters of the first model and / or model parameters of the second model.
[0009] According to an embodiment of the present disclosure, running a basic model including model parameters of a first model and / or model parameters of a second model includes: in response to identification information indicating the first model or the second model, unloading the model parameters loaded by the first interface, and loading the model parameters indicated by the identification information; or loading the model parameters of the first model using the first interface, loading the model parameters of the second model using the second interface, switching the first interface and the second interface based on the identification information, and loading the model parameters indicated by the identification information.
[0010] According to an embodiment of the present disclosure, the method also includes: obtaining input text; performing semantic processing on the input text to obtain identification information corresponding to the input text; in response to the identification information indicating a first model, obtaining model parameters of the first model through a first interface using a task scheduler; running a basic model including model parameters of a basic model and model parameters of the first model; and in response to the identification information indicating a second model, uninstalling the model parameters of the first model, and obtaining model parameters of the second model through the first interface using a task scheduler.
[0011] Another aspect of the present disclosure also provides a device, including: a processing device, including: a first loading module for loading model parameters of a basic model; a creation module for creating a first interface of the basic model; a second loading module for loading the model parameters of the first model based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the first model.
[0012] Another aspect of the present disclosure provides an electronic device, comprising: a storage device for caching model parameters; one or more processors for loading model parameters of a base model; creating a first interface for the base model; and loading the model parameters of the first model based on the first interface so that the base model has the model parameters of the base model and the model parameters of the first model.
[0013] Another aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the above method.
[0014] Another aspect of the present disclosure provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0015] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0017] Figure 1 is a schematic diagram of an exemplary system architecture to which the processing method and apparatus according to one embodiment of the present disclosure may be applied;
[0018] Figure 2 A flowchart of a processing method according to an embodiment of the present disclosure is schematically shown;
[0019] Figure 3 A schematic diagram schematically illustrates a basic model according to an embodiment of the present disclosure;
[0020] Figure 4 A flowchart schematically illustrates a processing method according to another embodiment of the present disclosure;
[0021] Figure 5 A flowchart schematically illustrates a processing method according to another embodiment of the present disclosure;
[0022] Figure 6 A flowchart schematically illustrates a processing method according to another embodiment of the present disclosure;
[0023] Figure 7 A block diagram schematically illustrates a structure of a processing device according to an embodiment of the present disclosure; and
[0024] Figure 8 A schematic block diagram of an electronic device that can be used to implement the processing method of an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0025] Embodiments of the present disclosure are described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted from the following description.
[0026] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of the data involved (including but not limited to user personal information) comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0027] Figure 1 is a schematic diagram of an exemplary system architecture to which the processing method and apparatus according to one embodiment of the present disclosure can be applied. It should be noted that, Figure 1The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0028] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0029] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104, downloading large models from server 105 or updating large model functions. Terminal devices 101, 102, and 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers.
[0030] Server 105 can be a server that provides various services, such as a cloud server (for example only) that trains and distributes large models used by users on terminal devices 101, 102, and 103. The cloud server can feed back model files of different versions or functions to the terminal device based on received user requests.
[0031] It should be noted that the processing method provided in the embodiment of the present disclosure can generally be executed by the terminal devices 101, 102, and 103. Accordingly, the processing device provided in the embodiment of the present disclosure can generally be set in the terminal devices 101, 102, and 103.
[0032] Figure 2 The flowchart of the processing method according to the embodiment of the present disclosure is schematically shown.
[0033] like Figure 2 As shown, the processing method of this embodiment includes operations S210 to S230.
[0034] In operation S210 , model parameters of a base model are loaded.
[0035] In the embodiments of the present disclosure, a base model may refer to a model pre-trained using data and suitable for general tasks or domains. For example, it may be a large language model (LLM) with certain language understanding and generation capabilities. The model parameters of the base model may refer to pretrained weights.
[0036] In operation S220 , a first interface of the base model is created.
[0037] The first interface can be an interface for the base model to mount the weights of other modules. Through the first interface, the weights of other modules can be added to the corresponding layers in the base model and run.
[0038] In operation S230 , model parameters of the first model are loaded based on the first interface, so that the base model has the model parameters of the base model and the model parameters of the first model.
[0039] In embodiments of the present disclosure, the first model can be obtained by fine-tuning a base model. The base model is fine-tuned using a specified dataset. During the fine-tuning process, the base model parameters are frozen, and low-rank matrix parameters are added to the base model using LoRA (Low-Rank Adaptation) technology and fine-tuned. The model parameters of the first model may refer to the low-rank matrix parameters adjusted during the fine-tuning process.
[0040] Compared to the base model, the first model has a smaller data set and is suitable for specific tasks. By loading the first model's model parameters through the base model's first interface, the first model's model parameters can be injected into a specified layer of the base model, allowing the base model to have both the base model's model parameters and the first model's model parameters, making it suitable for both its own general domain and the specific domain corresponding to the first model.
[0041] In an embodiment of the present disclosure, the cloud server can generate a fine-tuned large model by fine-tuning the pre-trained basic model. The parameters of the large model include the model parameters of the basic model and the model parameters of the fine-tuned first model. The fine-tuned large model file is split to obtain a basic model file and a first model file respectively. The cloud server can send the basic model file and the first model file to the user. By loading the basic model file, the user can obtain a basic model with basic functions. By loading the model parameters of the first model through the first interface of the basic model, the fine-tuned large model can be obtained.
[0042] According to the embodiments of the present disclosure, by taking the existing large model as the base model and loading the model parameters of the fine-tuning model through the interface, the function of the base model can be expanded, the flexibility and scalability of the use of the local model can be improved, and this model configuration solution based on cloud distribution and local management can improve the loading speed and usage efficiency of the local model, and optimize the user experience.
[0043] According to an embodiment of the present disclosure, the method also includes: creating a second interface of the base model; loading the model parameters of the second model based on the second interface so that the base model also has the model parameters of the second model; the model parameters of the first model are different from the model parameters of the second model.
[0044] Next, combine Figure 3 The basic model of the present disclosure is further explained.
[0045] Figure 3 A schematic diagram of a basic model according to an embodiment of the present disclosure is schematically shown.
[0046] like Figure 3 As shown, the base model Base may include a first interface in1 and a second interface in2. The host memory 310 may store model parameters for the base model Base, model parameters for the first model Lora1, model parameters for the second model Lora2, and model parameters for models for other tasks or functions. The base model Base may load the model parameters of the first model Lora1 through the first interface in1, so that the base model Base has the model parameters of the base model Base and the model parameters of the first model Lora1.
[0047] In the embodiment of the present disclosure, the first model and the second model may be models fine-tuned based on different data sets. The base model may have different functions when loading the first model or the second model.
[0048] For example, the first model can be a model for intent understanding tasks, and the second model can be a model for sentiment analysis tasks. The base model can load the model parameters of the first model through the first interface to enable intent understanding, and can also load the model parameters of the second model through the second interface to enable sentiment analysis.
[0049] According to an embodiment of the present disclosure, the method further includes: unloading model parameters of the first model; and loading model parameters of the second model based on the first interface, so that the base model has model parameters of the base model and model parameters of the second model.
[0050] In the embodiment of the present disclosure, if local resources are limited or user demand is small, the function of the basic model can be adjusted by unloading and loading other model parameters and switching the model parameters on the interface in a manner similar to hot plugging.
[0051] For example, the first model can be a model for the intent understanding task, and the second model can be a model for the sentiment analysis task. The base model loads the model parameters of the first model through the first interface to provide the user with the intent understanding function. If the user does not need the intent understanding function but requires sentiment analysis, the base model can unload the parameters loaded by the first interface and reload the model parameters of the second model through the first interface, thereby providing the user with the sentiment analysis function.
[0052] According to the embodiments of the present disclosure, by switching different fine-tuning models, the basic model can be dynamically adjusted in a timely manner to adapt to different application scenarios, thereby improving the flexibility and scalability of model use and optimizing user experience.
[0053] In the embodiments of the present disclosure, the cloud server can fine-tune the base model in different directions based on LoRA to obtain a large model that meets different requirements, and then split the fine-tuned model file into a base model and different fine-tuned models (i.e., the first model and the second model described above). If the base model is installed locally, the cloud server can directly send the fine-tuned model file with specific functions and less data to the local computer, so that users can select the corresponding fine-tuned model based on their needs and load it into the base model through the interface for use.
[0054] For example, through fine-tuning, the cloud server can generate a base model file and a fine-tuned model file. The base model file occupies 3GB of memory, and the fine-tuned model file occupies 50MB of memory. If the base model is already installed locally, the cloud server can distribute the fine-tuned model file, which occupies less memory. This reduces network bandwidth consumption while ensuring model performance and data security, enabling functional updates to the local base model and optimizing the user experience.
[0055] In the embodiment of the present disclosure, the processing method can also be applied to the scenario of local model update. Figure 4 Further explanation of model updating.
[0056] Figure 4 The flowchart schematically shows a processing method according to another embodiment of the present disclosure.
[0057] like Figure 4 As shown, the processing method of this embodiment includes operations S410 to S430 and operations S441 to S443, wherein operations S410 to S430 are similar to operations S210 to S230 described above and are not described again for the sake of brevity.
[0058] In operation S410 , model parameters of a base model are loaded.
[0059] In operation S420 , a first interface of the base model is created.
[0060] In operation S430 , model parameters of the first model are loaded based on the first interface, so that the base model has the model parameters of the base model and the model parameters of the first model.
[0061] In operation S441, the model parameters of the third model are loaded. The third model has the same name as the first model, and the model parameters of the third model are different from the model parameters of the first model.
[0062] In the embodiments of the present disclosure, the first model may refer to a fine-tuned model obtained by fine-tuning the initial large model, and the third model may refer to a fine-tuned model obtained by updating the large model.
[0063] For example, by fine-tuning the base model using dataset A, a base model and a first model can be obtained. After dataset A is updated, the base model is fine-tuned again using the updated dataset to obtain a base model and a third model. The third model is applicable to the same scenarios as the first model and has the same name. The base model uses the same calling path and parameter injection method when loading the first and third models. However, the model parameters of the third model and the first model are different, and the output results of the large model obtained after loading based on different fine-tuned models when processing user input also have certain differences.
[0064] In operation S442 , the model parameters of the first model are overwritten with the model parameters of the third model.
[0065] In operation S443 , model parameters of the third model are loaded based on the first interface.
[0066] In an embodiment of the present disclosure, the base model loads the model parameters of the first model through the first interface to obtain a fine-tuned large model. When the large model needs to be updated, the base model can be fine-tuned to obtain a large model that meets new requirements or has new functions. The new large model can include the model parameters of the base model and the model parameters of the updated third model. Since the model parameters of the base model and the model parameters of the first model are already stored locally, only the model parameters of the third model including the new function can be loaded to obtain the updated large model.
[0067] For example, combined with Figure 3 , the model parameters of the third model can be used to overwrite the model parameters of the first model in the host memory 310. The overwritten first model Lora1 includes the model parameters of the third model. When the basic model calls the first model Lora1 through the same path, the obtained large model is the updated large model.
[0068] In an embodiment of the present disclosure, the model parameters of the first model and the model parameters of the updated third model may be loaded separately based on different interfaces. The version of the large model may be selected based on user needs or settings.
[0069] According to the embodiments of the present disclosure, the model is updated by obtaining an updated third model file from the cloud server to replace the first model file, which greatly improves the update efficiency of the model, avoids downloading the complete model file at each update, ensures the high-efficiency update of the local model, and optimizes the user experience.
[0070] In the embodiment of the present disclosure, the processing method can also implement switching of model functions based on user needs.
[0071] According to an embodiment of the present disclosure, the method also includes: obtaining input text; determining identification information of the input text based on the input text, the identification information including at least one of a domain identification and a function identification; in response to the identification information indicating the first model and / or the second model, running a basic model including model parameters of the first model and / or model parameters of the second model.
[0072] The input text can be text information containing user-entered questions or task requests. Based on the input text, the user's usage intent, task, or domain information can be determined. The domain identifier can represent the domain information of the input text, and the function identifier can represent the user's usage intent or task information.
[0073] For example, the user input text is "Please help me generate the answers to this CET-6 test paper." It can be determined that the identification information of the input text includes a domain identifier and a function identifier, where the domain identifier is "College English" and the function identifier is "Answer Generation."
[0074] In the embodiments of the present disclosure, the base model can refer to a large model that includes basic language understanding and generation capabilities based on identification information and possesses basic common sense. A fine-tuning model can be a LoRa module focused on a specific knowledge domain, with high professionalism and precision. The fine-tuning model can be trained, updated, and released independently of the base model. The base model can load smaller, more sophisticated fine-tuning models for different domains and functions through an interface to meet diverse demand scenarios.
[0075] For example, the fine-tuning model can be a small fine-tuning model in a variety of fields such as physics, chemistry, social knowledge, military, history, geography, programming, law, medicine, middle school knowledge, and university professional knowledge.
[0076] The basic model can load the fine-tuning model indicated by the identification information through the interface, and process the user's input accordingly by plugging in the fine-tuning model.
[0077] For example, based on the domain identifier "College English," the first model can be identified as a College English fine-tuning module. Based on the function identifier "Answer Generation," the second model can be identified as an Answer Generation fine-tuning module. The scheduler uses the identifier information to retrieve the corresponding model file from the fine-tuning model library and provide it to the base model. The base model then loads the model parameters of the first and second models to generate a large model capable of processing user input text and outputting a result corresponding to the input text.
[0078] According to the embodiments of the present disclosure, the base model automatically identifies and dynamically loads fine-tuning models from related fields. This allows it to address complex cross-domain problems through multi-domain collaborative reasoning while meeting diverse demand scenarios. Because the fine-tuning models providing different functions are relatively small in size, their application significantly improves model loading speed and scalability, optimizing the user experience.
[0079] In the embodiments of the present disclosure, users can manually select, uninstall, and install new fine-tuning models based on actual needs, and customize the combination of fine-tuning models to meet the user's personalized needs.
[0080] Figure 5 The flowchart schematically shows a processing method according to another embodiment of the present disclosure.
[0081] like Figure 5 As shown, the processing method of this embodiment includes operations S510-S520 and operations S531-S532. Operations S510-S520 are similar to the operations S210-S220 described above, and are not repeated for the sake of simplicity.
[0082] In operation S510 , model parameters of a base model are loaded.
[0083] In operation S520 , a first interface and a second interface of a base model are created.
[0084] In operation S531 , in response to the identification information indicating the first model and / or the second model, model parameters of the first model are loaded using the first interface, and model parameters of the second model are loaded using the second interface.
[0085] In an embodiment of the present disclosure, a corresponding interface may be allocated to the fine-tuning model, and when the identification information indicates that a fine-tuning model with a corresponding function is required, the model is loaded through the corresponding interface.
[0086] For example, the first model is a fine-tuned model for the biology domain, and the second model is a fine-tuned model for the chemistry domain. The base model can load the model parameters of the fine-tuned model for the biology domain through the first interface, and load the model parameters of the fine-tuned model for the chemistry domain through the second interface. In this case, the base model can be a large model that incorporates knowledge from both the biology and chemistry domains.
[0087] In operation S532 , switching is performed between the first interface and the second interface based on the identification information, and the model parameters indicated by the identification information are loaded.
[0088] In the disclosed embodiments, different interfaces of the base model are loaded with fine-tuned models for different domains. After determining the identification information based on the user's input text, the model parameters of the fine-tuned model can be directly loaded by switching to the corresponding interface, resulting in a large model that meets the user's specific domain needs.
[0089] According to the embodiments of the present disclosure, by flexibly loading LoRA models in different fields, it can be applied to different application scenarios. Compared with using models with all domain knowledge, the present disclosure can improve resource utilization while avoiding performance waste by modularizing the loading and deployment of model parameters, effectively improving the loading speed and usage efficiency of the model, and optimizing the user experience.
[0090] Figure 6 The flowchart schematically shows a processing method according to another embodiment of the present disclosure.
[0091] like Figure 6 As shown, the processing method of this embodiment includes operations S601 to S605.
[0092] In operation S601 , input text is acquired.
[0093] In operation S602, semantic processing is performed on the input text to obtain identification information corresponding to the input text.
[0094] In an embodiment of the present disclosure, the domain information or function information of the input text may be determined by semantic processing.
[0095] For example, a small intent recognition and classification model can be deployed to judge user needs based on the input text, thereby determining the identification information of the input text.
[0096] For example, the user's needs may be judged in combination with the context information of the historical conversation, thereby determining the identification information of the input text.
[0097] For example, it is also possible to determine the identification information of the input text by scanning or extracting the input text through preset keywords or sentence patterns, and performing matching based on the extracted information to judge the user's needs.
[0098] For example, identification information of the input text may also be determined by a front end or other detection modules.
[0099] In operation S603 , in response to the identification information indicating the first model, the task scheduler is used to obtain model parameters of the first model through the first interface.
[0100] In an embodiment of the present disclosure, after determining the identification information of the input text, the current task type can be determined, and the task scheduler can be used to obtain the corresponding model file from the fine-tuning module library based on the identification information.
[0101] For example, if the identification information represents "intent recognition", the task scheduler can be used to obtain a fine-tuning model for intent recognition from a fine-tuning model library as the first model, and the model parameters of the first model can be loaded through the first interface.
[0102] In operation S604 , the base model including the model parameters of the base model and the model parameters of the first model is executed.
[0103] In operation S605 , in response to the identification information indicating the second model, the model parameters of the first model are uninstalled, and the model parameters of the second model are acquired through the first interface using the task scheduler.
[0104] In an embodiment of the present disclosure, after a user has solved a certain need through a model, they may have a need for a new field or a new task. In this case, the execution process of the large model can be monitored through the resource manager, and the current model parameters can be automatically unloaded when the previous event is detected to be completed. When the task scheduler detects that the identification information has changed, it can unload the loaded fine-tuning model parameters, remove the weight of the fine-tuning model, restore the parameters to the model parameters of the base model, and load new fine-tuning model parameters to meet the user's needs.
[0105] For example, a user might enter "Please list the core formulas required for advanced mathematics." This input can be used to invoke the advanced mathematics fine-tuning model. If the large model outputs an answer and the user doesn't ask any follow-up questions about it, the task manager can simply uninstall the advanced mathematics fine-tuning model. Subsequently, the user might enter "Please help me solve a chemical equation." This second input can be used to invoke the chemistry fine-tuning model, ensuring that the large model outputs a result that meets the user's needs.
[0106] According to the embodiments of the present disclosure, by detecting user input, a specific fine-tuning model is mounted based on the current task requirements through the task scheduler, and the dynamic mounting and unmounting of the fine-tuning model is achieved through the resource manager, thereby optimizing system resource allocation and improving the utilization efficiency and loading speed of large models.
[0107] Figure 7 The structural block diagram of the processing device according to an embodiment of the present disclosure is schematically shown.
[0108] like Figure 7 As shown, the processing device 700 of this embodiment includes a first loading module 710 , a first creating module 720 and a second loading module 730 .
[0109] The first loading module 710 is used to load the model parameters of the basic model. The first loading module 710 can be used to perform the operation S210 described above, which will not be repeated here.
[0110] The first creation module 720 is used to create a first interface of the basic model. The first creation module 720 can be used to perform the operation S220 described above, which will not be repeated here.
[0111] The second loading module 730 is used to load the model parameters of the first model based on the first interface so that the base model has the model parameters of the base model and the model parameters of the first model. In one embodiment, the second loading module 730 can be used to perform the operation S230 described above, which will not be repeated here.
[0112] According to an embodiment of the present disclosure, the processing device 700 further includes a second creation module and a third loading module.
[0113] The second creation module is used to create a second interface of the basic model. The third loading module is used to load the model parameters of the second model based on the second interface, so that the basic model also has the model parameters of the second model; the model parameters of the first model are different from the model parameters of the second model.
[0114] According to an embodiment of the present disclosure, the processing device 700 further includes a fourth loading module, an overlay module, and a fifth loading module.
[0115] The fourth loading module is configured to load model parameters of a third model, wherein the third model has the same name as the first model, but the model parameters of the third model are different from those of the first model. The overwriting module is configured to overwrite the model parameters of the first model with the model parameters of the third model. The fifth loading module is configured to load the model parameters of the third model based on the first interface, so that the base model has the model parameters of the base model and the replaced model parameters of the third model.
[0116] According to an embodiment of the present disclosure, the processing device 700 further includes an unloading module and a sixth loading module.
[0117] The unloading module is used to unload the model parameters of the first model. The sixth loading module is used to load the model parameters of the second model based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the second model.
[0118] According to an embodiment of the present disclosure, the processing device 700 further includes a first acquisition module, an identification determination module, and a first operation module.
[0119] The first acquisition module is configured to acquire input text. The identification determination module is configured to determine identification information of the input text based on the input text, where the identification information includes at least one of a domain identification and a function identification. The first operation module is configured to operate a base model including model parameters of the first model and / or model parameters of the second model in response to the identification information indicating the first model and / or the second model.
[0120] According to an embodiment of the present disclosure, the first running module includes a first loading submodule and a second loading submodule.
[0121] The first loading submodule is configured to, in response to the identification information indicating the first model or the second model, unload the model parameters loaded by the first interface and load the model parameters indicated by the identification information. The second loading submodule is configured to load the model parameters of the first model using the first interface, load the model parameters of the first model using the second interface, switch between the first interface and the second interface based on the identification information, and load the model parameters indicated by the identification information.
[0122] According to an embodiment of the present disclosure, the processing device 700 further includes a second acquisition module, a semantic processing module, a first scheduling module, a second execution module, and a second scheduling module.
[0123] The second acquisition module is used to acquire input text. The semantic processing module is used to perform semantic processing on the input text to obtain identification information corresponding to the input text. The first scheduling module is used to, in response to the identification information indicating the first model, obtain model parameters of the first model using the task scheduler through the first interface. The second running module is used to run the basic model including the model parameters of the basic model and the model parameters of the first model. The second scheduling module is used to, in response to the identification information indicating the second model, uninstall the model parameters of the first model and obtain model parameters of the second model using the task scheduler through the first interface.
[0124] Figure 8 A schematic block diagram of an electronic device that can be used to implement the method of an embodiment of the present disclosure is shown. The electronic device includes: a storage device for caching model parameters; one or more processors for loading model parameters of a base model; creating a first interface for the base model; and loading model parameters of the first model based on the first interface, so that the base model has the model parameters of the base model and the model parameters of the first model.
[0125] The basic model and the first model work as a target model as a whole, including when the basic model and the first model work as a target model as a whole to respond to the user's input information, based on the model parameters of the basic model in the storage device (i.e., the memory of the local electronic device or the local terminal) and the model parameters of the first model, the processor calculates and / or infers the user's input information and gives feedback information on the input information.
[0126] The basic model described in the embodiments of the present application and the target model of the basic model and the first model as a whole are both machine learning models that can recognize natural language and / or other inputs (such as audio, video, images, tables, etc.) input into the intelligent model, and perform comprehensive language processing tasks such as analyzing semantics and answering questions, thereby generating output related to the input and / or responding to the input.
[0127] The base model described in the embodiments of this application, as well as the target model as a whole, uses training on large amounts of diverse data to learn the characteristics and patterns of natural language, thereby enabling understanding and generating natural language. These models typically have hundreds of millions to hundreds of billions of model parameters (model parameters are variables that control the behavior of intelligent models) and are capable of capturing complex relationships and patterns in natural language.
[0128] The base model described in the embodiments of this application, as well as the target model for the base model and the first model as a whole, can be generative models or generative language models (GLMs). For example, they can include large language models (LLMs) and GPT (Generative Pre-trained Transformer). The models involved in the embodiments of this application can be general large models or expert large models obtained by fine-tuning based on needs, which are not limited in the embodiments of this application.
[0129] The electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage unit 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0130] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and RAM 803. The processor 801 may also implement the method provided by the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0131] According to an embodiment of the present disclosure, electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to bus 804. Electronic device 800 may also include one or more of the following components connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 808 including a hard disk; and a communication section 809 including a network interface card such as a LAN card or modem. Communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. Removable media 811, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 810 as needed, so that computer programs read from the removable media can be installed into storage section 808 as needed.
[0132] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0133] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above, and / or one or more memories other than ROM 802 and RAM 803.
[0134] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiments of the present disclosure.
[0135] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the processor 801 executes the computer program. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0136] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0137] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from a removable medium 811. When the computer program is executed by the processor 801, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0138] It should be noted that the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information in the technical solutions disclosed herein comply with relevant laws and regulations, employ necessary confidentiality measures, and do not violate public order and good morals. In the technical solutions disclosed herein, user authorization or consent is obtained before obtaining or collecting user personal information.
[0139] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0141] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings fall within the scope of this disclosure.
[0142] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A processing method comprising: Load the model parameters of the base model; Creating a first interface of the base model; Model parameters of the first model are loaded based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the first model.
2. The method according to claim 1, further comprising: Creating a second interface of the base model; Model parameters of the second model are loaded based on the second interface, so that the basic model also has the model parameters of the second model; the model parameters of the first model are different from the model parameters of the second model.
3. The method according to claim 1, further comprising: Loading model parameters of a third model, where the third model has the same name as the first model, and the model parameters of the third model are different from the model parameters of the first model; Overwriting the model parameters of the first model with the model parameters of the third model; The model parameters of the third model are loaded based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the replaced third model.
4. The method according to claim 1, further comprising: Unloading model parameters of the first model; The model parameters of the second model are loaded based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the second model.
5. The method according to claim 2 or 4, further comprising: Get input text; Determining identification information of the input text according to the input text, the identification information including at least one of a domain identifier and a function identifier; In response to the identification information indicating the first model and / or the second model, a base model including model parameters of the first model and / or model parameters of the second model is executed.
6. The method according to claim 5, wherein the running of the base model including the model parameters of the first model and / or the model parameters of the second model comprises: In response to the identification information indicating the first model or the second model, unloading the model parameters loaded by the first interface and loading the model parameters indicated by the identification information; or The model parameters of the first model are loaded using the first interface, the model parameters of the second model are loaded using the second interface, the first interface and the second interface are switched based on the identification information, and the model parameters indicated by the identification information are loaded.
7. The method according to claim 1, further comprising: Get input text; Performing semantic processing on the input text to obtain identification information corresponding to the input text; In response to the identification information indicating the first model, obtaining model parameters of the first model through the first interface using a task scheduler; running a base model including the model parameters of the base model and the model parameters of the first model; and In response to the identification information indicating the second model, the model parameters of the first model are uninstalled, and the model parameters of the second model are acquired through the first interface by using the task scheduler.
8. A processing device comprising: A first loading module is used to load model parameters of the basic model; A creation module, configured to create a first interface of the basic model; The second loading module is configured to load the model parameters of the first model based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the first model.
9. An electronic device comprising: Storage device for caching model parameters One or more processors for loading model parameters of the base model; Creating a first interface of the base model; Model parameters of the first model are loaded based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the first model.
10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the following method: Load the model parameters of the base model; Creating a first interface of the base model; Model parameters of the first model are loaded based on the first interface, so that the basic model has the model parameters of the basic model and the model parameters of the first model.