Method and apparatus for generating target business model and processing data based on large model

By distilling knowledge on pre-trained large models, generating base models and further generating target business models, the computing resources and deployment complexity problems when applying large models in multiple business types in enterprise office scenarios are solved, and efficient and low-cost data processing effect is achieved.

CN119312943BActive Publication Date: 2025-06-17BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411302931.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-06-17
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

In the enterprise office scenario, how to apply large language models efficiently and at low cost, especially when facing multiple business types, the computing resource overhead and complex deployment.

Method used

By distilling knowledge on at least two pre-trained large models, generating a base model for the target scenario, and then distilling knowledge on the base model to obtain a target business model for the target business type, which is used to process data for the target business type.

Benefits of technology

It realizes the low-cost and efficient generation of target business models suitable for corporate office scenarios, improves the generalization and accuracy of the model, and reduces resource occupation and calculation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312943B_ABST
    Figure CN119312943B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for generating a target business model and data processing based on a large model, which relates to the field of artificial intelligence technology, specifically in the technical fields of intelligent office, big data, large models, etc. The method for generating a target business model based on a large model includes: performing knowledge distillation on at least two pre-trained large models to obtain a base model for a target scenario; each pre-trained large model corresponds to one of at least two business types included in the target scenario; performing knowledge distillation on the base model to obtain a target business model for a target business type among the at least two business types; the target business model is used to process data of the target business type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as intelligent office, big data, large models, etc., and particularly relates to a method and device for generating a target business model and processing data based on a large model. Background Art

[0002] Currently, with the rapid development of artificial intelligence technology, large language models (LLMs, simply referred to as large models) have been widely applied in various scenarios.

[0003] In the enterprise office scenario, how to efficiently and low-costly apply large models is a problem that needs to be solved. Summary of the Invention

[0004] The present disclosure provides a method and device for generating a target business model and processing data based on a large model.

[0005] According to one aspect of the present disclosure, a method for generating a target business model based on a large model is provided, including: performing knowledge distillation on at least two pre-trained large models to obtain a base model for a target scenario; each pre-trained large model corresponds to one of at least two business types included in the target scenario; performing knowledge distillation on the base model to obtain a target business model for a target business type among the at least two business types; the target business model is used to process data of the target business type.

[0006] According to another aspect of the present disclosure, a data processing method is provided, including: obtaining target data of a target business type; using the target business model corresponding to the target business type to process the target data to obtain a data processing result; wherein, the target business model is generated by using the method described in any item of any aspect above.

[0007] According to another aspect of the present disclosure, a device for generating a target business model based on a large model is provided, including: a first generation module for performing knowledge distillation on at least two pre-trained large models to obtain a base model for a target scenario; each pre-trained large model corresponds to one of at least two business types included in the target scenario; a second generation module for performing knowledge distillation on the base model to obtain a target business model for a target business type among the at least two business types; the target business model is used to process data of the target business type.

[0008] According to another aspect of the present disclosure, there is provided a data processing apparatus, including: an acquisition module configured to acquire target data of a target service type; a processing module configured to process the target data by using a target service model corresponding to the target service type to obtain a data processing result; wherein, the target service model is generated by using the method described in any one of the above aspects.

[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in any one of the above aspects.

[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in any one of the above aspects.

[0011] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method described in any one of the above aspects.

[0012] According to the embodiments of the present disclosure, data processing can be performed efficiently and at low cost.

[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings are used to better understand the solution and do not limit the present disclosure. Among them:

[0015] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0016] Figure 2 is a schematic diagram of an application scenario for implementing the embodiments of the present disclosure;

[0017] Figure 3 is a schematic diagram according to the second embodiment of the present disclosure;

[0018] Figure 4 is a schematic diagram of knowledge distillation in the first stage provided according to the embodiments of the present disclosure;

[0019] Figure 5Schematic diagram of the second stage of knowledge distillation provided according to an embodiment of the present disclosure;

[0020] Figure 6 Schematic diagram according to the third embodiment of the present disclosure;

[0021] Figure 7 Schematic diagram according to the fourth embodiment of the present disclosure;

[0022] Figure 8 Schematic diagram according to the fifth embodiment of the present disclosure;

[0023] Figure 9 Schematic diagram of an electronic device for implementing the method for generating a target business model or data processing method based on a large model according to an embodiment of the present disclosure. Detailed implementation manners

[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0025] In a vertical scenario, usually, a pre-trained large model is fine-tuned to obtain a fine-tuned model corresponding to the vertical scenario, and the fine-tuned model is used for inference.

[0026] The enterprise office scenario is different from the vertical scenario. The vertical scenario usually only deals with a single type of data processing, such as code generation; while the enterprise office scenario faces multiple business types, for example, including the following business types: customer service conversations, enterprise document processing, intelligent recommendation, etc. If each business type fine-tunes the pre-trained large model to obtain a fine-tuned model corresponding to the corresponding type, there will be problems such as large computational resource overhead and complex deployment.

[0027] Therefore, in the enterprise office scenario, it is necessary to solve the problem of how to efficiently and low-costly obtain a target model suitable for the target business type.

[0028] To obtain a model suitable for the enterprise office scenario efficiently and at low cost, the present disclosure provides the following embodiments.

[0029] Figure 1 Schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a method for generating a target business model based on a large model, as Figure 1 shown, the method includes:

[0030] 101. Perform knowledge distillation on at least two pre-trained large models to obtain a base model for the target scenario; each pre-trained large model corresponds to one of at least two business types included in the target scenario.

[0031] 102. Perform knowledge distillation on the base model to obtain a target business model for the target business type among the at least two business types; the target business model is used to process data of the target business type.

[0032] Among them, a pre-trained large model refers to an existing large model obtained through pre-training.

[0033] In the embodiments of the present disclosure, processing is performed based on multiple (at least two) pre-trained large models.

[0034] These pre-trained large models include, for example: GPT-o, GPT-4, EB4, Qwen3.5, etc. GPT-o and GPT-4 are respectively one of the General Pre-trained Transformer (GPT) series of models; EB4 is one of the Wenxin Yiyan series of models; Qwen3.5 is one of the Qwen series of models.

[0035] The core idea of knowledge distillation is to use a trained complex model (teacher model) to guide the training of a simple model (student model), so that the student model is close to the teacher model in performance, but the number of parameters is greatly reduced.

[0036] In the first stage, use the pre-trained large model as the teacher model and the base model of the target scenario as the student model, and through knowledge distillation, generate the base model of the target scenario from multiple existing pre-trained large models.

[0037] The target scenario is the scenario where the model is to be applied, and the target scenario involves multiple business types.

[0038] Taking the target scenario as the enterprise office scenario as an example, this scenario involves multiple business types, such as customer service conversations, enterprise document processing, intelligent recommendation, etc.

[0039] If for each business type, fine-tuning the pre-trained large model to obtain the business model corresponding to the business type, there will be problems such as low efficiency and high cost.

[0040] Therefore, in this embodiment, a base model shared by different business types can be obtained based on the pre-trained large model first, and then the business model corresponding to each business type can be obtained based on the base model.

[0041] Since the base model is shared by different business types, it needs to have a certain degree of generalization to meet the data processing requirements of different business types.

[0042] To make the base model have generalization ability, knowledge distillation is carried out on multiple pre-trained large models to learn knowledge from each pre-trained large model and improve the generalization ability of the base model.

[0043] Multiple pre-trained large models respectively correspond to different business types, which is conducive to the base model learning knowledge of different business types.

[0044] Suppose the target scenario includes three business types. For example, the above enterprise office scenario includes the following three business types: customer service dialogue, enterprise document processing, and intelligent recommendation. Then, three pre-trained large models can be selected, corresponding to the above three business types respectively. For example, Qwen3.5 has good dialogue ability, so the pre-trained large model Qwen3.5 can be used as the pre-trained large model corresponding to the customer service dialogue business type.

[0045] In this way, the base model can learn different knowledge from the pre-trained large models corresponding to different business types, improving the generalization ability and accuracy of the base model.

[0046] After obtaining the base model, knowledge distillation can also be used to perform knowledge distillation on the base model of the target scenario to obtain the target business model of the target business type in the target scenario.

[0047] That is, in the second stage, the base model is used as the teacher model, and the target business model is used as the student model. Through knowledge distillation, the target business model is generated from the base model.

[0048] The target business type is one or more of the multiple business types included in the target scenario.

[0049] For example, samples corresponding to the first business type can be used to perform knowledge distillation on the base model to obtain the first business model, which is used for the inference process of the first business type. Specifically, if the first business type is customer service dialogue, then the first business model for customer service dialogue is generated. In the inference stage, this first business model is used for customer service dialogue. Another example is to use samples corresponding to the second business type to perform knowledge distillation on the base model to obtain the second business model, which is used for the inference process of the second business type. Specifically, if the second business type is enterprise document processing, then the second business model for enterprise document processing is generated. In the inference stage, this second business model is used for enterprise document processing, such as document classification and document abstract generation.

[0050] Through the knowledge distillation process, a student model with a smaller scale than the teacher model can be obtained, but this student model can inherit the performance of the teacher model, so that a model with better performance and smaller scale can be obtained.

[0051] In the model application stage, the target business model can be used for inference. For example, it can process business data of types such as text, images, and Application Programming Interfaces (APIs) to obtain accurate and efficient data processing results.

[0052] In this embodiment, by performing knowledge distillation on a pre-trained large model, a base model for the target scenario is obtained, and then knowledge distillation is performed on this base model to obtain a target business model for the target business type, which can generate the target business model at low cost and efficiently. In addition, by performing knowledge distillation on multiple pre-trained large models to obtain the base model, and each pre-trained large model corresponds to a business type, the base model can learn knowledge of different business types, improving the generalization and accuracy of the base model, and further improving the accuracy of the target business model.

[0053] In specific implementation, the model generation method and the data processing method based on the generated target business model can be specifically executed by a physical device, which can be a terminal device, a server, a processing unit, etc. Since the target business model is obtained through two stages of knowledge distillation, the scale of the target business model is small, occupying less resources. And due to its small scale, it can also improve the data processing efficiency, thus saving the storage resources and computing resources of the physical device, improving the processing speed, and enhancing the internal performance of the physical device.

[0054] Taking text data as an example, the following solution can be adopted to obtain a target business model suitable for text data:

[0055] Based on the first text data sample, perform knowledge distillation on at least two pre-trained large models to obtain a base model for the target scenario; each pre-trained large model corresponds to one of at least two business types included in the target scenario; and,

[0056] Based on the second text data sample, perform knowledge distillation on the base model to obtain a target business model for the target business type among the at least two business types; the target business model is used to process data of the target business type;

[0057] Wherein, the first text data sample includes: text data samples corresponding to the at least two business types; the second text data sample includes: text data samples corresponding to the target business type.

[0058] Specifically, in the first stage: Assume that the target scenario involves two business types, namely customer service conversations and document processing. Then, a pre-trained large model suitable for customer service conversations can be selected as the first pre-trained large model, and a large model suitable for document processing can be selected as the second pre-trained large model. Specifically, for example, since Qwen3.5 has good conversation capabilities, Qwen3.5 is used as the first pre-trained large model. Since EB4 has good document processing capabilities such as document classification, EB4 is used as the second pre-trained large model.

[0059] In addition, samples of various business types can be collected as the first text data samples. For example, existing conversation data samples and document classification samples are collected to form the first text data sample.

[0060] After determining the first pre-trained large model and the second pre-trained large model, and obtaining the first text data sample, the first text data sample is input into the above two pre-trained large models respectively, and into the base model to be trained. These models process the first text data, such as extracting the features of the first text data, and performing data generation or classification based on the extracted features, etc., to obtain corresponding output results. For example, the first pre-trained large model obtains the first conversation prediction result, the second pre-trained large model obtains the first document processing prediction result, and the base model will obtain prediction results corresponding to multiple business types, which are respectively called the second conversation prediction result and the second document processing prediction result. These conversation prediction results and document processing prediction results are all text data. Then, a first loss function can be constructed based on the first conversation prediction result and the second conversation prediction result, and a second loss function can be constructed based on the first document processing prediction result and the second document processing prediction result. The sum of the first loss function and the second loss function is used as the loss function for the first stage. The parameters of the base model are adjusted using this loss function until a preset end condition is reached, and the final base model is obtained.

[0061] What is involved in this process is text data and its processing process. Since pre-trained large models corresponding to multiple business types are involved in processing the text data, and the base model will obtain text prediction results corresponding to multiple business types, a base model suitable for processing text data of multiple business types can be obtained, improving the processing accuracy and generalization ability of the base model for text data of multiple business types.

[0062] In the second stage: Assume that the target business type is customer service conversation. Then, existing conversation data samples can be collected as the second text data samples. The second text data samples are respectively input into the base model obtained in the first stage and the target business model corresponding to the customer service conversation to be trained. The base model and the target business model process the second text data, such as extracting text features of the second text data and generating conversations based on the extracted text features, to obtain corresponding output results. For example, the base model obtains a third conversation prediction result, and the target business model obtains a fourth conversation prediction result. These conversation prediction results are all text data. Then, a loss function for the second stage can be constructed based on the third conversation prediction result and the fourth conversation prediction result, and the parameters of the target business model are adjusted using this loss function until a preset end condition is reached, obtaining the final target business model. When applied, this target business model is used for customer service conversations.

[0063] In this way, through processing such as text data generation using the second text data samples (conversation data samples), a target business model suitable for customer service conversations can be obtained, improving the accuracy of this target business model.

[0064] To better understand the present disclosure, the application scenarios involved in the present disclosure are described as follows:

[0065] Figure 2 It is a schematic diagram for implementing the application scenario of the embodiments of the present disclosure.

[0066] As Figure 2 shown, this application scenario includes: a pre-trained large model 201, a base model 202 for the target scenario, and business models 203 corresponding to each business type.

[0067] Among them, a large model library can be pre-constructed, and existing selectable pre-trained large models are stored in this large model library. When targeting a certain target scenario, multiple corresponding pre-trained large models can be selected according to the multiple business types included in this target scenario.

[0068] In this embodiment, taking the target scenario as the enterprise office scenario as an example, in the enterprise office scenario, assume that the multiple business types it involves include: customer service conversations, enterprise document processing, and intelligent recommendation. Therefore, pre-trained large models suitable for the above three business types can be selected from the large model library. Assume that the selected pre-trained large models are respectively represented as the first pre-trained large model to the third pre-trained large model. In this way, the base model for the target scenario can learn knowledge from multiple teacher models (pre-trained large models), improving the accuracy, stability, and generalization ability of the base model. The business models corresponding to the above business types are respectively called the first business model, the second business model, and the third business model.

[0069] The base model is obtained by performing knowledge distillation on a pre-trained large model.

[0070] The business model, in its initial state, is obtained by performing knowledge distillation on the base model. Further, to improve the performance (such as accuracy or stability) of the business model, each business model can also perform an evolution operation to achieve self-evolution.

[0071] Model evolution, at its core, lies in continuously optimizing and improving the model to adapt to new requirements or environmental changes. This process may involve improvements to the model structure and / or parameters. In this embodiment, the business model can specifically perform the evolution operation using a preset evolution algorithm.

[0072] Combined with the above application scenarios, the present disclosure also provides the following embodiments.

[0073] Figure 3 FIG. is a schematic diagram according to the second embodiment of the present disclosure. This embodiment provides a method for generating a target business model based on a large model, and the method includes:

[0074] 301. Perform knowledge distillation on at least two pre-trained large models to obtain a base model for the target scenario; each pre-trained large model corresponds to one of at least two business types included in the target scenario.

[0075] Specifically, each pre-trained large model can be used to process the first training samples of the at least two business types to obtain a first output result; the base model can be used to process the first training samples to obtain at least two second output results; each second output result corresponds to one of the business types; a first loss function is constructed based on the first output result and the second output results; and the model parameters of the base model are adjusted based on the first loss function.

[0076] Among them, the knowledge distillation from the pre-trained large model to the base model can be referred to as the knowledge distillation in the first stage. Figure 4 FIG. is a schematic diagram of the knowledge distillation in the first stage provided according to the embodiment of the present disclosure.

[0077] As Figure 4 shown, the training samples used in the knowledge distillation in the first stage are called first training samples, and the first training samples include data of multiple business types involved in the target scenario.

[0078] For example, if the target scenario is an enterprise office scenario, and the multiple business types involved include: customer service conversations, enterprise document processing, and intelligent recommendations, then it is necessary to collect data of these three business types as the first training samples.

[0079] Specifically, various types of business data can be collected from the enterprise's log data to obtain the first training sample.

[0080] In addition, to improve data quality, the collected raw data can be preprocessed, and the preprocessed data is used as the training sample. The preprocessing includes, for example, removing sensitive data, removing garbled characters or data with inconsistent formats, semantic deduplication, removing data with poor quality (such as overly short or long sentences, or sentences with insufficient semantic content), etc.

[0081] After obtaining the first training sample, the first training sample is input into each pre-trained large model and the base model to be trained. The output of the pre-trained large model is called the first output result, and the output of the base model is called the second output result.

[0082] Each pre-trained large model corresponds to a type of business. Then, the first output result of each pre-trained large model is one, corresponding to a type of business. For example, the first pre-trained large model outputs a dialogue result, and the second pre-trained large model outputs a document classification result, etc.

[0083] The base model can process data of multiple business types. Therefore, the second output results are multiple, corresponding to each type of business respectively. For example, the base model can obtain a dialogue result and a document classification result, etc. Specifically, the base model can include multiple output layers, and each output layer is used to obtain the output result corresponding to a type of business.

[0084] As Figure 4 shown, taking three types of business as an example, therefore, the pre-trained large models are the first pre-trained large model to the third pre-trained large model respectively, and each pre-trained large model outputs one first output result; the base model outputs three second output results.

[0085] After that, a first loss function is constructed based on the first output result and the second output result. Specifically, the loss function for each type of business can be constructed using the first output result and the second output result corresponding to each type of business, and then the first loss function is constructed according to the loss functions for each type of business.

[0086] For example, the first output results output by the three pre-trained large models are respectively denoted as y11, y12, y13, and the three second output results output by the base model are respectively denoted as p21, p22, p23. Assuming that y11 and p21 correspond to the same type of business, such as both being dialogue results, and the rest are similar, then loss1 can be constructed based on y11 and p21, loss2 can be constructed based on y12 and p22, and loss3 can be constructed based on y13 and p23. After that, the first loss function is obtained by directly adding or weighted adding loss1, loss2, and loss3.

[0087] After obtaining the first loss function, algorithms such as backpropagation can be used to adjust the model parameters of the base model using the first loss function until a preset end condition is reached, and the base model when the end condition is reached is used as the final base model.

[0088] In this embodiment, based on the first output results output by each pre-trained large model and the second output results of the corresponding business types output by the base model, a first loss function is constructed. Using the first loss function to adjust the parameters of the base model can enable the base model to accurately learn the knowledge of each pre-trained large model without the need to understand the internal structure of the pre-trained large model, thereby efficiently and accurately generating the base model.

[0089] 302. Perform knowledge distillation on the base model to obtain a target business model for the target business type among the at least two business types; the target business model is used to process data of the target business type.

[0090] Assume that the target scenario includes three business types and a business model for the first business type needs to be obtained. Then, knowledge distillation can be performed on the base model to obtain the first business model.

[0091] Specifically, the base model can be used to process the second training samples to obtain a third output result; the second training samples include: data of the target business type; the target business model is used to process the second training samples to obtain a fourth output result; a second loss function is constructed based on the third output result and the fourth output result; based on the second loss function, the model parameters of the target business model are adjusted.

[0092] Among them, the knowledge distillation from the base model to the business model can be called the knowledge distillation in the second stage. Figure 5 It is a schematic diagram of the knowledge distillation in the second stage provided according to the embodiments of the present disclosure.

[0093] Such as Figure 5 shown, the training samples used in the knowledge distillation in the second stage are called second training samples, and the second training samples include data of the target business type.

[0094] For example, the target scenario is an enterprise office scenario, and the multiple business types involved include: customer service conversations, enterprise document processing, and intelligent recommendations. Assume that the target business type is customer service conversations, then data of the customer service conversation business type needs to be collected as the second training sample.

[0095] Specifically, data of the required business type can be collected from the enterprise's log data to obtain the second training sample.

[0096] In addition, to improve data quality, the collected raw data can be preprocessed, and the preprocessed data can be used as training samples. The preprocessing may include removing sensitive data, removing garbled characters or data with inconsistent formats, semantic deduplication, removing data with poor quality (such as too short or too long sentences, or insufficient semantic content), etc.

[0097] As Figure 5 shown, after obtaining the second training sample, the second training sample is input into the base model obtained by the first-stage knowledge distillation and the business model to be trained (target business model). The output of the base model is called the third output result, and the output of the target business model is called the fourth output result.

[0098] After that, a second loss function is constructed based on the third output result and the fourth output result. Specifically, the third output result can be specifically selected as the output result corresponding to the target business type of the base model. For example, if the base model has three output layers, each output layer corresponds to a business type. Assuming that the target business type is the first business type, then the output result of the output layer corresponding to the first business type can be selected as the third output result. Each business model corresponds to a business type, so each target business model can output a fourth output result corresponding to the business type. In this way, the third output result and the fourth output result correspond to the same business type, such as both being dialogue results. Therefore, a second loss function corresponding to each target business model can be constructed according to the third output result and the fourth output result, and then the corresponding target business model can be adjusted according to the second loss function, such as adjusting the business model of customer service dialogue.

[0099] The specific adjustment process may include: using the second loss function to adjust the model parameters of the target business model by algorithms such as backpropagation until a preset end condition is reached, and taking the target business model when the end condition is reached as the final target business model.

[0100] In this embodiment, by constructing a second loss function based on the third output result output by the base model and the fourth output result output by the target business model, and using the second loss function to adjust the parameters of the target business model, the target business model can accurately learn the knowledge of the base model without the need to understand the internal structure of the base model, thereby efficiently and accurately generating the target business model.

[0101] The above process involves the generation of the base model and the target business model. In some cases, the base model and the target business model can also be updated.

[0102] This update process can be called an evolution process.

[0103] Further, for the base model, its evolution process mainly involves re - performing knowledge distillation based on a new pre - trained large model to obtain an updated base model.

[0104] For the business model, its evolution process mainly involves self - evolving according to the inference results (data processing results) of the business model using a preset evolution algorithm to obtain an updated business model.

[0105] To this end, in some embodiments, it may further include:

[0106] 303. In response to determining that at least some of the at least two pre - trained large models have been updated, perform knowledge distillation on the updated pre - trained large models to obtain an updated base model; perform knowledge distillation on the updated base model to obtain an updated target business model.

[0107] For example, there are three pre - trained large models, and at least one of them has been updated. Then, the updated pre - trained large models can be re - distilled, such as the knowledge distillation in the first stage as described above, to obtain an updated base model.

[0108] Since the base model has been updated, the updated base model can be further re - distilled to obtain an updated target business model.

[0109] Further, in the first stage, to reduce the computational load, the same first training sample can be used in the update process. In this way, only the output results of the updated pre - trained large models need to be adjusted, and the other output results remain unchanged, thereby reducing the computational load and saving resources.

[0110] For example, referring to Figure 4 , assume that only the first pre - trained large model has been updated. Then, the first training sample can be input into the updated first pre - trained large model to obtain a new first output result output by the first pre - trained large model, while the output results of the other pre - trained large models and the benchmark model remain unchanged and the previously calculated results can be directly used without recalculation, thereby reducing the computational load and saving resources.

[0111] In addition, after the base model is updated, in the second stage, the same second training sample can also be used, so that the output results of the previously calculated target business model can be directly used, thereby reducing the computational load and saving resources.

[0112] In this embodiment, updating the base model based on the updated pre - trained large model and updating the target business model according to the updated base model can achieve timely updating of the target business model and improve the accuracy of the target business model.

[0113] 304. Process the data of the target business type using the target business model to obtain a data processing result; in response to determining that the data processing result meets a preset evolution condition, evolve the target business model to obtain an evolved target business model.

[0114] After generating the target business model, it can be deployed. Then, in the online process, use this target business model for data processing to obtain a data processing result. For example, use the business model of customer service conversations to execute the customer service conversation process to obtain a customer service conversation result.

[0115] After obtaining the data processing result, if the data processing result triggers an evolution operation, the target business model will perform a self-evolution operation.

[0116] Specifically, evolution conditions can be preset. For example, when the accuracy rate of the data processing result is lower than a certain set threshold, at this time, the target business model can perform an evolution operation according to a preset evolution algorithm. The evolution algorithm, for example, includes a process for updating model parameters. Based on this evolution algorithm, the automatic update (evolution) of the target business model can be achieved, thereby obtaining an evolved target business model.

[0117] After that, use this evolved target business model for subsequent data processing of the target business type.

[0118] In this embodiment, when the data processing result triggers an evolution operation, evolving the target business model can achieve the timely update of the target business model and improve the accuracy of the target business model.

[0119] Figure 6 This is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides a data processing method, including:

[0120] 601. Obtain target data of the target business type.

[0121] 602. Process the target data using the target business model corresponding to the target business type to obtain a data processing result.

[0122] Among them, the target business model is generated by using the method shown in any of the above embodiments.

[0123] For example, for customer service conversation data, the customer service conversation model can be used for processing to obtain a customer service conversation result.

[0124] In this embodiment, since the target business model is obtained through knowledge distillation and has a small scale, it can reduce resource occupancy and improve processing efficiency; since the target business model has the advantage of high accuracy, the accuracy of the data processing result can be improved based on this model.

[0125] Figure 7 It is a schematic diagram according to the fourth embodiment of the present disclosure. In this embodiment, a target business model generation device based on a large model is provided. The device 700 includes: a first generation module 701 and a second generation module 702.

[0126] The first generation module 701 is used to perform knowledge distillation on at least two pre-trained large models to obtain a base model for the target scenario; each pre-trained large model corresponds to one of at least two business types included in the target scenario; the second generation module 702 is used to perform knowledge distillation on the base model to obtain a target business model for the target business type among the at least two business types; the target business model is used to process data of the target business type.

[0127] In this embodiment, by performing knowledge distillation on the pre-trained large models, a base model for the target scenario is obtained, and then knowledge distillation is performed on this base model to obtain a target business model for the target business type, which can generate the target business model at low cost and efficiently; in addition, by performing knowledge distillation on multiple pre-trained large models to obtain the base model, and each pre-trained large model corresponds to one business type, the base model can learn knowledge of different business types, improving the generalization and accuracy of the base model, and further improving the accuracy of the target business model.

[0128] In some embodiments, the first generation module 701 is further used to:

[0129] Use each pre-trained large model to process the first training sample to obtain a first output result; the first training sample includes: data of the at least two business types;

[0130] Use the base model to process the first training sample to obtain at least two second output results; each second output result corresponds to one of the business types;

[0131] Construct a first loss function based on the first output result and the second output results;

[0132] Adjust the model parameters of the base model based on the first loss function.

[0133] In this embodiment, by constructing a first loss function based on the first output result output by each pre-trained large model and the second output results corresponding to the business types output by the base model, and using the first loss function to adjust the parameters of the base model, the base model can accurately learn the knowledge of each pre-trained large model without knowing the internal structure of the pre-trained large model, thus efficiently and accurately generating the base model.

[0134] In some embodiments, the second generation module 702 is further configured to:

[0135] Process the second training sample using the base model to obtain a third output result; the second training sample includes: data of the target business type;

[0136] Process the second training sample using the target business model to obtain a fourth output result;

[0137] Construct a second loss function based on the third output result and the fourth output result;

[0138] Adjust the model parameters of the target business model based on the second loss function.

[0139] In this embodiment, by constructing a second loss function based on the third output result of the base model and the fourth output result of the target business model, and using the second loss function to adjust the parameters of the target business model, the target business model can accurately learn the knowledge of the base model without the need to understand the internal structure of the base model, thus efficiently and accurately generating the target business model.

[0140] In some embodiments, the apparatus 700 may further include:

[0141] A first update module, configured to perform knowledge distillation on the updated pre-trained large model to obtain an updated base model in response to determining that at least some of the at least two pre-trained large models are updated; perform knowledge distillation on the updated base model to obtain an updated target business model.

[0142] In this embodiment, updating the base model based on the updated pre-trained large model and updating the target business model according to the updated base model can achieve timely update of the target business model and improve the accuracy of the target business model.

[0143] In some embodiments, the apparatus 700 may further include:

[0144] A second update module, configured to process the data of the target business type using the target business model to obtain a data processing result; in response to determining that the data processing result meets a preset evolution condition, evolve the target business model to obtain an evolved target business model.

[0145] In this embodiment, evolving the target business model when the data processing result triggers an evolution operation can achieve timely update of the target business model and improve the accuracy of the target business model.

[0146] Figure 8It is a schematic diagram according to the fifth embodiment of the present disclosure. This embodiment provides a data processing device, and the device 800 includes: an acquisition module 801 and a processing module 802.

[0147] The acquisition module 801 is used to acquire target data of a target service type; the processing module 802 is used to process the target data by using a target service model corresponding to the target service type to obtain a data processing result.

[0148] Among them, the target service model is generated by using the method shown in any of the above embodiments.

[0149] For example, for customer service dialogue data, a customer service dialogue model can be used for processing to obtain a customer service dialogue result.

[0150] In this embodiment, since the target service model is obtained through knowledge distillation and has a small scale, it can reduce resource occupancy and improve processing efficiency; since the target service model has the advantage of high accuracy, the accuracy of the data processing result can be improved based on this model.

[0151] It can be understood that in the embodiments of the present disclosure, the same or similar content in different embodiments can be referred to each other.

[0152] It can be understood that the "first", "second", etc. in the embodiments of the present disclosure are only used for distinction and do not represent the level of importance, the sequence of time, etc.

[0153] It can be understood that if there is no special limitation on the sequence of steps involved in the process, it means that the timing relationship between these steps is not limited.

[0154] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0155] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0156] Figure 9FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device 900 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0157] As Figure 9 shown, the electronic device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0158] A plurality of components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0159] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the model generation method or the data processing method. For example, in some embodiments, the model generation method or the data processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the model generation method or the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the model generation method or the data processing method by any other suitable means (e.g., by means of firmware).

[0160] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable task processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0163] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0164] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0165] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system or a server combined with a blockchain.

[0166] It should be understood that various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0167] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for generating a target business model based on a large model, comprising: Perform knowledge distillation on at least two pre-trained large models to obtain a base model shared by different business types in the target scenario; The target scenario refers to a scenario to which the model is to be applied, and the target scenario involves at least two business types; each pre-trained large model corresponds to one of the at least two business types; Performing knowledge distillation on the base model to obtain a target business model of a target business type among the at least two business types; the target business model is used to process data of the target business type; The step of performing knowledge distillation on at least two pre-trained large models to obtain a base model shared by different business types in the target scenario includes: Using each of the pre-trained large models, a first training sample is processed to obtain a first output result; the first training sample includes: data of the at least two business types; Using the base model, processing the first training sample to obtain at least two second output results; each second output result corresponds to the one business type; Constructing a first loss function based on the first output result and the second output result; Based on the first loss function, model parameters of the base model are adjusted.

2. The method according to claim 1, wherein: The performing knowledge distillation on the base model to obtain a target business model of a target business type among the at least two business types includes: The base model is used to process a second training sample to obtain a third output result; the second training sample includes: data of the target business type; Using the target business model, processing the second training sample to obtain a fourth output result; Constructing a second loss function based on the third output result and the fourth output result; Based on the second loss function, the model parameters of the target business model are adjusted.

3. The method according to claim 1, further comprising: In response to determining that at least part of the at least two pre-trained large models is updated, performing knowledge distillation on the updated pre-trained large models to obtain an updated base model; Knowledge distillation is performed on the updated base model to obtain an updated target business model.

4. The method according to claim 1, further comprising: Processing the data of the target business type using the target business model to obtain a data processing result; In response to determining that the data processing result meets a preset evolution condition, the target business model is evolved to obtain an evolved target business model.

5. A data processing method, comprising: Obtain target data of target business type; Using a target business model corresponding to the target business type, the target data is processed to obtain a data processing result; Wherein, the target business model is generated by using the method according to any one of claims 1-4.

6. A target business model generation device based on a large model, comprising: A first generation module is used to perform knowledge distillation on at least two pre-trained large models to obtain a base model shared by different business types in a target scenario; The target scenario refers to a scenario to which the model is to be applied, and the target scenario involves at least two business types; each pre-trained large model corresponds to one of the at least two business types; A second generating module is used to perform knowledge distillation on the base model to obtain a target business model of a target business type among the at least two business types; the target business model is used to process data of the target business type; Wherein, the first generating module is further used for: Using each of the pre-trained large models, a first training sample is processed to obtain a first output result; the first training sample includes: data of at least two business types; Using the base model, processing the first training sample to obtain at least two second output results; each second output result corresponds to the one business type; Constructing a first loss function based on the first output result and the second output result; Based on the first loss function, model parameters of the base model are adjusted.

7. The device according to claim 6, wherein: The second generating module is further used for: Using the base model, processing the second training sample to obtain a third output result; The second training sample includes: data of the target service type; Using the target business model, processing the second training sample to obtain a fourth output result; Constructing a second loss function based on the third output result and the fourth output result; Based on the second loss function, the model parameters of the target business model are adjusted.

8. The apparatus according to claim 6, further comprising: The first updating module is used to, in response to determining that at least part of the at least two pre-trained large models are updated, perform knowledge distillation on the updated pre-trained large model to obtain an updated base model; and perform knowledge distillation on the updated base model to obtain an updated target business model.

9. The apparatus according to claim 6, further comprising: A second updating module, used to process the data of the target business type using the target business model to obtain a data processing result; In response to determining that the data processing result meets a preset evolution condition, the target business model is evolved to obtain an evolved target business model.

10. A data processing device, comprising: An acquisition module, used to acquire target data of a target business type; A processing module, used to process the target data using a target business model corresponding to the target business type to obtain a data processing result; Wherein, the target business model is generated by using the method according to any one of claims 1-4.

11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.

13. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Model generation method and device, image classification method and device, equipment and medium

    CN116797829A

  • Classification method and device, equipment, storage medium and product

    CN117312934A