Model training method and device, text processing method and device, electronic equipment and storage medium
By utilizing key information from high-parameter, large models to guide the training of lightweight models, the problems of high-parameter model resource consumption and low training efficiency are solved, enabling efficient deployment and rapid adaptation of lightweight models in customer service call center operations.
Patent Information
- Application Number
- CN202510983875.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-11-25
AI Technical Summary
High-parameter, large-scale models require significant resources to deploy in customer service call center operations, have low training efficiency, and low-parameter models are not adaptable to various scenarios, making it difficult to meet the needs of real-time and high-concurrency scenarios.
By using the key information output by the first major model as a supervision signal to guide the training of the lightweight second major model, the lightweight model is used to replace the high-parameter model, reducing the number of parameters and performing fine-tuning training. Supervised fine-tuning and knowledge distillation techniques are employed to reduce resource requirements.
It enables efficient deployment of lightweight models in real-time and high-concurrency scenarios, reduces resource consumption, improves training efficiency and scenario adaptability, and meets the requirements of real-time and high-concurrency calls.
Smart Images

Figure CN121009364A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a model training method and device, a text processing method and device, an electronic device and a storage medium. BACKGROUND
[0002] The customer service field has the characteristics of large data volume, unstructured, and more pure dialogue text data. In the traditional customer service operation, extracting key information requires relying on manual experience to determine the key information, which is inefficient and inaccurate.
[0003] In order to improve efficiency and accuracy, a high-parameter large model can be used to extract key information in customer service operation. However, the high-parameter large model requires a large amount of deployment resources, and in the case of limited resources, it is often difficult to meet the needs of real-time scenarios and high-concurrency calling scenarios, and the existing large model fine-tuning training method relies on manual labeling, which is time-consuming and laborious, and the efficiency of making data sets is low, and cannot adapt to the rapidly changing business environment. SUMMARY
[0004] In view of the above problems, the embodiments of the present application provide a model training method and device, a text processing method and device, an electronic device and a storage medium, to solve the problems of large resource requirement of high-parameter large model deployment, low training efficiency, and poor adaptability of low-parameter untrained large model scenarios.
[0005] According to an aspect of an embodiment of the present application, a model training method is provided, the method comprising:
[0006] inputting first prompt information and first sample text information into a first large model to obtain first key information in the first sample text information output by the first large model;
[0007] using the first prompt information, the first sample text information and the first key information as a first training sample, and using the first key information as a label corresponding to the first sample text information;
[0008] training a second large model to be trained using the first training sample to obtain a trained second large model, the parameter quantity of the second large model being less than that of the first large model.
[0009] Optionally, the first prompt information includes a text information analysis method, a text information analysis example and a key information output format.
[0010] Optionally, before the first prompt information and the first sample text information are input into the first large model, the method further comprises: extracting first dialogue summary information corresponding to the first dialogue text information, and using the first dialogue summary information as the first sample text information.
[0011] Optionally, before the training the second large model using the first training sample, the method further comprises: taking the second prompt information, second sample text information and first rejection information corresponding to the second sample text information as a second training sample, the first rejection information as a label of the second sample text information; and the training the second large model using the first training sample comprises: training the second large model using the first training sample and the second training sample.
[0012] Optionally, the method further comprises:
[0013] taking the second prompt information, third sample text information and second rejection information corresponding to the third sample text information as a first test sample, the second rejection information as a label of the third sample text information;
[0014] testing the trained second large model using the first test sample to obtain an output accuracy rate of the rejection information corresponding to the trained second large model.
[0015] Optionally, the method further comprises:
[0016] taking the first prompt information and fourth sample text information as a second test sample, testing the trained second large model using the second test sample to obtain an output qualified rate of the key information corresponding to the trained second large model;
[0017] and / or,
[0018] taking the first prompt information and fifth sample text information as a third test sample, testing the trained second large model using the third test sample to obtain a performance parameter of the trained second large model.
[0019] According to another aspect of embodiments of the present application, a text processing method is provided, the method comprising:
[0020] obtaining first prompt information and text information to be processed;
[0021] inputting the first prompt information and the text information to be processed into a trained second large model to obtain second key information in the text information to be processed output by the trained second large model;
[0022] wherein the trained second large model is trained by the model training method according to any one of the above.
[0023] According to another aspect of the embodiments of the present application, a model training apparatus is provided, the apparatus comprising:
[0024] a first processing module configured to input the first prompt information and the first sample text information into the first large model to obtain first key information in the first sample text information output by the first large model;
[0025] a first obtaining module configured to take the first prompt information, the first sample text information and the first key information as a first training sample, and take the first key information as a label corresponding to the first sample text information;
[0026] a training module configured to train a second large model to be trained by using the first training sample to obtain a trained second large model, a parameter quantity of the second large model being less than that of the first large model.
[0027] Optionally, the first prompt information comprises a text information analysis method, a text information analysis example and a key information output format.
[0028] Optionally, before the first prompt information and the first sample text information are input into the first large model, the apparatus further comprises an extraction module configured to extract first dialogue summary information corresponding to the first dialogue text information, and take the first dialogue summary information as the first sample text information.
[0029] Optionally, the apparatus further comprises a third obtaining module configured to take second prompt information, second sample text information and first rejection information corresponding to the second sample text information as a second training sample, and take the first rejection information as a label of the second sample text information; and the training module is specifically configured to train the second large model to be trained by using the first training sample and the second training sample.
[0030] Optionally, the apparatus further comprises a first testing module configured to take the second prompt information, third sample text information and second rejection information corresponding to the third sample text information as a first testing sample, take the second rejection information as a label of the third sample text information, test the trained second large model by using the first testing sample, and obtain a rejection information output accuracy rate corresponding to the trained second large model.
[0031] Optionally, the apparatus further comprises:
[0032] a second testing module configured to take the first prompt information and fourth sample text information as a second testing sample, test the trained second large model by using the second testing sample, and obtain a key information output pass rate corresponding to the trained second large model.
[0033] and / or,
[0034] a third test module, configured to take the first prompt information and the fifth sample text information as third test samples, test the second large model trained to obtain a performance parameter of the second large model trained.
[0035] According to another aspect of the embodiments of the present application, a text processing device is provided, the device comprising:
[0036] a second obtaining module, configured to obtain first prompt information and text information to be processed;
[0037] a second processing module, configured to input the first prompt information and the text information to be processed into a second large model trained to obtain second key information in the text information to be processed output by the second large model trained.
[0038] The second large model trained is obtained by the model training method according to any one of the preceding aspects.
[0039] According to another aspect of the embodiments of the present application, an electronic device is provided, comprising a processor and a computer readable storage medium, the computer readable storage medium storing a computer program; when the computer program is executed by the processor, the processor executes the model training method according to any one of the preceding aspects, or executes the text processing method according to any one of the preceding aspects.
[0040] According to another aspect of the embodiments of the present application, a computer readable storage medium is provided, the computer readable storage medium storing a computer program; when the computer program is executed by a processor, the processor executes the model training method according to any one of the preceding aspects, or executes the text processing method according to any one of the preceding aspects.
[0041] In the embodiment of the present application, the first prompt information and the first sample text information are input into the first large model to obtain the first key information in the first sample text information output by the first large model; the first prompt information, the first sample text information and the first key information are taken as the first training sample, and the first key information is taken as the label corresponding to the first sample text information; the second large model to be trained is trained by using the first training sample to obtain the trained second large model, and the parameter quantity of the second large model is less than that of the first large model. Therefore, in the embodiment of the present application, when training the second large model for extracting key information in text information, the first key information output by the first large model is used as a supervision signal to guide the training of the second large model, so that manual labeling of the training sample is not required, the parameter quantity of the second large model is less than that of the first large model, that is, the second large model is a lightweight model, thereby improving the scene adaptation degree of the second large model and reducing the demand of the second large model for deployment resources, so as to meet the needs of real-time scenarios and high-concurrency calling scenarios.
[0042] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the content of the specification can be implemented, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following will describe the specific embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some of the drawings of the present application, and other drawings can also be obtained by those skilled in the art without creating laborious work.
[0044] Figure 1 is a step flow chart of a model training method of an embodiment of the present application;
[0045] Figure 2 is a schematic diagram of a data preprocessing process of an embodiment of the present application;
[0046] Figure 3 is a schematic diagram of a model training process of an embodiment of the present application;
[0047] Figure 4 is a step flow chart of a text processing method of an embodiment of the present application;
[0048] Figure 5 is a schematic diagram of a hot word word cloud and a hot word ranking list of an embodiment of the present application;
[0049] Figure 6 is a schematic diagram of a keyword word cloud of an embodiment of the present application;
[0050] Figure 7 is a schematic diagram of a keyword trend of an embodiment of the present application;
[0051] Figure 8 is a structural block diagram of a model training device of an embodiment of the present application;
[0052] Figure 9 is a structural block diagram of a text processing device of an embodiment of the present application;
[0053] Figure 10 is a structural block diagram of an electronic device of an embodiment of the present application;
[0054] Figure 11 is a structural block diagram of a computer-readable storage medium of an embodiment of the present application. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0056] In the embodiments of the present application, the large model technology is used to realize the intelligent extraction of key information from text information. A second large model with low parameter quantity is proposed to replace a first large model with high parameter quantity, aiming to solve the problems of resource consumption, real-time performance and concurrent capability, the need for manual annotation, and low training efficiency when using a large model with high parameter quantity.
[0057] Referring to Figure 1 , a step flowchart of a model training method of an embodiment of the present application is shown.
[0058] As shown in Figure 1 , the model training method can include the following steps:
[0059] Step 101, input the first prompt information and the first sample text information into the first large model to obtain the first key information in the first sample text information output by the first large model.
[0060] In an embodiment of the present application, for the business scenario of extracting key information from text information, text information capable of being extracted key information under the business scenario is obtained as first sample text information. The first sample text information can be obtained according to relevant logs under the business scenario, and the first sample text information is text information that is expected to be extracted key information. The business scenario can be any applicable business scenario such as customer service traffic key information extraction.
[0061] Exemplarily, for the business scenario of customer service traffic key information extraction, historical dialogue text information between a customer service and a customer can be taken as first dialogue text information, and then first dialogue summary information corresponding to the first dialogue text information is extracted, and the first dialogue summary information is taken as the first sample text information. In this way, the first dialogue summary information can concisely and accurately summarize the first dialogue text information, so that the processing efficiency of the first large model and the second large model on the first sample text information is higher.
[0062] Exemplarily, the process of extracting the first dialogue summary information corresponding to the first dialogue text information can include: inputting the third prompt information and the first dialogue text information into a third large model, the third large model extracts summary information from the first dialogue text information according to the third prompt information, and the third large model outputs the first dialogue summary information corresponding to the first dialogue text information. The third prompt information is used to guide the third large model to extract summary information from the first dialogue text information, so as to output the first dialogue summary information corresponding to the first dialogue text information. The third large model can be any applicable large language model, and the present embodiment does not limit this.
[0063] In an embodiment of the present application, the first prompt information is set for the business scenario of extracting key information from text information. The first prompt information is used to guide a large model (the first large model or the second large model) to extract key information from input text information, so as to extract key information in the text information.
[0064] Exemplarily, the first prompt information can include text information analysis mode, text information analysis example, and key information output format, etc. The text information analysis mode is used to guide the large model to analyze the text information according to the text information mode, so that the analysis process of the large model is more standardized and specific; the text information analysis example is used to provide specific examples for the large model for reference in analysis, so as to further improve the accuracy of the analysis of the large model; and the key information output format is used to guide the large model to output the extracted key information according to the key information output format, so as to standardize the output format and facilitate subsequent use.
[0065] For example, for the business scenario of customer service call key information extraction, the following structured first prompt information can be set:
[0066] Please extract the "rootword" for the core needs and then extract the "keywords" for association. Please follow the steps I gave you and proceed with the task:
[0067] Text content: "%s"
[0068] When extracting and analyzing hot words, please follow the following steps:
[0069] 1. Perform text preprocessing, including removing punctuation and converting to uniform character encoding.
[0070] 2. Apply dependency algorithm technology to decompose the text into individual lexical units.
[0071] 3. Filter out common stop words in the text, such as "of", "and", "is", etc.
[0072] 4. Perform part-of-speech recognition to identify nouns, verbs, etc.
[0073] 5. Pay special attention to the subject-object structure formed by nouns and verbs, as well as their dependency relationships.
[0074] 6. According to word frequency, dependency relationship, and text relevance, output "rootword". Note that "rootword" must be a verb-object word, and it must be related to communication services and only select one group, such as "renew the bill".
[0075] 7. The selected rootword should be as specific as possible, avoiding general situations such as "cancel service". Please be specific, such as "cancel MMS service" or "cancel TV service".
[0076] 8. Only output the most common "rootword" and 1-3 associated "keywords" containing "rootword", such as "send bill" and "bill reminder". Note that the association should be displayed from high to low.
[0077] 9. Please strictly follow the steps I gave you and proceed with the task:
[0078] 10. See the following example content: If the text information is "The user reflects that the bill has not been received, and the customer service representative indicates that the bill will be re-sent again, and confirms the user's delivery address and name. The user asks if it needs to be reminded every month, and the customer service representative confirms that the monthly bill has been set. The user is worried that the bill has not been received, and the customer service representative promises to re-sent it again and suggests that the user pay attention. At the end of the call, the customer service representative asks if there are any other inquiries and thanks for the call." For this text content, I expect the output result to be "rootword": "re-sent bill", "keywords": ["send bills", "bill reminders"].
[0079] 11. Your output should only contain JSON (JavaScript Object Notation), and the output value uses Chinese, value1 represents "rootword", and value2 represents "keywords"
[0080] Output format:
[0081] ```json
[0082] {
[0083] \"rootword\" : \"value1\"
[0084] \"keywords\" : [\"value2\"]
[0085] }
[0086] ```
[0087] Through the steps in the above first prompt information, the hot words (rootword) and 1 to 3 groups of keywords (keywords) in the text information can be effectively extracted, and the output is required to be in a standard JSON format, which is convenient for subsequent business interfaces to directly call.
[0088] It should be noted that the hot words and keywords mentioned above are the extracted key information. The first prompt information described above is only used for example illustration and does not limit the embodiments of the present application.
[0089] After obtaining the first prompt information and the first sample text information, the first prompt information and the first sample text information are input into the first large model, and the first large model extracts the key information from the first sample text information under the guidance of the first prompt information, that is, the first key information in the first sample text information output by the first large model is obtained.
[0090] In the embodiments of the present application, the first large model can be any applicable large language model, and the present embodiments do not limit this. For example, a high-parameter large model (such as a 32B large model) can be selected, which is a model after instruction optimization and quantization. The high-parameter large model has good performance for the first prompt information, and in actual business deployment, compared with a non-quantized model, it is more resource-saving, faster in inference speed, and stronger in real-time performance.
[0091] In step 102, the first prompt information, the first sample text information and the first key information are taken as first training samples, and the first key information is taken as a label corresponding to the first sample text information.
[0092] For each first sample text information, the first prompt information and the first sample text information are input into the first large model to obtain the first key information in the first sample text information output by the first large model. After that, the first prompt information, the first sample text information and the first key information can be taken as first training samples, and the first key information can be taken as a label corresponding to the first sample text information. The label refers to the key information that the first sample text information hopes to extract.
[0093] In the embodiments of the present application, in order to further improve the accuracy of the data, a large amount of real running logs of the first large model can be collected, the input and output of the first large model can be retained, and a table can be made as source data, and data preprocessing can be performed on these source data.
[0094] Referring to Figure 2 , a schematic diagram of a data preprocessing process of an embodiment of the present application is shown.
[0095] As Figure 2 indicated, the data preprocessing process can include:
[0096] Format conversion:
[0097] The output of the first large model is parsed in JSON format, and data that cannot be parsed or fails to be interpreted is deleted. For example, rows with failed interface calls (i.e., output as error) are removed. The interface call failure cases include call timeout, call request number exceeding the concurrency limit, and abnormality caused by network fluctuations.
[0098] Data cleaning:
[0099] For the output that can be parsed in JSON format, it is verified whether the two keys of the output JSON are rootword and keywords. If not, it is cleaned. Sometimes, the model output has hallucinations, which can cause this situation, resulting in unsuccessful interface calls.
[0100] For the output that can be parsed by the JSON format, it is verified whether the rootword quantity and the keywords quantity specified in the prompt information in the JSON of the output are met, and if not, the output is cleaned.
[0101] For the output that can be parsed by the JSON format, the data with empty or NULL rootword and keywords due to insufficient understanding ability of the model is cleaned.
[0102] It should be noted that other data preprocessing methods can also be used in the embodiments of the present application, which are not limited.
[0103] After the above data preprocessing steps, high-quality first sample text information and first key information can be obtained, and the first prompt information, the first sample text information and the first key information after data preprocessing are used as first training samples.
[0104] For example, the JSON format template of the first training sample is as follows:
[0105] {"system":"0","query":"1","response":"2"}
[0106] Among them, "system" represents the first prompt information, "query" represents the first sample text information, and "response" represents the first key information. Through the code program, each row of data cleaned from the table is filled into the above template, and then spliced into the first training sample data set file.
[0107] In step 103, the first training sample is used to train the second large model to be trained, and a trained second large model is obtained, and the parameter quantity of the second large model is less than that of the first large model.
[0108] In the embodiments of the present application, the first training sample is used to train the second large model to be trained, and a trained second large model is obtained. The parameter quantity of the second large model is less than that of the first large model.
[0109] First, the second large model is selected. In the embodiments of the present application, the second large model can be any applicable large language model, and the present embodiment does not limit it. The parameter quantity of the second large model is less than that of the first large model, but the parameter quantity cannot be too small, otherwise it is difficult to learn true knowledge in the model fine-tuning stage, for example, a 7B large model can be selected.
[0110] Then, the training parameter setting is performed. In the embodiment of the application, any applicable training framework can be used, the fine-tuning method can adopt supervised fine-tuning (SFT), the fine-tuning strategy can be set (such as Lora fine-tuning strategy), the learning rate can be set (such as 1e-4), the batch size can be set (such as 1), the number of training rounds can be set (such as 4), and the like. The dataset participating in the training is the first training sample described above, the maximum truncation length of the input sequence is set (such as 4096), the hyperparameters of the fine-tuning strategy are set (such as low-rank matrix lora_rank=8 and scaling factor lora_alpha=32), the cosine annealing learning rate decay method is adopted, and the like.
[0111] Finally, the model training is performed and the training process monitoring is performed. In the process of training the second large model to be trained by using the first training sample, the first prompt information and the first sample text information are input into the second large model to be trained, the second large model to be trained extracts the third key information in the first sample text information according to the first prompt information, obtains the third key information output by the second large model to be trained, calculates a loss function according to the first key information in the first sample text information and the third key information in the first sample text information, adjusts the parameters of the second large model to be trained for continuous training if the loss function does not meet a target condition, and stops until the loss function meets the target condition, thereby obtaining the second large model after training. During the training, tools such as Tensorboard can be used for training process monitoring. When it is observed that the loss function curve of the training converges to a smaller fluctuation, the training is interrupted in time to prevent overfitting.
[0112] The loss function can be selected from any applicable form such as a cross-entropy loss function and a logarithmic loss function. The target condition can be that the loss function is less than a set loss threshold or that the loss function converges.
[0113] In the embodiment of the application, when the second large model for extracting key information in text information is trained, the first key information output by the first large model is used as a supervision signal to guide the training of the second large model, so that manual annotation of the training sample is not required, the parameter quantity of the second large model is less than that of the first large model, that is, the second large model is a lightweight model, thereby improving the scene adaptation degree of the second large model and reducing the demand for deployment resources of the second large model, so as to meet the needs of real-time scenarios and high-concurrency calling scenarios.
[0114] In an optional implementation, before the training of the second large model using the first training sample, the method further includes: taking the second prompt information, the second sample text information, and the first rejection information corresponding to the second sample text information as a second training sample, and taking the first rejection information as a label of the second sample text information. Correspondingly, the process of training the second large model using the first training sample includes: training the second large model using the first training sample and the second training sample. The second prompt information is used to guide the second large model to analyze the input text information to obtain the rejection information.
[0115] In the embodiments of the present application, it is considered that the text information may contain some sensitive information. For such text information, in order to ensure information security, the second large model can be trained to output rejection information for such text information. Therefore, the second prompt information, the second sample text information, and the first rejection information corresponding to the second sample text information can be obtained, and the second sample text information is the text information that is expected to be output as rejection information.
[0116] In the implementation, the second sample text information and the first rejection information corresponding to the second sample text information can be manually written or written by using a tool, or collected from related open source information.
[0117] In the embodiment of the present application, in the process of training the second large model to be trained by using the first training sample and the second training sample, for the first training sample, the first prompt information and the first sample text information are input into the second large model to be trained, the second large model to be trained extracts key information from the first sample text according to the first prompt information, and obtains third key information in the first sample text information output by the second large model to be trained; for the second training sample, the second prompt information and the second sample text information are input into the second large model to be trained, the second large model to be trained analyzes the second sample text according to the second prompt information, and obtains third refusal information corresponding to the second sample text information output by the second large model to be trained, and according to the first key information in the first sample text information and the third key information in the first sample text information, and the first refusal information corresponding to the second sample text information and the third refusal information corresponding to the second sample text information, a loss function is calculated, if the loss function does not meet the target condition, the parameters of the second large model to be trained are adjusted for continuous training until the loss function meets the target condition, and the second large model trained is obtained. Wherein, the first key information and the third key information are a group of loss function calculation, and the first refusal information and the third refusal information are a group of loss function calculation.
[0118] In the embodiment of the present application, the second large model trained can be further tested to analyze the use effect of the second large model.
[0119] In an optional implementation, the second large model trained is tested for business capability.
[0120] Exemplarily, fourth sample text information similar to the first sample text information can be obtained, and specifically, the second dialogue summary information corresponding to the second dialogue text information can be extracted, and the second dialogue summary information is taken as the fourth sample text information. The first prompt information and the fourth sample text information are taken as a second test sample, the second large model trained is tested by using the second test sample, and a key information output qualified rate corresponding to the second large model trained is obtained.
[0121] Exemplarily, the first prompt information and the fourth sample text information can be input into the second large model trained, the second large model trained extracts key information from the fourth sample text information according to the first prompt information, obtains fourth key information in the fourth sample text information output by the second large model trained, judges whether the fourth key information in the fourth sample text information is qualified, calculates a ratio of qualified fourth key information to all fourth key information, and obtains a key information output qualified rate corresponding to the second large model trained.
[0122] The process of judging whether the fourth key information in the fourth sample text information is qualified can include judging whether the format of the fourth key information conforms to a key information output format contained in the first prompt information, and if so, determining that the fourth key information is qualified, otherwise, determining that the fourth key information is not qualified.
[0123] For example, the fourth sample text information not participating in training is taken, the second large model trained is used to perform batch inference according to the first prompt information and the fourth sample text information, and fourth key information in the fourth sample text information is output. For the fourth key information, first, JSON format is parsed, if the parsing is passed, and two keys of JSON are rootword and keywords, and keywords are 1 to 3 groups, and rootword and keywords values are not empty or NULL, it is indicated that the fourth key information is qualified.
[0124] In an optional implementation, the second large model trained is subjected to deployment performance testing.
[0125] Exemplarily, fifth sample text information similar to the first sample text information can be obtained, and specifically, third dialogue summary information corresponding to the third dialogue text information can be extracted, and the third dialogue summary information is taken as the fifth sample text information. The first prompt information and the fifth sample text information are taken as a third test sample, the second large model trained is tested by using the third test sample, and a performance parameter of the second large model trained is obtained.
[0126] Exemplarily, the first prompt information and the fifth sample text information can be input into the second large model trained, the second large model trained extracts key information from the fifth sample text information according to the first prompt information, obtains fifth key information in the fifth sample text information output by the second large model trained, and calculates performance parameters such as resource consumption parameters, average calling time, limit concurrency, and throughput in the process of calling the second large model trained, so as to analyze the deployment performance of the second large model trained.
[0127] In an optional implementation, the second large model trained is subjected to security capability testing.
[0128] Exemplarily, third sample text information and second rejection information corresponding to the third sample text information can be obtained, the third sample text information is similar to the second sample text information, and the second rejection information is similar to the first rejection information. The second prompt information, the third sample text information, and the second rejection information corresponding to the third sample text information are taken as a first test sample, and the second rejection information is taken as a label of the third sample text information. The first test sample is used to test the second large model trained, and an output accuracy rate of rejection information corresponding to the second large model trained is obtained.
[0129] Exemplarily, the second prompt information and the third sample text information can be input into the second large model trained, the second large model trained analyzes rejection information of the third sample text information according to the second prompt information, obtains fourth rejection information corresponding to the third sample text information output by the second large model trained, calculates an error between the second rejection information corresponding to the third sample text information and the fourth rejection information corresponding to the third sample text information, and then obtains the output accuracy rate of rejection information corresponding to the second large model trained according to an average value of errors corresponding to all fourth sample text information.
[0130] The security capability testing is mainly to resist prompt injection and data leakage attacks, guarantee model system reliability, help enterprises meet data security regulation requirements, and avoid negative impact on society.
[0131] Referring to Figure 3 , a schematic diagram of a model training process of an embodiment of the present application is shown. Figure 3 Take the business scenario of key information extraction of customer service traffic as an example for description.
[0132] As Figure 3 shown, the model training process includes:
[0133] Prompt information design and data preprocessing: design the first structured prompt information corresponding to the key information extraction business scenario of customer service, select a first large model with high parameter quantity, run the first large model to generate a running log, and perform data preprocessing on the running log.
[0134] Model training: according to the preprocessed data, a first training sample is extracted, and the second large model with low parameter quantity is fine-tuned and trained using the first training sample.
[0135] Post-training test verification: the second large model after training is tested for business capability, security capability, and deployment performance.
[0136] For specific introduction of each process, please refer to the relevant description above.
[0137] In the embodiments of the present application, the second large model is trained by model fine-tuning and knowledge distillation. Model fine-tuning refers to a technical solution that collects professional field data to optimize model parameters locally or globally based on a pre-trained large model to adapt to specific task requirements. Its core lies in adjusting key parameter layers (such as attention weights and fully connected layers) to improve task precision while preserving the general representation ability of the model. Typical application scenarios include natural language processing and image classification. Knowledge distillation guides the training of a lightweight model (student model) by using the output of a large model (teacher model) as a supervision signal to achieve knowledge transfer. This technology can reduce model computational complexity while maintaining high precision, and is commonly used in efficient inference in resource-constrained scenarios.
[0138] Referring to Figure 4 , a step flowchart of a text processing method according to an embodiment of the present application is shown.
[0139] As shown in Figure 4 , the text processing method can include the following steps:
[0140] Step 401, obtaining first prompt information and text information to be processed.
[0141] The first prompt information is the same as the first prompt information in the model training process.
[0142] The text information to be processed is similar to the first sample text information in the model training process. The dialogue text information between the customer service and the customer can be obtained as the dialogue text information to be processed, and then the dialogue summary information corresponding to the dialogue text information to be processed is extracted, and the dialogue summary information to be processed is taken as the text information to be processed.
[0143] Step 402, input the first prompt information and the to-be-processed text information into the trained second large model, to obtain second key information in the to-be-processed text information output by the trained second large model.
[0144] The first prompt information and the to-be-processed text information are input into the trained second large model, and the trained second large model extracts key information from the to-be-processed text information under the guidance of the first prompt information, that is, the second key information in the to-be-processed text information output by the trained second large model is obtained.
[0145] For example, for the business scenario of customer service call key information extraction, the speech recognition is first performed on each call between the customer service and the customer, the to-be-processed call text information is recognized, and then the to-be-processed call summary information corresponding to the to-be-processed call text information is extracted by combining the large model, as the to-be-processed text information. The trained second large model is used as a base model, the to-be-processed text information is input through the designed first prompt information, and the second key information in the to-be-processed text information is generated, which includes rootword (hot word) and keywords (key word).
[0146] In the embodiment of the application, the rootword (hot word) and the keywords (key word) can also be visualized.
[0147] For example, all hot words of the day can be collected to form a hot word cloud chart and a hot word ranking list, so as to analyze the customer service call hot words in real time and understand the call hot word trend of the day. The schematic diagram of the hot word cloud chart and the hot word ranking list is as shown in Figure 5 Figure 5 The left side of the hot word cloud chart is the hot word cloud chart, and the right side is the hot word ranking list.
[0148] For example, on the basis of the hot word cloud chart and the hot word ranking list described above, any hot word can be clicked to view the subordinate keyword cloud chart. The schematic diagram of the keyword cloud chart is as shown in Figure 6 Figure 6 For example, the keyword cloud chart after clicking the hot word "repair" is taken as an example.
[0149] For example, on the basis of the keyword cloud chart described above, the keyword trend in 24 hours and 30 days can be displayed according to the statistical analysis of the keyword cloud. The schematic diagram of the keyword trend is as shown in Figure 7 Figure 7 The keyword quantity (vertical axis) of each time point (horizontal axis) is displayed.
[0150] In the embodiment of the application, after the traffic hotspot extraction using the second large model, traffic hotspot analysis can be performed, customer concerns can be automatically identified, demand changes can be quickly captured, product and service optimization of enterprises can be assisted, core business and product problems can be refined from scattered traffic hotspots, and product improvement can be accurately focused on. At the same time, traffic hotspot trend analysis can also be performed, customer trends can be obtained according to recent hotspot words, hotspot trends can be analyzed, timely warning can be realized, and quick processing and hotspot feedback can be performed.
[0151] Referring to Figure 8 , a structural block diagram of a model training device in an embodiment of the application is shown.
[0152] As shown in Figure 8 , the model training device can include the following modules:
[0153] The first processing module 801 is configured to input the first prompt information and the first sample text information into the first large model to obtain first key information in the first sample text information output by the first large model;
[0154] The first obtaining module 802 is configured to use the first prompt information, the first sample text information, and the first key information as a first training sample, and use the first key information as a label corresponding to the first sample text information.
[0155] The training module 803 is configured to train a second large model to be trained using the first training sample to obtain a trained second large model, and the parameter quantity of the second large model is less than that of the first large model.
[0156] Optionally, the first prompt information includes a text information analysis method, a text information analysis example, and a key information output format.
[0157] Optionally, before the first prompt information and the first sample text information are input into the first large model, the device further includes an extraction module configured to extract first dialogue summary information corresponding to the first dialogue text information and use the first dialogue summary information as the first sample text information.
[0158] Optionally, the device further includes a third obtaining module configured to use second prompt information, second sample text information, and first rejection information corresponding to the second sample text information as a second training sample, and use the first rejection information as a label of the second sample text information; and the training module 803 is specifically configured to train the second large model to be trained using the first training sample and the second training sample.
[0159] Optionally, the apparatus further comprises a first test module configured to use the second prompt information, third sample text information, and second rejection information corresponding to the third sample text information as first test samples, the second rejection information as a label of the third sample text information; test the trained second large model using the first test samples to obtain an output accuracy rate of the rejection information corresponding to the trained second large model.
[0160] Optionally, the apparatus further comprises:
[0161] a second test module configured to use the first prompt information and fourth sample text information as second test samples, test the trained second large model using the second test samples to obtain an output qualified rate of the key information corresponding to the trained second large model;
[0162] and / or,
[0163] a third test module configured to use the first prompt information and fifth sample text information as third test samples, test the trained second large model using the third test samples to obtain a performance parameter of the trained second large model.
[0164] Referring to Figure 9 , a structural block diagram of a text processing apparatus according to an embodiment of the present application is shown.
[0165] As Figure 9 shown, the text processing apparatus can comprise the following modules:
[0166] a second obtaining module 901 configured to obtain first prompt information and to-be-processed text information;
[0167] a second processing module 902 configured to input the first prompt information and the to-be-processed text information into a trained second large model to obtain second key information in the to-be-processed text information output by the trained second large model;
[0168] The trained second large model is obtained by the model training method according to any one of the above embodiments.
[0169] In the embodiments of the present application, when training the second large model for extracting key information in text information, the first key information output by the first large model is used as a supervision signal to guide the training of the second large model, so that manual labeling of training samples is not required, and the parameter quantity of the second large model is smaller than that of the first large model, that is, the second large model is a lightweight model, thereby improving the scene adaptation degree of the second large model, reducing the demand for deployment resources of the second large model, and thus meeting the needs of real-time scenarios and high-concurrency calling scenarios.
[0170] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant part can be referred to the part of the method embodiment.
[0171] Referring to Figure 10 , a structural block diagram of an electronic device of an embodiment of the present application is shown. As shown in Figure 10 , the electronic device 11 includes a processor 111 and a computer readable storage medium 112, and the computer readable storage medium 112 stores a computer program 1121.
[0172] The processor 111 is configured to execute the computer program 1121 stored in the computer readable storage medium 112, and the processor 111 executes the model training method according to any one of the above embodiments or executes the text processing method according to any one of the above embodiments when executing the computer program 1121, and the same technical effects can be achieved. To avoid repetition, it will not be described here.
[0173] The processor 111 mentioned above can include but is not limited to a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and the like.
[0174] The computer readable storage medium 112 mentioned above can include but is not limited to a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), an electrically erasable programmable read-only memory (EEPROM), a hard disk, a floppy disk, a flash memory, and the like.
[0175] Referring to Figure 11 , a structural block diagram of a computer readable storage medium of an embodiment of the present application is shown. As shown in Figure 11As shown, the computer readable storage medium 21 stores a computer program 211, which can be executed by a processor of an electronic device, and when the computer program 211 is executed by the processor, the processor executes the model training method according to any one of the above embodiments, or executes the text processing method according to any one of the above embodiments, and achieves the same technical effects. To avoid repetition, details are not described here.
[0176] The various embodiments in the specification are interrelated, and each embodiment is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.
[0177] It should be noted that all actions of obtaining signals, information or data in the present application are carried out in accordance with the corresponding data protection regulations and policies of the place, and with the authorization of the owner of the corresponding device.
[0178] It should be noted that in this document, relational terms such as first and second and the like are used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or terminal device. Without further limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0179] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium, includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application.
[0180] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims.
[0181] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0182] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0183] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0184] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected to achieve the purpose of the embodiments of the present application according to actual needs.
[0185] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0186] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. In summary, the content of the present specification should not be understood as the limitation of the present application.
Claims
1. A model training method, characterized in that, The method comprises: inputting the first prompt information and the first sample text information into the first large model to obtain first key information in the first sample text information output by the first large model; using the first training sample to train the second large model to be trained to obtain the second large model after training, and the parameter quantity of the second large model is less than that of the first large model. The first prompt information includes text information analysis method, text information analysis example and key information output format.
2. The method of claim 1, wherein, Before the first prompt information and the first sample text information are input into the first large model, the method further comprises:
3. The method of claim 1, wherein, extracting first dialogue summary information corresponding to the first dialogue text information, and taking the first dialogue summary information as the first sample text information.
4. The method of claim 1, wherein, Before the first training sample is used to train the second large model to be trained, the method further comprises: taking the second prompt information, the second sample text information and the first rejection information corresponding to the second sample text information as a second training sample, and the first rejection information as a label of the second sample text information; The training of the second large model to be trained using the first training sample comprises training the second large model to be trained using the first training sample and the second training sample. The method further comprises:
5. The method of claim 4, wherein, taking the second prompt information, the third sample text information and the second rejection information corresponding to the third sample text information as a first test sample, and the second rejection information as a label of the third sample text information; using the first test sample to test the second large model after training to obtain the rejection information output accuracy rate corresponding to the second large model after training. The method further comprises:
6. The method of claim 1, wherein, taking the first prompt information and the fourth sample text information as a second test sample, using the second test sample to test the second large model after training to obtain the key information output qualified rate corresponding to the second large model after training; and / or, taking the first prompt information and the fifth sample text information as a third test sample, using the third test sample to test the second large model after training to obtain the performance parameter of the second large model after training. The method comprises:
7. A text processing method characterized by, obtaining first prompt information and text information to be processed; inputting the first prompt information and the text information to be processed into the second large model after training to obtain second key information in the text information to be processed output by the second large model after training; wherein the second large model after training is obtained by the model training method of any one of claims 1 to 6. The device comprises:
8. A model training apparatus, comprising: The first processing module is configured to input the first prompt information and the first sample text information into a first large model to obtain first key information in the first sample text information output by the first large model; The first obtaining module is configured to take the first prompt information, the first sample text information, and the first key information as a first training sample, and take the first key information as a label corresponding to the first sample text information; The training module is configured to train a second large model to be trained by using the first training sample to obtain a trained second large model, and a parameter quantity of the second large model is less than that of the first large model.
9. A text processing apparatus characterized by comprising: The device comprises: The second obtaining module is configured to obtain first prompt information and text information to be processed; The second processing module is configured to input the first prompt information and the text information to be processed into the trained second large model to obtain second key information in the text information to be processed output by the trained second large model; The trained second large model is obtained by the model training method in any one of claims 1 to 6.
10. An electronic device, comprising: The electronic device comprises a processor and a computer readable storage medium, and the computer readable storage medium stores a computer program; When the computer program is executed by the processor, the processor executes the model training method in any one of claims 1 to 6, or executes the text processing method in claim 7.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and when the computer program is executed by the processor, the processor executes the model training method in any one of claims 1 to 6, or executes the text processing method in claim 7.