Dialogue processing method and apparatus, and electronic device
By performing machine learning and channel importance evaluation on the historical dialogue dataset and iterative pruning combined with preset pruning rules, the existing dialogue model structure complex and parameter redundant is solved, and efficient and accurate dialogue prediction is achieved.
Patent Information
- Application Number
- PCT/CN2024/135306
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-08
AI Technical Summary
The dialogue model in existing intelligent online customer service has complex structure, with parameter redundancy and redundant channels, which leads to low efficiency and low accuracy of dialogue prediction. In addition, artificial intervention is required to determine the pruning rate when pruning the model, resulting in inaccuracy.
By machine learning the initial dialogue model based on the historical dialogue dataset, the first dialogue model is obtained, and the redundant channels are determined based on the channel importance value in each network layer, and iteratively pruning is performed according to the preset pruning rules to obtain the target dialogue model.
The dialogue model is streamlined and optimized, the dialogue prediction efficiency and accuracy are improved, and the over-pruning and under-pruning problems caused by human intervention are avoided. While maintaining the model accuracy, the parameter quantity and floating-point calculation quantity are greatly compressed.
Smart Images

Figure CN2024135306_08052025_PF_FP_ABST
Abstract
Description
Dialogue processing method, device and electronic equipment
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 2023114258802, filed on October 30, 2023, entitled “Dialogue Processing Method, Device and Electronic Device,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the field of artificial intelligence, and more specifically, to a method, device, and electronic device for processing conversations. Background Art
[0004] Currently, conversational models in the field of intelligent online customer service are able to intelligently understand user intent, interact in real time based on user needs, and provide specific textual responses to user questions. However, these conversational models often have complex structures, contain a certain degree of parameter redundancy and / or redundant channels, and require manual intervention during model pruning to determine the pruning rate. This can lead to inaccurate conversational model predictions, resulting in low conversational prediction efficiency and accuracy.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0006] Embodiments of the present application provide a conversation processing method, device, and electronic device.
[0007] According to one aspect of an embodiment of the present application, a conversation processing method is provided, comprising: performing machine learning on an initial conversation model based on a historical conversation dataset to obtain a first conversation model, wherein the first conversation model includes multiple network layers, each network layer includes multiple channels, and the multiple channels respectively correspond to importance values, wherein the importance values are used to indicate the degree of importance of the corresponding channels, and the historical conversation dataset includes multiple groups of historical input texts and historical reply texts corresponding to the multiple groups of historical input texts; based on the importance values corresponding to the multiple channels included in each network layer, redundant channels in each network layer are determined; and redundant channels in each network layer are iteratively pruned according to preset pruning rules to obtain a target conversation model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the proportion of the number of pruned parameters to the total number of parameters in the corresponding channel.
[0008] Optionally, the performing machine learning on the initial dialogue model based on the historical dialogue dataset to obtain the first dialogue model includes: determining an initial loss function; adding an L1 regularization term for a target hyperparameter on the basis of the initial loss function to obtain a target loss function, wherein the target hyperparameter is located in a batch normalization layer in each network layer; and performing machine learning on the initial dialogue model based on the historical dialogue dataset and the target loss function to obtain the first dialogue model.
[0009] Optionally, the performing machine learning on the initial dialogue model based on the historical dialogue dataset and the target loss function to obtain the first dialogue model includes: performing machine learning on the initial dialogue model based on the historical dialogue dataset and the target loss function; and outputting the first dialogue model when the importance values corresponding to a predetermined number of channels among the multiple channels included in each network layer are within a preset interval.
[0010] Optionally, the iterative pruning of redundant channels in each network layer according to preset pruning rules to obtain a target dialogue model includes: iteratively pruning redundant channels in each network layer according to the preset pruning rules; and outputting the target dialogue model when all network layers in the first dialogue model are pruned and preset pruning constraints are met.
[0011] Optionally, the iterative pruning of redundant channels in each network layer according to the preset pruning rule includes: taking each network layer in the first dialogue model as the current network layer in turn according to a preset pruning order, and looping through the following operations until all network layers in the first dialogue model are pruned and the preset pruning constraints are satisfied: determining an initial pruning rate corresponding to the current network layer based on the preset pruning rule; pruning the current network layer according to the initial pruning rate to obtain a pruned dialogue model; obtaining a model accuracy loss based on the model accuracy of the pruned dialogue model and the model accuracy corresponding to the first dialogue model; adjusting the initial pruning rate according to the model accuracy loss to obtain a new pruning rate; and continuing to prune the pruned dialogue model based on the new pruning rate in the same manner as the above operation.
[0012] Optionally, the initial pruning rate is adjusted according to the model accuracy loss to obtain a new pruning rate, including: when the model accuracy loss is less than or equal to a preset accuracy loss tolerance value, increasing the initial pruning rate according to a preset ratio to obtain the new pruning rate; when the model accuracy loss is greater than the preset accuracy loss tolerance value, using the initial pruning rate as the new pruning rate.
[0013] Optionally, the preset pruning constraint is:
[0014] Among them, S i represents the parameter amount of any one of the multiple network layers, S represents the total parameter amount in the first dialogue model, L represents the number of the multiple network layers, i represents any one of the multiple network layers, p g is the preset global pruning rate.
[0015] Optionally, the method further includes: obtaining a text to be input; and obtaining a target reply text based on the text to be input and using the target dialogue model.
[0016] According to another aspect of an embodiment of the present application, a dialogue processing device is also provided, including: a machine learning module, configured to perform machine learning on an initial dialogue model based on a historical dialogue dataset to obtain a first dialogue model, wherein the first dialogue model includes multiple network layers, each network layer includes multiple channels, and the multiple channels respectively correspond to importance values, wherein the importance values are used to indicate the importance of the corresponding channels, and the historical dialogue dataset includes multiple groups of historical input texts and historical reply texts corresponding to the multiple groups of historical input texts; a determination module, configured to determine redundant channels in each network layer based on the importance values corresponding to the multiple channels included in each network layer; and a pruning module, configured to iteratively prune the redundant channels in each network layer according to preset pruning rules to obtain a target dialogue model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the proportion of the number of pruned parameters to the total number of parameters in the corresponding channel.
[0017] According to another aspect of an embodiment of the present application, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the dialogue processing methods described. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.
[0019] FIG1 is a schematic diagram of a conversation processing method provided according to some embodiments of the present application;
[0020] FIG2 is a schematic diagram of an optional conversation processing method provided according to some embodiments of the present application;
[0021] FIG3 is a schematic diagram of another optional conversation processing method provided according to some embodiments of the present application;
[0022] FIG4 is a schematic diagram of another optional conversation processing method provided according to some embodiments of the present application;
[0023] FIG5 is a schematic diagram of a conversation processing device provided according to some embodiments of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] According to an embodiment of the present application, a method embodiment of dialogue processing is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0027] FIG1 is a flow chart of a method for processing a conversation according to an embodiment of the present application. As shown in FIG1 , the method includes the following steps:
[0028] Step S102: Based on the historical conversation dataset, machine learning is performed on the initial conversation model to obtain a first conversation model, wherein the first conversation model includes multiple network layers, each network layer includes multiple channels, and the multiple channels respectively correspond to importance values, wherein the importance values are used to indicate the importance of the corresponding channels, and the historical conversation dataset includes multiple groups of historical input texts and historical reply texts corresponding to the multiple groups of historical input texts.
[0029] Optionally, an initial conversation model can be sparsely trained based on a historical conversation dataset to obtain a first conversation model. This initial conversation model can be a TextCNN text classification model, which includes multiple network layers, each followed by a batch normalization layer. The importance value of each channel can be determined based on the target hyperparameter γ in the batch normalization layer. Through sparse training, the target hyperparameter of the batch normalization layer (BN) gradually approaches 0 during sparse training, resolving the problem in which the BN layer weights in conventionally trained models tend to be too close to 0, facilitating subsequent channel screening.
[0030] In an optional embodiment, based on the historical conversation dataset, machine learning is performed on the initial conversation model to obtain a first conversation model, including: determining an initial loss function; adding an L1 regularization term for a target hyperparameter on the basis of the initial loss function to obtain a target loss function, wherein the target hyperparameter is located in a batch normalization layer in each network layer; based on the historical conversation dataset and the target loss function, machine learning is performed on the initial conversation model to obtain the first conversation model.
[0031] Optionally, the importance values corresponding to the multiple channels included in each network layer are determined based on the target hyperparameters in the batch normalization layer in each network layer. The initial loss function can be a cross-entropy loss function, which uses the target hyperparameter γ (i.e., the scaling factor) of the BN layer as a parameter to measure the importance of the channel. The closer the target hyperparameter γ is to 0, the smaller the importance of the channel, and vice versa. An L1 regularization term for γ is added to the initial loss function. After adding the L1 regularization term, training is performed again until the loss converges, and the first dialogue model is output. At the end of training, the γ value of the BN layer corresponding to the first dialogue model is determined as the importance value of each channel to facilitate subsequent channel screening.
[0032] In an optional embodiment, machine learning is performed on the initial dialogue model based on the historical dialogue dataset and the target loss function to obtain the first dialogue model, including: performing machine learning on the initial dialogue model based on the historical dialogue dataset and the target loss function; and outputting the first dialogue model when importance values corresponding to a predetermined number of channels among multiple channels included in each network layer are within a preset range.
[0033] Optionally, the preset interval range may be a certain interval range near 0, that is, when the target hyperparameters of the batch normalization layer (BN) layer gradually approach 0 during sparse training, the first dialogue model is output to solve the problem that the weights of the BN layer of the conventionally trained model will not be too close to 0.
[0034] Step S104 : determining redundant channels in each network layer based on importance values corresponding to the multiple channels included in each network layer.
[0035] Optionally, the smaller the importance value and the closer it is to 0, the less important the channel is in the model, and the parameters in the channel can be appropriately pruned to streamline the model structure. Channels with importance values less than a preset importance threshold (e.g., close to 0) can be treated as redundant channels and their parameters can be appropriately pruned to streamline the model structure.
[0036] In step S106, redundant channels in each network layer are iteratively pruned according to preset pruning rules to obtain a target dialogue model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the ratio of the number of pruned parameters to the total number of parameters in the corresponding channel.
[0037] Optionally, the preset pruning rules specify a method for dynamically adjusting the pruning rate. Specifically, during the iterative pruning of each network layer in the first network model, the pruning rate dynamically changes with the model accuracy corresponding to each round of pruning. In this manner, during the model pruning process, the pruning rate is dynamically adjusted as the model accuracy changes, thereby improving the efficiency and effectiveness of model pruning.
[0038] It should be noted that in the related art, when pruning a model, manual intervention is required to adjust the pruning rate. By manually specifying the importance threshold of the network channel in advance, channels below the set threshold will be removed. However, setting the channel importance threshold based on experience often requires a lot of manpower, and different network layers have different sensitivities to pruning. Therefore, setting the threshold too high may cause over-pruning, resulting in significant accuracy loss; setting the threshold too low may cause under-pruning, resulting in suboptimal pruning results, and the inability to completely prune redundant parameters, affecting model compression efficiency. Based on this, the embodiment of the present application first accurately evaluates the importance of the channels included in each layer of the network, and then uses this technology to dynamically determine the optimal pruning rate for each channel based on the model accuracy corresponding to each round of pruning. Without manual intervention, the tedious process of manually setting thresholds in existing methods is eliminated, and redundant parameters are eliminated. To a certain extent, the problems of over-pruning and under-pruning in existing methods are avoided, while maintaining accuracy. The number of parameters and floating-point calculations of the first dialogue model are significantly reduced, achieving inference acceleration and improving the efficiency of solving user problems.
[0039] In an optional embodiment, redundant channels in each network layer are iteratively pruned according to preset pruning rules to obtain a target dialogue model, including: iteratively pruning redundant channels in each network layer according to preset pruning rules; and outputting the target dialogue model when all network layers in the first dialogue model have been pruned and preset pruning constraints are met.
[0040] Optionally, preset pruning constraints are:
[0041] Among them, S i represents the parameter amount of any network layer in multiple network layers, S represents the total parameter amount in the first dialogue model, L represents the number of multiple network layers, i represents any network layer in multiple network layers, p g is the preset global pruning rate.
[0042] Optionally, the pruning rate of each channel in each network layer is adaptively adjusted according to preset pruning rules, removing redundant channels from each channel. Once all redundant channels in each layer have been pruned and the pruned model meets the preset pruning conditions, the final target dialogue model is output. By setting these constraints, the pruned model can be made more suitable for real-world scenarios while ensuring maximum model compression and optimal streamlining.
[0043] In an optional embodiment, redundant channels in each network layer are iteratively pruned according to a preset pruning rule, including: taking each network layer in the first dialogue model as the current network layer in turn according to a preset pruning order, and looping through the following operations until all network layers in the first dialogue model are pruned and meet preset pruning constraints: determining an initial pruning rate corresponding to the current network layer based on the preset pruning rule; pruning the current network layer according to the initial pruning rate to obtain a pruned dialogue model; obtaining a model accuracy loss based on a model accuracy of the pruned dialogue model and a model accuracy corresponding to the first dialogue model; adjusting the initial pruning rate according to the model accuracy loss to obtain a new pruning rate; and continuing to prune the pruned dialogue model based on the new pruning rate in the same manner as the above operation.
[0044] Optionally, redundant channels of each network layer are iteratively pruned according to preset pruning rules. Specifically, for a certain network layer in the first dialogue model, the channels of the network layer are pruned according to the set initial pruning rate, and the model accuracy after each pruning is tested to obtain the model accuracy loss. The pruning rate is adjusted according to the adjustment direction of the initial pruning rate based on the model accuracy loss, and the model is pruned according to the new pruning rate obtained after the adjustment. In the above manner, the pruning rate is adaptively adjusted according to the model accuracy loss, thereby improving the model compression accuracy while ensuring the model pruning efficiency.
[0045] In an optional embodiment, the initial pruning rate is adjusted according to the model accuracy loss to obtain a new pruning rate, including: when the model accuracy loss is less than or equal to the preset accuracy loss tolerance value, increasing the initial pruning rate according to a preset ratio to obtain a new pruning rate; when the model accuracy loss is greater than the preset accuracy loss tolerance value, using the initial pruning rate as the new pruning rate.
[0046] Optionally, if the model accuracy loss is within the set accuracy loss tolerance, that is: acc(M)-acc(M')≤η
[0047] Where acc(M) represents the accuracy of the unpruned network model, i.e., the model accuracy of the first dialogue model; acc(M') represents the accuracy of the network model after adaptive channel pruning, i.e., the model accuracy of the dialogue model after pruning; η represents the accuracy loss tolerance, which is initialized to 0.005. The pruning rate is gradually increased according to the set proportional factor to achieve adaptive adjustment. Specifically, the pruning rate update process, i.e., the adaptive adjustment process, is as follows: i =p i +αp i
[0048] Among them, p i It represents the channel pruning rate of the i-th layer network, with an initialization value of 0.01, and α is the set scaling factor, with an initialization value of 0.1.
[0049] The above expression makes the channel pruning rate of the network layer gradually increase by a small amount until it reaches the optimal value, that is, the accuracy loss of the network model is minimized after pruning.
[0050] Optionally, if the model accuracy loss is greater than the set accuracy loss tolerance value, that is: acc(M)-acc(M')>η
[0051] It means that the current pruning rate is the optimal pruning rate of the network layer, and the pruning process of the network layer is implemented. The specific expression is as follows: γ i =γ i -p i γ i
[0052] Among them, γ i represents the channel importance parameter of the i-th layer network. Then, the adaptive channel pruning method is used to iteratively prune other unpruned network layers until all network layers in the first dialogue model have completed the pruning process.
[0053] In an optional embodiment, the method further includes: obtaining a text to be input; and obtaining a target reply text using a target dialogue model based on the text to be input.
[0054] Optionally, after obtaining the target conversational model, it can be deployed in actual application scenarios. The user first asks a question as input text, such as "Why hasn't my top-up credit arrived yet?", "When will my voucher be returned?", or "Can I change the delivery address for my product?" This question is then fed into the target conversational model. The model can more quickly answer these three user questions, providing prompt responses such as "Hello, top-up credit takes some time. Please be patient.", "Hello, your voucher will be returned within three business days after verification. Please be patient.", or "Hello, your product has been shipped. Changing the delivery address is currently unavailable. Thank you for your understanding." This reduces user waiting time, improves problem-solving efficiency, and further reduces labor costs. Compared to the original uncompressed model, it significantly accelerates inference.
[0055] Through the above steps S102 to S106, the purpose of adaptively determining the optimal channel pruning rate in each network layer according to the preset pruning rules and the model accuracy corresponding to the model after each round of pruning can be achieved, and the trained dialogue model can be pruned and optimized, thereby achieving the technical effect of adaptive optimization and streamlining the model structure, thereby improving the dialogue prediction efficiency and prediction accuracy, and thus solving the technical problems of complex dialogue model structure, the presence of a large number of redundant channels, and the determination of the pruning rate by human intervention during model pruning, which leads to inaccurate dialogue model determination and thus low dialogue prediction efficiency and accuracy.
[0056] Based on the above embodiment and optional embodiment, the present application proposes an optional implementation method. FIG2 is a flowchart of an optional conversation processing method according to the embodiment of the present application. FIG3 is a flowchart of another optional conversation processing method according to the embodiment of the present application. FIG4 is a flowchart of another optional conversation processing method according to the embodiment of the present application. As shown in FIG2 to FIG4, the method includes:
[0057] Step S1: Parameter redundancy in the TextCNN text classification model often exists in the backbone network. First, the TextCNN text classification model is extracted as the initial dialogue model. The number of network layers of this model is L. The i-th layer network can be expressed as L i , where the channel can be represented as C, and the number of channels is represented as n i , then the whole can be used Subsequent model compression and acceleration are carried out on the initial dialogue model.
[0058] In step S2, the initial dialogue model is trained using a sparse training method. A channel importance evaluation method is constructed to obtain the channel importance value γ for each layer in the first dialogue model (i.e., the trained dialogue model). The scaling factor γ of the BN layer is used as a parameter to measure channel importance. An L1 regularization term for γ is added to the original objective function. After adding the L1 regularization term, training is performed again until the loss converges. This process is called sparse training. The specific steps are shown in Figure 3:
[0059] In step S21, based on the conventional loss function of the first dialogue model, an L1 regularization loss is applied to the target hyperparameter γ in the batch normalization (BN) layer. The target hyperparameter γ represents the multiplication factor in the BN layer. The final loss function expression is as follows: y'=f D (x,W)
[0060] Among them, Loss represents the model loss, l represents the cross-entropy loss function commonly used in classification tasks; x represents the input data, that is, any set of historical input texts among multiple sets of historical input texts; y represents the true label, that is, the historical reply text corresponding to any set of historical input texts; y' represents the prediction result on the historical dialogue dataset D, that is, the prediction result corresponding to any set of historical input texts; W represents the trainable parameter, and λ is used to balance the sparsity of the network during training.
[0061] In step S22, the target hyperparameter γ of the batch normalization layer (BN) gradually approaches 0 during sparse training, so as to solve the problem that the weight of the BN layer of the conventionally trained model will not be too close to 0.
[0062] In step S23, the accuracy of the initial dialogue model and the sparsity of the BN layer gradually reach a balance.
[0063] In step S24, the first dialogue model and the channel importance value γ of each layer of the first dialogue model are obtained. The channel whose corresponding output in the BN layer is closest to 0 will be regarded as a redundant channel and become the target of the next adaptive channel pruning.
[0064] Step S3: construct an adaptive channel pruning method to automatically determine the optimal channel pruning rate for each layer of the initial dialogue model. The specific steps are shown in Figure 4:
[0065] In step S31, the channel importance values obtained in step S2 are sorted in ascending order. The values reflect the importance of the network layer channels and provide a decision basis for subsequent adaptive channel pruning.
[0066] Step S32: Iteratively prune redundant channels at each layer. Specifically, for a certain network layer in the first dialogue model, prune the channels of the network layer according to the set initial pruning rate, and test the model accuracy after each pruning to obtain the model accuracy loss.
[0067] Step S33: If the model accuracy loss is within the set accuracy loss tolerance, that is: acc(M)-acc(M')≤η
[0068] Where acc(M) represents the accuracy of the unpruned network model, i.e., the model accuracy of the first dialogue model; acc(M') represents the accuracy of the network model after adaptive channel pruning, i.e., the model accuracy of the dialogue model after pruning; η represents the accuracy loss tolerance, which is initialized to 0.005. The pruning rate is gradually increased according to the set proportional factor to achieve adaptive adjustment. Specifically, the pruning rate update process, i.e., the adaptive adjustment process, is as follows: i =p i +αp i
[0069] Among them, p i It represents the channel pruning rate of the i-th layer network in the locked network layer, with an initialization value of 0.01, and α is the set scaling factor, with an initialization value of 0.1.
[0070] The above expression makes the channel pruning rate of the network layer gradually increase by a small amount until it reaches the optimal value, that is, the accuracy loss of the network model is minimized after pruning.
[0071] Step S34: If the model accuracy loss is greater than the set accuracy loss tolerance value, that is: acc(M)-acc(M')>η
[0072] It means that the current pruning rate is the optimal pruning rate of the network layer, and the pruning process of the network layer is implemented. The specific expression is as follows: γ i =γ i -p i γ i
[0073] Among them, γ i represents the channel importance parameter of the i-th layer network. Then, the adaptive channel pruning method is used to iteratively prune other unpruned network layers until all network layers in the first dialogue model have completed the pruning process and meet the preset pruning conditions, where the preset pruning conditions are:
[0074] Among them, S iIndicates the parameter amount of any network layer in multiple network layers, S represents the total parameter amount in the first dialogue model, L represents the number of multiple network layers, p g is the preset global pruning rate.
[0075] The advantage of the pruning method proposed in this application is that it does not require human intervention, and effectively solves the problems of over-pruning and under-pruning that exist in the existing methods of manually setting thresholds. The above adaptive channel pruning process can ensure that redundant parameters are maximized while maintaining high model accuracy.
[0076] In step S4, the first dialogue model pruned in step S3 is fine-tuned to further restore the accuracy of the obtained lightweight pruned first dialogue model to the accuracy level of the first dialogue model before pruning, thereby obtaining a target dialogue model.
[0077] In step S5, the final target dialogue model is a lightweight TextCNN text classification model, which significantly compresses the number of parameters and floating-point calculations of the TextCNN text classification model while maintaining accuracy, thereby achieving the purpose of inference acceleration.
[0078] In step S6, during the user-AI customer service interaction scenario, the user first asks a question, such as, "Why hasn't my recharged phone bill arrived yet?", "When will my voucher be returned?", or "Can I change the delivery address for my product?" This question is fed into the target dialogue model. The target dialogue model can more quickly answer these three user questions, providing prompt responses such as, "Hello, recharging your phone bill takes some time. Please be patient.", "Hello, your voucher will be returned within three business days after verification. Please be patient.", or "Hello, your product has been shipped. The delivery address cannot be changed at this time. Thank you for your understanding." This reduces user waiting time, improves problem-solving efficiency to a certain extent, further saves labor costs, and significantly accelerates inference compared to the original, uncompressed model.
[0079] In the embodiment of the present application, an L1 regularization loss is applied to the target hyperparameter γ in the BN layer, and a sparse training model is used to accurately assess the importance of channels, providing a decision basis for the adaptive channel pruning of the subsequent TextCNN text classification model. In addition, the embodiment of the present application introduces an accuracy loss tolerance mechanism to achieve adaptive adjustment of the pruning rate, thereby deriving the optimal channel pruning rate for each layer of the TextCNN text classification model, ensuring that redundant parameters are eliminated as much as possible while maintaining high model accuracy.
[0080] The embodiments of the present application can achieve at least one of the following effects: 1) The channel importance evaluation method proposed in the embodiments of the present application can accurately evaluate the channel importance of each layer of the TextCNN text classification model, and has a higher degree of generalization. 2) The advantage of the adaptive channel pruning method proposed in the embodiments of the invention is that it automatically determines the optimal channel pruning rate of each layer of the TextCNN text classification model without human intervention, eliminating the tedious process of manually setting thresholds in the existing methods, eliminating redundant parameters, and to a certain extent avoiding the problems of over-pruning and under-pruning in the existing methods. While maintaining accuracy, it greatly compresses the number of parameters and floating-point calculations of the TextCNN text classification model, realizes inference acceleration, improves the efficiency of solving user problems, further saves labor costs, and effectively promotes the practical application of the TextCNN text classification model.
[0081] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0082] This embodiment also provides a conversation processing device for implementing the above-mentioned embodiments. Details already described will not be repeated. As used below, the terms "module" and "device" may refer to a combination of software and / or hardware that implements a predetermined function. The devices described in the following embodiments may be implemented in software, hardware, or a combination of software and hardware. Implementation is also possible and contemplated.
[0083] According to an embodiment of the present application, an embodiment of a device for implementing the above-mentioned dialog processing method is also provided. FIG5 is a schematic structural diagram of a dialog processing device according to an embodiment of the present application. As shown in FIG5 , the above-mentioned dialog processing device includes: a machine learning module 500, a determination module 502, and a pruning module 504, wherein:
[0084] A machine learning module 500 is configured to perform machine learning on an initial dialogue model based on a historical dialogue dataset to obtain a first dialogue model, wherein the first dialogue model includes multiple network layers, each network layer includes multiple channels, and each channel has an importance value corresponding to each channel, wherein the importance value indicates the importance of the corresponding channel, and the historical dialogue dataset includes multiple sets of historical input texts and historical response texts corresponding to each set of historical input texts.
[0085] a determination module 502, connected to the machine learning module 500, configured to determine redundant channels in each network layer based on importance values corresponding to the plurality of channels included in each network layer;
[0086] The pruning module 504 is connected to the determination module 502 and is used to iteratively prune redundant channels in each network layer according to preset pruning rules to obtain a target dialogue model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the ratio of the number of pruned parameters to the total number of parameters in the corresponding channel.
[0087] In an embodiment of the present application, a machine learning module 500 is set to perform machine learning on an initial dialogue model based on a historical dialogue data set to obtain a first dialogue model, wherein the first dialogue model includes multiple network layers, each network layer includes multiple channels, and the multiple channels respectively correspond to importance values, wherein the importance values are used to indicate the importance of the corresponding channels, and the historical dialogue data set includes multiple groups of historical input texts and historical reply texts corresponding to the multiple groups of historical input texts; a determination module 502 is connected to the machine learning module 500 and is used to determine redundant channels in each network layer based on the importance values corresponding to the multiple channels included in each network layer; a pruning module 504 is connected to the determination module 502 and is used to iterate the redundant channels in each network layer according to preset pruning rules. The target dialogue model is obtained by pruning the model on behalf of the network, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the proportion of the number of pruned parameters to the total number of parameters in the corresponding channel. The purpose of adaptively determining the optimal channel pruning rate in each network layer according to the preset pruning rules and the model accuracy corresponding to the model after each round of pruning is achieved, and pruning and optimizing the trained dialogue model is performed, thereby achieving the technical effect of adaptive optimization and streamlining the model structure, thereby improving the dialogue prediction efficiency and prediction accuracy, and thus solving the technical problems that the dialogue model structure is complex, there are a large number of redundant channels, and the pruning rate is determined by human intervention during model pruning, resulting in inaccurate dialogue model determination, and thus leading to low dialogue prediction efficiency and low accuracy.
[0088] In an optional embodiment, the machine learning module includes: a first determination submodule for determining an initial loss function; a first acquisition submodule for adding an L1 regularization term regarding a target hyperparameter to the initial loss function to obtain a target loss function, wherein the target hyperparameter is located in a batch normalization layer in each network layer; a first machine learning submodule for performing machine learning on the initial conversation model based on a historical conversation dataset and a target loss function to obtain a first conversation model; and a second determination submodule for determining importance values corresponding to multiple channels included in each network layer according to the target hyperparameter in the batch normalization layer in each network layer.
[0089] In an optional embodiment, the first machine learning submodule includes: a second machine learning submodule, which is used to perform machine learning on the initial dialogue model based on the historical dialogue data set and the target loss function; and a first output submodule, which is used to output the first dialogue model when the importance values corresponding to a predetermined number of channels among the multiple channels included in each network layer are within a preset range.
[0090] In an optional embodiment, the pruning module includes: a first pruning submodule, which is used to iteratively prune redundant channels in each network layer according to preset pruning rules; and a second output submodule, which is used to output a target dialogue model when all network layers in the first dialogue model have been pruned and preset pruning constraints are met.
[0091] In an optional embodiment, the second output submodule includes: according to a preset pruning order, taking each network layer in the first dialogue model as the current network layer in turn, and looping the following operations until all network layers in the first dialogue model are pruned and meet preset pruning constraints: determining an initial pruning rate corresponding to the current network layer based on a preset pruning rule; pruning the current network layer according to the initial pruning rate to obtain a pruned dialogue model; obtaining a model accuracy loss based on the model accuracy of the pruned dialogue model and the model accuracy corresponding to the first dialogue model; adjusting the initial pruning rate according to the model accuracy loss to obtain a new pruning rate; and continuing to prune the pruned dialogue model based on the new pruning rate in the same manner as the above operation.
[0092] In an optional embodiment, the initial pruning rate is adjusted according to the model accuracy loss to obtain a new pruning rate, including: a second pruning sub-module, which is used to increase the initial pruning rate according to a preset ratio to obtain a new pruning rate when the model accuracy loss is less than or equal to a preset accuracy loss tolerance value; and a pruning rate update sub-module, which is used to use the initial pruning rate as the new pruning rate when the model accuracy loss is greater than the preset accuracy loss tolerance value.
[0093] In an optional embodiment, the preset pruning constraint condition is:
[0094] Among them, S i represents the parameter amount of any network layer in multiple network layers, S represents the total parameter amount in the first dialogue model, L represents the number of multiple network layers, i represents any network layer in multiple network layers, p g is the preset global pruning rate.
[0095] In an optional embodiment, the method further includes: a second acquisition submodule, used to acquire the text to be input; and a prediction submodule, used to obtain the target reply text based on the text to be input and using the target dialogue model.
[0096] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0097] It should be noted that the machine learning module 500, determination module 502, and pruning module 504 described above correspond to steps S102 to S106 in the embodiment. The examples and application scenarios implemented by these modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment. It should be noted that these modules, as part of the device, can be run on a computer terminal.
[0098] It should be noted that the optional implementation methods of this embodiment can be found in the relevant descriptions in the embodiments and will not be repeated here.
[0099] The above-mentioned dialogue processing device may also include a processor and a memory. The above-mentioned machine learning module 500, determination module 502, pruning module 504, etc. are all stored in the memory as program modules, and the processor executes the above-mentioned program modules stored in the memory to realize corresponding functions.
[0100] The processor includes a core, which retrieves corresponding program modules from memory. There can be one or more cores. Memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.
[0101] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is further provided. Optionally, in this embodiment, the non-volatile storage medium includes a stored program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute any of the above-mentioned conversation processing methods.
[0102] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group, and the non-volatile storage medium includes a stored program.
[0103] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following functions: based on the historical conversation data set, machine learning is performed on the initial conversation model to obtain a first conversation model, and importance values corresponding to the multiple channels included in each network layer in the multiple network layers included in the first conversation model, wherein the importance value is used to indicate the degree of importance of the corresponding channel, and the historical conversation data set includes multiple groups of historical input texts, and historical reply texts corresponding to the multiple groups of historical input texts; based on the importance values corresponding to the multiple channels included in each network layer, redundant channels in each network layer are determined; redundant channels in each network layer are iteratively pruned according to preset pruning rules to obtain a target conversation model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the proportion of the number of pruned parameters to the total number of parameters in the corresponding channel.
[0104] According to an embodiment of the present application, a processor embodiment is also provided. Optionally, in this embodiment, the processor is used to run a program, wherein the program executes any one of the above-mentioned dialogue processing methods when it is run.
[0105] According to an embodiment of the present application, an embodiment of a computer program product is also provided, which, when executed on a data processing device, is suitable for executing a program that initializes any one of the steps of the above-mentioned dialogue processing method.
[0106] Optionally, the above-mentioned computer program product, when executed on a data processing device, is suitable for executing a program initialized with the following method steps: based on the historical conversation data set, performing machine learning on the initial conversation model to obtain a first conversation model, and importance values corresponding to the multiple channels included in each network layer in the multiple network layers included in the first conversation model, wherein the importance value is used to indicate the degree of importance of the corresponding channel, and the historical conversation data set includes multiple groups of historical input texts and historical reply texts corresponding to the multiple groups of historical input texts; based on the importance values corresponding to the multiple channels included in each network layer, determining the redundant channels in each network layer; iteratively pruning the redundant channels in each network layer according to preset pruning rules to obtain the target conversation model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the proportion of the number of pruned parameters to the total number of parameters in the corresponding channel.
[0107] An embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the following steps are implemented: based on a historical conversation data set, machine learning is performed on an initial conversation model to obtain a first conversation model, and importance values corresponding to multiple channels included in each network layer in multiple network layers included in the first conversation model, wherein the importance values are used to indicate the degree of importance of the corresponding channel, and the historical conversation data set includes multiple groups of historical input texts and historical reply texts corresponding to the multiple groups of historical input texts; based on the importance values corresponding to the multiple channels included in each network layer, redundant channels are determined in each network layer; and redundant channels in each network layer are iteratively pruned according to preset pruning rules to obtain a target conversation model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the proportion of the number of pruned parameters to the total number of parameters in the corresponding channel.
[0108] The above sequence of the embodiments of the present application is for description only and does not represent the superiority or inferiority of the embodiments.
[0109] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0110] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the above modules can be a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, modules or indirect coupling or communication connection of modules, which can be electrical or other forms.
[0111] The modules described above as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0112] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.
[0113] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a non-volatile storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned non-volatile storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.
[0114] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0115] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A conversation processing method, comprising: Based on the historical conversation data set, the initial conversation model is machine-learned to obtain a first conversation model, wherein the first conversation model includes a plurality of network layers, each network layer includes a plurality of channels, and the plurality of channels respectively correspond to importance values, wherein the importance values are used to indicate the importance of the corresponding channel, and the historical conversation data set includes a plurality of groups of historical input texts, and historical reply texts respectively corresponding to the plurality of groups of historical input texts; Determine redundant channels in each network layer based on the importance values respectively corresponding to the multiple channels included in each network layer; The redundant channels in each network layer are iteratively pruned according to preset pruning rules to obtain a target dialogue model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the proportion of the number of pruned parameters to the total number of parameters in the corresponding channel.
2. The method according to claim 1, wherein: The method of performing machine learning on the initial dialogue model based on the historical dialogue data set to obtain the first dialogue model includes: Determine the initial loss function; Adding an L1 regularization term about a target hyperparameter to the initial loss function to obtain a target loss function, wherein the target hyperparameter is located in a batch normalization layer in each network layer; Based on the historical dialogue data set and the target loss function, machine learning is performed on the initial dialogue model to obtain the first dialogue model.
3. The method according to claim 2, wherein: The performing machine learning on the initial dialogue model based on the historical dialogue data set and the target loss function to obtain the first dialogue model includes: Based on the historical dialogue data set and the target loss function, performing machine learning on the initial dialogue model; In the case where the importance values corresponding to a predetermined number of the multiple channels included in each network layer are within a preset interval, the first dialogue model is output.
4. The method according to claim 1, wherein: The iterative pruning of redundant channels in each network layer according to a preset pruning rule to obtain a target dialogue model includes: Iteratively pruning the redundant channels in each network layer according to the preset pruning rule; When all network layers in the first dialogue model are pruned and preset pruning constraints are met, the target dialogue model is output.
5. The method according to claim 4, wherein: The iterative pruning of the redundant channels in each network layer according to the preset pruning rule includes: According to the preset pruning order, each network layer in the first dialogue model is used as the current network layer in turn, and the following operations are performed cyclically until all network layers in the first dialogue model are pruned and the preset pruning constraints are met: Based on the preset pruning rule, determine the initial pruning rate corresponding to the current network layer; prune the current network layer according to the initial pruning rate to obtain a pruned dialogue model; obtain the model accuracy loss based on the model accuracy of the pruned dialogue model and the model accuracy corresponding to the first dialogue model; adjust the initial pruning rate according to the model accuracy loss to obtain a new pruning rate; based on the new pruning rate, continue to prune the pruned dialogue model in the same processing manner as the above operation.
6. The method according to claim 5, wherein: The adjusting the initial pruning rate according to the model accuracy loss to obtain a new pruning rate includes: When the model accuracy loss is less than or equal to a preset accuracy loss tolerance value, increasing the initial pruning rate according to a preset ratio to obtain the new pruning rate; When the model accuracy loss is greater than the preset accuracy loss tolerance value, the initial pruning rate is used as the new pruning rate.
7. The method according to claim 4, wherein: The preset pruning constraints are: Among them, S i represents the parameter amount of any one of the multiple network layers, S represents the total parameter amount in the first dialogue model, L represents the number of the multiple network layers, i represents any one of the multiple network layers, p g is the preset global pruning rate.
8. The method according to any one of claims 1 to 7, wherein: The method further comprises: Get the text to be input; Based on the text to be input, the target dialogue model is adopted to obtain a target reply text.
9. A dialogue processing device, comprising: a machine learning module, configured to perform machine learning on an initial dialogue model based on a historical dialogue data set to obtain a first dialogue model, wherein the first dialogue model includes a plurality of network layers, each network layer includes a plurality of channels, and the plurality of channels respectively correspond to importance values, wherein the importance values are used to indicate the importance of the corresponding channel, and the historical dialogue data set includes a plurality of groups of historical input texts, and historical reply texts respectively corresponding to the plurality of groups of historical input texts; A determination module, configured to determine a redundant channel in each network layer based on the importance values respectively corresponding to the multiple channels included in each network layer; A pruning module is used to iteratively prune the redundant channels in each network layer according to preset pruning rules to obtain a target dialogue model, wherein the preset pruning rules are used to indicate the corresponding relationship between the model accuracy and the pruning rate corresponding to the model after each round of pruning, and the pruning rate is used to indicate the proportion of the number of pruned parameters to the total number of parameters in the corresponding channel.
10. An electronic device comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein: When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Model processing method, federal learning method and related equipment
CN113469340A
Neural network model compression method based on structure search and channel pruning
CN114330644A
Automatic question and answer library updating method and device for open domain science popularization
CN116361306A
Model compression method and device based on channel pruning, equipment and medium
CN116502695A
Conversation processing method and device and electronic equipment
CN117453878A