Power business customer service question and answer model training method and computer system
By performing random disturbances and disturbance removal on the customer service question and answer network of the power business, combined with quality evaluation indicators and style constraint information, the problems of unprofessional response and insufficient training in traditional systems are solved, and the accuracy of response and user satisfaction are improved.
Patent Information
- Application Number
- CN202510354129.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-19
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional power service customer service Q&A systems are difficult to generate response texts that meet specific requirements based on user needs, lack professionalism and accuracy, and lack effective optimization mechanisms in the training process, resulting in a decrease in user satisfaction.
By obtaining the training binary data of the power service customer service Q&A network, random perturbation and perturbation removal are performed, and network parameters are adjusted in combination with quality evaluation indicators and response style constraint information to generate predicted response text that meets the style and content requirements.
It improves the accuracy and professionalism of the customer service Q&A system of the power business, enhances user satisfaction, and optimizes the pertinence and effectiveness of the training process.
Smart Images

Figure CN120296122A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a method for training an electric power service customer service question-answering model and a computer system. Background Art
[0002] In the current electric power service field, with the continuous growth of electric power service demand and the improvement of users' expectations for service quality, the electric power service customer service question-answering system faces many challenges.
[0003] The electric power service covers a wide range of knowledge fields, including complex business areas such as power supply, power equipment maintenance, electricity bill pricing and payment, and power safety. The types of questions raised by users are numerous and highly diverse, which requires the customer service question-answering system to accurately understand the users' questions and give appropriate answers. However, when dealing with these complex and diverse questions, the existing systems often have difficulty generating response texts that meet specific requirements according to different user needs. For example, in terms of style, it may not be able to match the electric power service scenario. For some questions that require a professional and rigorous style, the response may be too colloquial or inaccurate, lacking the due professionalism. Or in terms of content quality, it cannot provide comprehensive, accurate, and effective answers to the users' questions well, and may omit important information or provide wrong answers, thus affecting users' satisfaction with the electric power service. On the other hand, the existing customer service question-answering systems lack an effective optimization mechanism during the training process. Due to the complexity of the electric power service, the system needs to be continuously adjusted and optimized to adapt to various situations, but the traditional training methods are difficult to comprehensively and targeted tune the network. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for training an electric power service customer service question-answering model and a computer system. The present invention is implemented as follows:
[0005] In a first aspect, the present invention provides a method for training a power service customer service Q&A model, including: obtaining training binary group data of a power service customer service Q&A network, where the training binary group data includes training service texts and response style constraint information of the training service texts; randomly perturbing the training service texts to obtain perturbed representation vectors of the training service texts, and obtaining the number of perturbation removals relied on during the iterative calibration process of the power service customer service Q&A network; wherein the iterative calibration process of the power service customer service Q&A network includes one or more calibration links, and different calibration links have different numbers of perturbation removals; removing perturbations from the perturbed representation vectors according to the number of perturbation removals to obtain perturbation-removed representation vectors, and generating predicted response texts of the training service texts based on the perturbation-removed representation vectors; obtaining quality evaluation indicators of the predicted response texts, generating a backpropagated optimization error based on the quality evaluation indicators, the predicted response texts, and the response style constraint information, and adjusting network parameters of the power service customer service Q&A network based on the backpropagated optimization error to obtain a trained power service customer service Q&A network.
[0006] In a second aspect, the present invention provides a computer system, including: one or more processors; a memory; one or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method described above is implemented.
[0007] When the present invention adjusts the network parameters of the power service customer service Q&A network, it obtains the training binary group data for network parameter adjustment in the real-time calibration link of the power service customer service Q&A network, so as to predict through the power service customer service Q&A network according to the training service text in the training binary group data and the corresponding response style constraint information of the training service text, and obtain the predicted response text corresponding to the training service text under the restriction of the response style constraint information. After obtaining the predicted response text, evaluate the quality of the predicted response text according to the corresponding evaluation score in the process of network parameter adjustment in the real-time calibration link, obtain the construction quality of the predicted response text constructed by the power service customer service Q&A network, and determine the training status of the power service customer service Q&A network according to the construction quality. Therefore, after obtaining the construction quality of the predicted response text, the training status of the power service customer service Q&A network reflected by the construction quality can be transmitted back to the power service customer service Q&A network, so that the power service customer service Q&A network can be further adjusted according to the feedback of the construction quality, and the text reconstruction effect of the power service customer service Q&A network can be improved. When transmitting the construction quality back to the power service customer service Q&A network and adjusting the parameters of the power service customer service Q&A network, according to the construction quality of the predicted response text, the predicted response text generated by the power service customer service Q&A network, and the response style constraint information guiding the generation of the predicted response text, determine the corresponding feedback optimization error for the power service customer service Q&A network. In this way, the power service customer service Q&A network can be iterated using the feedback optimization error, and the parameter adjustment of the power service customer service Q&A network can be completed according to the text reconstruction quality of the power service customer service Q&A network. By obtaining different training binary group data and performing feedback iteration on the power service customer service Q&A network according to different calibration links, since the focus of feedback iteration on the power service customer service Q&A network in different calibration links is different, the effect of the power service customer service Q&A network completing text reconstruction according to the style constraint information can be strengthened in different dimensions, the training quality of the power service customer service Q&A network can be improved, and thus the quality of text reconstruction can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a flowchart of a method for training a power service customer service Q&A model provided by an embodiment of the present invention.
[0009] Figure 2 is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0010] In the embodiments of the present invention, the execution subject of the method for training a power service customer service Q&A model is a computer system, including but not limited to servers, personal computers, laptops, tablets, etc. The computer system includes user equipment and network equipment. As Figure 1 shown, the method includes:
[0011] Step S100: Obtain the training binary tuple data of the power service customer service Q&A network. The training binary tuple data includes training service texts and response style constraint information for the training service texts.
[0012] In the embodiment of the present invention, the training binary tuple data consists of two parts, namely the training service text and the response style constraint information for the training service text.
[0013] The training service text is the basic data used by the computer system to train the power service customer service Q&A network. It is the response text for the content of the user's conversation to be replied. For example, in the power service, the user can ask a question about electricity bill inquiry, such as "How to inquire about this month's electricity bill?" Then a reasonable training service text for this question may be "You can log in to your account through the official mobile application of the power company, and view this month's electricity bill in the bill inquiry section. Or call the customer service hotline and provide your user number for inquiry." This text clearly answers the user's question about electricity bill inquiry and is a response text for a specific power service question.
[0014] The response style constraint information is the restriction and requirement for the training service text in terms of style. It can cover multiple aspects, such as language style (formal, concise, easy to understand, very professional, etc.), emotional tendency (positive, neutral, negative, etc.). Continuing with the above example of electricity bill inquiry, if the power company hopes that the customer service's answer gives the user a feeling of being kind, friendly and concise, then this style requirement is part of the response style constraint information. For example, the response style constraint information can be set as "The language style is easy to understand and friendly, avoid using overly complex professional terms, and the sentences are short and concise." Of course, the response style constraint information can also be represented by pre-set style tags.
[0015] A feasible way to obtain the training binary tuple data is to extract it from the existing power service customer service conversation record database. Assume that there is a historical conversation database of power service customer service, which contains a large number of interaction records between the customer service and the user. The computer system can screen out the qualified Q&A pairs from these conversation records as the training binary tuple data through data mining and information extraction technologies. Specifically, for each Q&A pair, the question part can be regarded as the user's input, and the answer part is used as the training service text. Then, according to the style requirements pre-set by the power company (such as analyzing the style characteristics reflected in the previous high-quality customer service answers), determine the response style constraint information corresponding to each training service text.
[0016] Another way to obtain training binary tuple data is through manual annotation. The computer system can first collect some text content related to power services, which may come from business manuals, promotional materials, etc. of power companies. Then, professional personnel (such as power service experts and customer service training personnel) are organized to annotate these texts. According to the characteristics of power services and the requirements of customer service responses, the annotators convert these texts into training service texts and add response style constraint information. For example, for a certain paragraph in a power safety knowledge promotional brochure: "A safe distance should be maintained around power equipment to avoid the risk of electric shock. When performing maintenance on power equipment, be sure to cut off the power supply first and use professional tools.", the annotator can convert it into a training service text: "Please keep a safe distance near power equipment to prevent electric shock. Cut off the power supply first and use professional tools when maintaining the equipment.", and at the same time annotate the response style constraint information as "Easy to understand, with a certain tone of warm reminder.".
[0017] Step S200: Randomly perturb the training service text to obtain a perturbed representation vector of the training service text, and obtain the number of perturbation removals relied on by the power service customer service Q&A network during the iterative tuning process; wherein, the iterative tuning process of the power service customer service Q&A network includes one or more tuning links, and the number of perturbation removals passed by different tuning links is different.
[0018] In the embodiment of the present invention, random perturbation is an operation to change the training service text. For example, in the power service customer service Q&A scenario, if the training service text is "You can query your electricity bill through the official website of the power company. The specific operation is to log in to your account and enter the bill page to view.", during random perturbation, some words can be randomly changed at certain positions in the text, or the order of words can be adjusted, etc., but it is necessary to ensure that the overall semantics is retained to a certain extent.
[0019] To achieve this random perturbation and obtain the perturbed representation vector, one way is based on the vector representation technology of text. The computer system first converts the training service text into a vector form, and this vector can mathematically represent the semantic information of the text. Assuming a word vector model (such as Word2Vec) is used, each word has a corresponding vector representation in a predefined vector space. For the entire training service text, the computer system can obtain an initial text vector representation by combining the word vectors therein according to certain rules (such as simple summation or weighted summation). Then, the computer system introduces random factors for perturbation. This can be achieved by adding random noise to certain dimensions of this vector. For example, assuming this text vector is V = [v1, v2, …, v n , the computer system generates a random noise vector N = [n1, n2, …, n n , where ni is a random number within a certain range (e.g., between -∈ and ∈, where ∈ is a pre-set small positive number). Then add these two vectors to get the perturbed vector V′ = V + N = [v1 + n1, v2 + n2, …, v n + n n . This V′ is the perturbed representation vector for training business texts.
[0020] Next is to obtain the number of perturbation removals relied on during the iterative calibration process of the power business customer service Q&A network. During the iterative calibration process of the power business customer service Q&A network, different calibration steps require different numbers of perturbation removals. This number of perturbation removals is determined by the computer system according to the calibration status and requirements of the network. For example, in the power business customer service Q&A network, there may be different calibration steps focusing on different aspects, such as the calibration of the accuracy of the answer content, the calibration of the answer style, etc. If it is a calibration step focusing on the accuracy of the answer content, it may require a smaller number of perturbation removals because the content accuracy may rely more on the original text information, and excessive perturbation removal may introduce unnecessary errors; while for the calibration step of the answer style, it may require a larger number of perturbation removals to better adjust the text style.
[0021] One implementation way for the computer system to determine the number of perturbation removals is based on pre-set rules and the current state of the network. Assume that the iterative calibration process of the power business customer service Q&A network is divided into several different stages. The computer system can determine the specific number according to which stage it is currently in and the pre-set range of the number of perturbation removals for each stage. For example, if the entire iterative calibration process is divided into three stages, namely the initial stage, the intermediate stage, and the final stage, and it is stipulated that the range of the number of perturbation removals in the initial stage is a1 to b1, in the intermediate stage is a2 to b2, and in the final stage is a3 to b3. When the computer system determines that it is currently in the intermediate stage, it can determine a specific number of perturbation removals within the range of a2 to b2 through a certain random selection algorithm (such as uniform random selection). If the iterative calibration process of the power business customer service Q&A network is represented as a function F(t), where t represents the iteration time step or stage, and the number of perturbation removals is represented as d(t), then d(t) = g(F(t), p), where g is a determined function, not specifically limited, and p is a set of pre-set parameters (such as the range of the number of perturbation removals for each stage, etc.).
[0022] This operation of randomly perturbing the training business text and determining the number of perturbation removals in step S200 increases the diversity of data by introducing random perturbations. At the same time, it determines the appropriate number of perturbation removals according to different calibration links, laying a foundation for accurately adjusting the parameters of the power business customer service Q&A network in the subsequent steps, and helping to improve the accuracy and adaptability of the network for power business customer service Q&A.
[0023] Step S300: Remove the perturbations from the perturbation representation vector according to the number of perturbation removals to obtain a perturbation-removed representation vector, and generate a predicted response text for the training business text based on the perturbation-removed representation vector.
[0024] The perturbation representation vector is obtained by randomly perturbing the training business text in step S200, and it contains the information of the original training business text after perturbation. The computer system removes the perturbations from the perturbation representation vector according to the number of perturbation removals. The number of perturbation removals is determined in step S200, and it reflects the number of operations required to remove the perturbations for this perturbation representation vector during the iterative calibration process of the current power business customer service Q&A network. The determination of this number is related to the calibration link of the network, and different calibration links may require different numbers of perturbation removals to achieve different optimization purposes.
[0025] The computer system can adopt a step-by-step reduction algorithm to perform perturbation removal. Suppose the perturbation representation vector V′ = [v′1, v′2, …, v′ n , and the number of perturbation removals is k. The computer system can define a reduction function R(V', k), and this function will reduce the perturbation components in V' according to certain rules each time it operates. For example, this rule can be that each time the values of certain dimensions in the vector are moved closer to the corresponding dimension values of the original unperturbed vector (suppose it is V = [v1, v2, …, v n ) by a pre-determined ratio α (0 < α < 1). After k such operations, the perturbation-removed representation vector V” is obtained. Specifically, at the i-th operation (1 ≤ i ≤ k), v″ j,i = (1 - α)v′ j,i-1 + αv j (here v″ j,i represents the value of the j-th dimension of the vector V” after the i-th operation, and v′ j,i-1 represents the value of the j-th dimension of the perturbation representation vector V' after the (i - 1)-th operation).
[0026] After obtaining the perturbation-removed representation vector V'', the computer system generates a predicted response text for the training service text based on it. This process can utilize a pre-trained language model or a neural network-based text generation algorithm. For example, the computer system can use a text generator based on a long short-term memory network (LSTM). Taking the perturbation-removed representation vector V'' as the input, this text generator gradually generates a sequence of words for the predicted response text according to its internal neuron structure and weight parameters. Suppose this text generator has a hidden state h t , at each time step t, according to the input vector V'' and the hidden state h t-1 at the previous time step, through an activation function σ (such as the tanh function) and some weight matrices W, U, b (these weight matrices are determined during pre-training or initialization), calculate the next hidden state h t = σ(WV″ + Uh t-1 + b). Then, based on the hidden state h t and an output matrix O and another activation function (such as the softmax function) to determine the probability distribution P(w t ) of the word output at this time step = softmax(O h t ), thus selecting the word with the highest probability as the next word in the predicted response text.
[0027] Step S400: Obtain a quality evaluation index for the predicted response text, generate a backpropagation optimization error based on the quality evaluation index, the predicted response text, and the response style constraint information, and adjust the network parameters of the power service customer service Q&A network based on the backpropagation optimization error to obtain a trained power service customer service Q&A network.
[0028] In the embodiment of the present invention, the computer system obtains a quality evaluation index for the predicted response text. The predicted response text is generated based on the perturbation-removed representation vector in step S300, and it is a response result of the power service customer service Q&A network for the training service text. The quality evaluation index is a quantitative or qualitative standard for measuring the quality of this predicted response text. For example, in the power service customer service Q&A scenario, if the training service text is an inquiry about the power failure repair process, and the predicted response text is "You can just call directly", this response is too brief and lacks key information, and the quality may be poor; while if the predicted response text is "You can call the customer service hotline, inform the customer service of your address, fault phenomenon, etc., and the maintenance personnel will be arranged to come to repair within the specified time", this answer is relatively complete and the quality is relatively high.
[0029] To obtain quality assessment metrics, a computer system can adopt various technical means. One approach is to conduct the assessment based on predefined rules and criteria. For example, for power business customer service Q&A, certain rules can be set regarding aspects such as integrity, accuracy, and professionalism. If a predicted response text contains all the necessary information (such as for a power cost query question, including the query channel, required information, etc.), a higher score can be given in terms of integrity; if the information in the answer matches the actual power business knowledge (such as the calculation method of electricity bills, etc.), a higher score is obtained in terms of accuracy; if the answer uses professional terms and conforms to the power industry norms, it performs well in terms of professionalism. If the assessments in aspects such as integrity, accuracy, and professionalism are represented as functions C(p), A(p), P(p) respectively (where p represents the predicted response text), then the quality assessment metric Q(p) can be a certain combination of these functions. For example, Q(p) = w1C(p) + w2A(p) + w3P(p), where w1, w2, and w3 are predefined weights indicating the importance of each aspect in the overall quality assessment.
[0030] After the computer system obtains the quality assessment metric, it generates a feedback optimization error based on this metric, the predicted response text, and the response style constraint information. The response style constraint information is obtained together with the training business text in step S100, and it stipulates the requirements for the style of the response text, such as language style (formal, colloquial, etc.), emotional tendency (positive, neutral, etc.). For example, for a power business customer service answer, if the required style is formal and positive, then the answer "We will handle it as soon as possible. Wish you a happy life" meets this style requirement; while "Well, I'll do it" does not.
[0031] When generating the feedback optimization error, the computer system can adopt different technical means according to different situations. If the predicted response text does not match the response style constraint information in terms of style, the computer system can calculate the degree of style deviation as part of the feedback optimization error. Suppose there is a style distance function S(p, s) (where p is the predicted response text and s is the response style constraint information), and the style deviation value calculated by this function is the component of the feedback optimization error in terms of style.
[0032] At the same time, if there are problems with the content of the predicted response text, such as inaccurate or incomplete information, the computer system also needs to reflect this part of the error in the feedback optimization error. For example, if the answer regarding the repair cost of power equipment in the predicted response text does not match the actual charging standard, the computer system needs to calculate a value based on this deviation as the component of the feedback optimization error in terms of content accuracy.
[0033] Based on this backpropagation optimization error, the computer system adjusts the network parameters of the power service customer service Q&A network. The power service customer service Q&A network is a model constructed based on technologies such as neural networks, and its network parameters determine the output results of the model. For example, the weights and biases in a neural network are typical network parameters. The computer system optimizes the performance of the network by adjusting these parameters. If the backpropagation optimization error represents the gap between the predicted response text and the ideal response, the computer system can use the backpropagation algorithm to adjust the network parameters.
[0034] As an implementation manner, in step S200, random perturbations are performed on the training service text to obtain a perturbed representation vector of the training service text, including:
[0035] Step S210: Perform a text embedding operation on the training service text to obtain a training service text representation vector of the training service text;
[0036] Step S220: Obtain an adjustment fluctuation parameter, and perform random perturbation on the training service text representation vector through the adjustment fluctuation parameter to obtain an initial perturbed representation vector of the training service text; wherein, the adjustment fluctuation parameter is a random mask, the random mask and the initial perturbed representation vector are in the same vector domain, and the random mask and the initial perturbed representation vector have the same vector dimension in the corresponding vector domain;
[0037] Step S230: Perform vector encoding on the initial perturbed representation vector to obtain a perturbed representation vector of the training service text.
[0038] In step S210, the computer system performs a text embedding operation on the training service text to obtain a training service text representation vector of the training service text. Text embedding is a technical means to convert text into a vector representation, aiming to enable the computer to process text in a more suitable way for calculation and analysis. In the embodiments of the present invention, for example, the training service text is "Power users can query electricity consumption details through an online platform". When the computer system performs a text embedding operation, each word in this text will be mapped to a predefined vector space.
[0039] Suppose the computer system uses a pre-trained word vector model, such as the Word2Vec model. In the Word2Vec model, each word is represented as a vector with a fixed dimension. For this training service text, the computer system will convert words such as "power", "user", "can", "through", "online platform", "query", "electricity consumption details" into corresponding vectors respectively. Then, in order to obtain the representation vector of the entire training service text, the computer system can adopt some combination strategies. A feasible strategy is a simple average strategy, that is, sum these word vectors and then divide by the number of words. If we use v1, v2, …, vn representing each word vector (where n is the number of words in the training business text), then the training business text representation vector can be expressed as Of course, in addition to the average strategy, other strategies such as weighted summation can also be adopted. When using weighted summation, different weights can be assigned according to factors such as the importance of words.
[0040] Next is step S220. The computer system obtains an adjustment fluctuation parameter, and randomly perturbs the training business text representation vector through the adjustment fluctuation parameter to obtain an initial perturbed representation vector of the training business text. The adjustment fluctuation parameter is used here as a random noise to perturb the vector. This adjustment fluctuation parameter is randomly generated and is in the same vector domain as the training business text representation vector, and the vector dimensions in the corresponding vector domain are the same.
[0041] Continuing with the above power business text as an example, assume the training business text representation vector V tb = [v tb1 , v tb2 , …, v tbm (where m is the dimension of the vector), and the adjustment fluctuation parameter (random mask) M generated by the computer system = [m1, m2, …, m m , where m i is a value randomly generated within a certain range. For example, m i can randomly take values between -∈ and ∈ (∈ is a preset positive number). Then, random perturbation is achieved by adding the training business text representation vector and the adjustment fluctuation parameter to obtain the initial perturbed representation vector V ip , that is, V ip = V tb + M = [v tb1 + m1, v tb2 + m2, …, v tbm + m m . The purpose of this random perturbation is to increase the diversity of data, enabling the model to learn more different situations. Just like in the power business, the ways users ask questions and their expression habits may vary in many ways, and through this random perturbation, the model can better handle various possible situations.
[0042] Finally, it is step S230. The computer system performs vector encoding on the initial perturbed representation vector to obtain the perturbed representation vector of the training business text. Vector encoding is a process of further processing the initial perturbed representation vector, and various technical means can be adopted for this process. One way is to use a neural network for encoding. For example, a multi-layer perceptron (MLP) can be used for vector encoding.
[0043] Assume the initial perturbed representation vector V ipEnter as input into a multi-layer perceptron with l layers. The neurons in the first layer receive the input vector V ip , and perform calculations through the weight matrix W_1 and the bias vector b_1. For the j-th neuron, its output o 1j can be expressed as where σ is the activation function, such as the tanh or ReLU function. Then, the output of the first layer serves as the input to the second layer, and so on. After l layers of calculations, the final output vector is obtained, and this output vector is the perturbation representation vector V of the training service text p . Through this vector encoding method, feature extraction and information integration can be performed on the initial perturbation representation vector, making the obtained perturbation representation vector more conducive to subsequent operations, such as performing perturbation removal operations on it according to the number of perturbation removals in step S300, etc. This series of operations on the training service text from text embedding to random perturbation and then to vector encoding is to enable the network to better adapt to different input situations and improve the generalization ability of the network when training the power service customer service Q&A network, so as to be able to answer various questions related to power services more accurately, whether it is about power usage, power charges, or power equipment maintenance, etc.
[0044] In the actual power service customer service scenario, this processing method helps to handle complex and changing user needs. For example, power users in different regions may have different ways of expressing the same power service concept, or in different service scenarios (such as residential electricity use and industrial electricity use), the focus and expression of the questions will also be different. The perturbation representation vector obtained by processing the training service text through steps S210 - S230 can enable the power service customer service Q&A network to better learn these different situations, so as to give more accurate and appropriate answers when actually answering questions. For example, for an industrial electricity user's inquiry about power load adjustment, the network can accurately answer the operation process, precautions, and relevant policies regarding power load adjustment based on the information of various similar situations learned before. Similarly, for a residential electricity user's question about electricity fee preferential policies, it can also answer accurately, improving the service quality and efficiency of the power service customer service.
[0045] For another example, during the promotion of power services, there may be issues related to the connection of new energy to the power grid. If the training service text is "The connection of new energy power generation equipment to the power grid needs to meet certain technical standards and safety requirements", the computer system processes it according to the above steps. In the text embedding operation of step S210, the words in this text are converted into vectors and combined to obtain the training service text representation vector. Then, in step S220, a random adjustment fluctuation parameter is added to obtain the initial perturbation representation vector, which simulates various expression changes that may occur in actual services. Finally, in step S230, the perturbation representation vector is obtained through vector encoding. This perturbation representation vector contains sufficient information to enable the power service customer service Q&A network to better learn the answering patterns for questions related to the connection of new energy to the power grid, so as to accurately answer users' questions about the technical standards, application processes, safety guarantees, etc. regarding the connection of new energy to the power grid in actual Q&A sessions.
[0046] In addition, during the training process of the power service customer service Q&A network, different training service texts may have different characteristics and requirements. For example, the training service text regarding power equipment fault reporting may need to pay more attention to accuracy and details, while the training service text regarding power service satisfaction surveys may focus more on language style and emotional tendency. By processing these different types of training service texts through steps S210 - S230, the power service customer service Q&A network can be better optimized for different types of questions in the subsequent training process, improving the overall Q&A quality. For example, for questions regarding power equipment fault reporting, the network can accurately identify key information such as the fault type and location and provide the correct reporting process and precautions; for questions regarding power service satisfaction surveys, the network can interact with users in an appropriate language style and positive emotional tendency, increasing user participation and satisfaction.
[0047] As an implementation method, in step S200, obtaining the number of perturbation removals relied on by the power service customer service Q&A network during the iterative calibration process includes:
[0048] Step S240: Obtain the real - time calibration link that the power service customer service Q&A network is currently in, and the total number of removals a when removing perturbations from the perturbation representation vector;
[0049] Step S250: Based on the total number of removals a, determine the corresponding real - time number range during the process of adjusting network parameters by the power service customer service Q&A network in the real - time calibration link; the number ranges corresponding to the process of adjusting network parameters by the power service customer service Q&A network in different calibration links are different; where a is a positive integer greater than 0;
[0050] Step S260: Determine an arbitrary value within the real-time number range to obtain a selected number, and determine the selected number as the number of disturbance removals relied on during the iterative tuning process of the power service customer service Q&A network in the real-time tuning session.
[0051] In step S240, the computer system obtains the real-time tuning session that the power service customer service Q&A network is currently in, and the total number of removals a when removing disturbances from the disturbance characterization vector. The real-time tuning session is different stages in the iterative tuning process of the power service customer service Q&A network, and each stage has different tuning focuses. For example, during the training process of the power service customer service Q&A network, there may be tuning sessions for answer accuracy, answer style matching, and answer tendency level, etc. The total number of removals a is a preset positive integer, indicating the total number of times to be carried out during the entire disturbance removal process.
[0052] Taking the Q&A related to electricity bill query in the power service as an example, if the network is currently being trained to optimize the accuracy of answers regarding electricity bill query, then at this time the computer system determines that it is in the real-time tuning session for answer accuracy. Suppose in this training scenario, after pre-analysis and setting, the total number of removals a is determined to be 10 times. The setting of the total number of removals a here may be based on various factors, such as the scale of the training data, the complexity of the network, and the expected training effect. If the training data scale is large and the network complexity is high, a relatively large total number of removals can be set to ensure that the network parameters can be fully adjusted, enabling the network to better learn the features and rules in the data.
[0053] Next, in step S250, the computer system determines the corresponding real-time number range during the process of adjusting the network parameters of the power service customer service Q&A network based on the total number of removals a. Since different calibration links have different goals and requirements, the corresponding number ranges during the process of adjusting network parameters in different calibration links are also different. Continuing with the above example of electricity bill query, if the current is in the real-time calibration link for answer accuracy, this link may pay more attention to the retention of original information and accurate expression. According to experience or pre-set experiments, the corresponding number range for this link may be a relatively small value range. Suppose this range is a numerical range greater than or equal to 3 and less than or equal to 5. This is because in the answer accuracy calibration link, excessive disturbance removal can introduce unnecessary errors, as accuracy often depends on the accurate information in the original training business text. If it is a calibration link for answer style matching, more adjustments to the text may be needed to make it meet the style requirements, and the corresponding number range can be larger, such as greater than 5 and less than or equal to 8. For the calibration link of the tendency level, its number range will also be set according to its own characteristics and requirements, and may partially overlap or be completely different from the answer style matching calibration link, depending on the specific training objectives and the characteristics of the power service customer service Q&A.
[0054] The computer system can determine this real-time number range through a predefined mapping relationship or function. For example, a function f(t,a) can be defined, where t represents the type of real-time calibration link (such as accuracy, style matching, tendency level, etc.), and a represents the total number of removals. This function returns the corresponding real-time number range according to different t values and a values. This mapping relationship can be pre-determined based on a large amount of experimental data or expert experience.
[0055] In step S260, the computer system determines an arbitrary value within the real-time number range to obtain the selected number, and determines the selected number as the number of disturbance removals relied on during the iterative calibration process of the power service customer service Q&A network in the real-time calibration link.
[0056] Still taking the example of electricity bill query, if the range of real-time times determined in the real-time calibration link of answer accuracy is greater than or equal to 3 and less than or equal to 5, the computer system can determine a selected number through a random selection algorithm (such as uniform random selection within this range). Assuming that the computer system randomly selects 4 as the selected number, then this 4 is determined as the number of disturbance removals relied on in the iterative calibration process of the power service customer service Q&A network in the real-time calibration link of answer accuracy. This selected number will play a key role in subsequent steps (such as the operation of removing disturbances from the disturbance characterization vector in step S300), which determines the degree of disturbance removal from the disturbance characterization vector, and thus affects the quality of the finally generated predicted response text.
[0057] As an implementation manner, the real-time calibration link is one of the following three links: the style matching calibration link, the content quality calibration link, and the tendency level calibration link;
[0058] Among them, in the process of adjusting the network parameters of the power service customer service Q&A network in the style matching calibration link, the corresponding number range is a numerical range greater than or equal to e and less than or equal to f, where both e and f are positive integers not greater than a;
[0059] In the process of adjusting the network parameters of the power service customer service Q&A network in the content quality calibration link, the corresponding number range is a numerical range greater than f and less than or equal to g, where g is a positive integer not greater than a;
[0060] In the process of adjusting the network parameters of the power service customer service Q&A network in the tendency level calibration link, the corresponding number range is a numerical range greater than f and less than or equal to h, where h is a positive integer not greater than a.
[0061] In the embodiment of the present invention, the real-time calibration link is an important part of the entire training process. It includes the style matching calibration link, the content quality calibration link, and the tendency level calibration link, and the computer system has different implementation manners for each link.
[0062] In the power service customer service Q&A scenario, the style matching calibration link aims to make the generated response text conform to the pre-set requirements in terms of style. For example, the power service customer service answer may need to maintain styles such as professional, concise, and friendly. When the computer system adjusts the network parameters of the power service customer service Q&A network in this link, it will operate according to a specific number range, which is a numerical range greater than or equal to e and less than or equal to f, where both e and f are positive integers not greater than a.
[0063] To achieve style matching and calibration, one way is to utilize predefined style templates or style feature libraries. Suppose there is a style feature library that contains characteristic information such as typical vocabulary and grammatical structures of various styles. For example, for a professional style, it may include professional terms such as "electric load", "voltage level", "kilowatt", etc., and rigorous sentence structures, such as longer compound sentences and fewer colloquial expressions. The computer system can compare the predicted response text with the features in the style feature library to calculate the degree of style difference. If the predicted response text is represented as T and the standard style in the style feature library is represented as S, a style difference function D(T, S) can be defined. This function can quantify the style difference by calculating factors such as the occurrence frequency of specific style vocabulary and the similarity of sentence structures in the text. For example, where n is the number of style features considered, w i is the weight of the i-th feature, f i (T) and f i (S) are the quantified values of the i-th feature in text T and standard style S respectively (such as the occurrence frequency of a certain vocabulary).
[0064] In this step, the range of the number of perturbation removals is set based on the characteristics of style adjustment. Since style adjustment relatively focuses more on the overall expression style and does not require overly in-depth large-scale reconstruction of the text, the range of the number of times is relatively moderate. For example, when a customer service representative in the power business answers a question about power equipment maintenance, if the style of the original training business text is relatively concise and clear, and the predicted response text is overly verbose and complex, the computer system in this style matching and calibration step adjusts the network parameters a limited number of times according to the set range of the number of times (such as e = 2, f = 4) to make the generated response text gradually approach the concise and clear style.
[0065] Next is the content quality calibration step. The content quality calibration step focuses on aspects such as the accuracy and integrity of the response text content. In the power business, accurate and complete information is crucial for users. For example, when a user asks about the power failure repair process, the response text must include accurate repair channels, required information, etc.
[0066] When the computer system adjusts network parameters in this step, the corresponding number range is a numerical range greater than f and less than or equal to g, where g is a positive integer not greater than a. To evaluate the content quality, the computer system can adopt technical means based on the knowledge graph. Suppose there is a knowledge graph of power services, which contains information such as the relationships between various power service concepts and operation processes. The computer system can match the content in the predicted response text with the relevant knowledge in the knowledge graph. For example, if the predicted response text mentions the calculation method of electricity charges, the computer system will search for the correct calculation method in the knowledge graph and compare the differences between the two. If C(T) represents the content information in the predicted response text T and K represents the relevant knowledge in the knowledge graph, then a content quality evaluation function Q(C(T), K) can be defined. This function can be quantified according to factors such as the degree of content matching and the inclusion of key information. For example, if the key information is completely matched, the function value may be 1, indicating high content quality; if key information is missing or there are incorrect information, the function value will decrease accordingly.
[0067] Since content quality calibration requires more in-depth adjustment of text content and may require more perturbation removal operations to ensure the accuracy and integrity of the content, its corresponding number range is greater than that of the style matching calibration step. For example, in a Q&A about power safety knowledge, if there are some missing or incorrect key information in the predicted response text, the computer system will adjust the network parameters according to this number range in this step (assuming f = 4, g = 6), and improve the content quality of the generated response text through multiple adjustments.
[0068] Finally, there is the inclination level calibration step. In the Q&A of power service customer service, the inclination level calibration step involves aspects such as the emotional inclination and service orientation of the response text. For example, the answers of power service customer service should have a positive emotional inclination to improve user satisfaction. When users have some questions or complaints about power services, the customer service answers should respond with a positive attitude to solve the problems.
[0069] When the computer system adjusts network parameters in the inclination level calibration step, the corresponding number range is a numerical range greater than f and less than or equal to h, where h is a positive integer not greater than a. The computer system can adopt sentiment analysis technology to achieve inclination level calibration. For example, using a pre-trained sentiment analysis model, this model can classify the emotional inclination in the text (such as positive, neutral, negative) and give the intensity value of the emotional inclination.
[0070] If the predicted response text T is input into the sentiment analysis model to obtain the emotional inclination intensity value E(T), it can be based on the preset positive inclination standard value E pTo evaluate the tendency level of the response text. For example, if E(T) ≥ E p , it indicates that the sentiment tendency of the response text meets the requirements; if E(T) < E p , adjustment is required.
[0071] The number range of the tendency level calibration link is similar to that of the content quality calibration link because adjusting the tendency level also requires a certain degree of in-depth modification to the text. For example, when responding to a power user's complaint about high electricity bills, if the predicted sentiment tendency of the response text is not positive enough, the computer system adjusts the network parameters according to the set number range in this link (assuming f = 4, h = 7), so that the generated response text has a more positive sentiment tendency. For example, express understanding of the user in the answer and provide some possible solutions to improve the user's satisfaction with the power service.
[0072] As an implementation manner, in step S300, the perturbation representation vector is subjected to perturbation removal according to the number of perturbation removals to obtain a perturbation-removed representation vector, including:
[0073] Step S310: Perform the corresponding number of perturbation removals on the perturbation representation vector based on the number of perturbation removals to obtain an initial perturbation-removed representation vector;
[0074] Step S320: According to the random perturbation process of the training service text, perform a vector reduction and reconstruction operation on the initial perturbation-removed representation vector to obtain the perturbation-removed representation vector corresponding to the perturbation representation vector.
[0075] In step S310, the computer system performs the corresponding number of perturbation removals on the perturbation representation vector based on the number of perturbation removals to obtain an initial perturbation-removed representation vector. The perturbation representation vector is a vector containing the perturbation information of the training service text obtained through previous steps (such as the operations in step S200). The number of perturbation removals is determined during the implementation of step S200, and this number determines the degree of adjustment of the perturbation representation vector.
[0076] For example, assume the perturbation representation vector is V p = [v p1 , v p2 , …, v pm (where m is the dimension of the vector) and the number of perturbation removals is k. The computer system can adopt a strategy of gradually restoring to perform perturbation removal. Each removal operation can be defined as a function R(V_p, i) (where i represents the i-th removal operation, 1 ≤ i ≤ k). This function can adjust the vector according to certain rules. A possible rule is to gradually reduce the perturbation components in the vector according to a preset ratio α (0 < α < 1).
[0077] Specifically, in the i-th operation, for each dimension j (1 ≤ j ≤ m) of the vector, the new vector value v p′j,i can be calculated by the formula v p′j,i =(1 - α)v pj,i-1 +αv j0 , where v pj,i-1 is the value of the j-th dimension of the vector V p after the (i - 1)-th operation, and v j0 is the value of the j-th dimension of the original unperturbed vector (assumed to be V0 = [v 01 , v 02 , …, v 0m ). After k such operations, the resulting vector is the initial perturbation removal representation vector V ipr .
[0078] Taking an example in the power business customer service Q&A, assume that the training business text is about the description of the maintenance time of power equipment. The original unperturbed vector representation may accurately correspond to the relevant semantic information of the maintenance time. The vector V p after perturbation may deviate from the original semantics in some dimensions. When performing perturbation removal, if the number of perturbation removal times k = 3, in each operation, the vector is gradually adjusted according to the above rules, making the vector gradually approach the original unperturbed semantic direction. The initial perturbation removal representation vector V ipr obtained after these 3 operations is closer to the original semantic information than the perturbed representation vector V p .
[0079] Next, in step S320, the computer system performs a vector reduction and reconstruction operation on the initial perturbation removal representation vector according to the random perturbation process of the training business text, to obtain the perturbation removal representation vector corresponding to the perturbed representation vector. This step further optimizes the initial perturbation removal representation vector on the basis of the previous steps, making it more in line with the semantics and requirements of the original training business text.
[0080] Since in the previous random perturbation process (the operation in step S200), the perturbation is applied to the training business text representation vector according to a certain pattern or rule, in this vector reduction and reconstruction operation, the inverse process or relevant information of this perturbation process can be used for the operation. For example, if in the random perturbation process, additive perturbation is performed on the vector according to a specific random mask (adjusting the fluctuation parameter) (such as V ip =V tb +M, where V ip is the initial perturbed representation vector, V tbIf it is to train the business text representation vector and M is the random mask, then in the vector restoration and reconstruction operation, the initial perturbation can be removed from the representation vector by inversely adjusting according to the relevant information of this random mask.
[0081] Suppose the random mask M = [m1, m2, …, m m , in the vector restoration and reconstruction operation, for the initial perturbation removal representation vector V ipr = [v ipr1 , v ipr2 , …, v iprm , it can be adjusted through a restoration function F(V ipr , M). A possible form of the restoration function can be v prj = v iprj - βm j (where v prj is the value of the j-th dimension of the perturbation removal representation vector V pr , β is a coefficient related to the perturbation application and removal, 1 ≤ j ≤ m). By adjusting each dimension of the initial perturbation removal representation vector in this way, the perturbation removal representation vector V pr is finally obtained.
[0082] Taking the example of the maintenance time of power equipment again, if during the random perturbation process, the random mask increases certain values in some dimensions, causing the vector to deviate from the original semantics, then in the vector restoration and reconstruction operation, according to the information of this random mask, a certain value (determined by β and m j ) is subtracted in the corresponding dimension, so that the obtained perturbation removal representation vector V pr can more accurately reflect the semantic information of the original training business text regarding the maintenance time of power equipment. This operation helps to improve the accuracy of the subsequent generated prediction response text, making the generated prediction response text closer to the requirements of the original training business text in terms of semantics and style, so as to provide more accurate and appropriate answers in the power business customer service Q&A, whether it is about power equipment maintenance, electricity bill inquiry or other power business-related questions.
[0083] As an implementation manner, in step S400, obtain the quality evaluation index of the prediction response text, including:
[0084] Step S410: Obtain the real-time evaluation score for content quality evaluation in the real-time calibration link; wherein, the evaluation score includes one of the following dimensions: style matching score, content quality score and tendency level, and one calibration link corresponds to one evaluation score;
[0085] Step S420: Obtain the score value of the prediction response text under the real-time evaluation score, and obtain the critical score value corresponding to the real-time evaluation score;
[0086] Step S430: Compare the score value under the real-time evaluation score with the critical score value, and obtain the quality evaluation index of the predicted response text based on the comparison result.
[0087] In step S410, the computer system obtains the real-time evaluation score for content quality evaluation in the real-time calibration session. In the context of power business customer service Q&A, the evaluation score covers different dimensions, including style matching score, content quality score, and tendency level, etc., and one calibration session corresponds to one evaluation score. These evaluation scores are designed to measure the degree of conformity between the predicted response text and the ideal response text from different aspects.
[0088] Taking the Q&A about electricity bill inquiry in the power business as an example, in the style matching calibration session, the style matching score is used to evaluate whether the style of the predicted response text conforms to the pre-set style requirements. If the power company stipulates that the customer service answer should adopt a concise, professional and friendly style, then for the predicted response text "You can log in to the power company's official website to inquire about your electricity bill. The operation is simple. Wish you a happy life.", the computer system will determine its style matching score according to certain rules. For example, the computer system can pre-define a set of style features, which includes features related to conciseness (such as short sentences, avoiding complex clauses, etc.), features related to professionalism (such as accurately using power industry terms, etc.), and features related to friendliness (such as using polite expressions, etc.). Then, by counting the number or proportion of these features that conform to the predicted response text, and combining with the pre-set weights (for example, the weight of conciseness is 0.3, the weight of professionalism is 0.4, and the weight of friendliness is 0.3), the style matching score is calculated according to the formula S = w1s1 + w2s2 + w3s3 (where S is the style matching score, wi is the weight of each feature, and si is the score of each feature).
[0089] In the content quality calibration session, the content quality score focuses on considering the accuracy, integrity, etc. of the content of the predicted response text. For example, for the predicted response text "You can inquire about your electricity bill through the official website or mobile APP. To query through the official website, you need to log in to your account. To query through the mobile APP, you need to register an account and bind your electricity user number.", the computer system needs to judge whether it contains the key information of electricity bill inquiry (such as the query channels and corresponding operation requirements). The computer system can evaluate based on the power business knowledge graph. Assuming that the knowledge graph contains the complete operation process, required information, etc. for electricity bill inquiry, the computer system compares the information in the predicted response text with the content in the knowledge graph. If the predicted response text contains most of the key information in the knowledge graph, the content quality score will be higher; on the contrary, if important information is missing or there are incorrect information, the score will be reduced.
[0090] For the propensity level calibration step, the propensity level reflects aspects such as the sentiment tendency and service orientation of the predicted response text. For example, when the user expresses dissatisfaction with the high electricity bill, the predicted response text "Understand your trouble. We will actively optimize the electricity pricing method. At the same time, you can reduce your electricity bill through energy-saving measures" shows a tendency to actively solve the problem. The computer system can determine the propensity level through a pre-trained sentiment analysis model. After inputting the text into the model, the model outputs a value representing the propensity level, for example, within the range of -1 to 1. The closer to 1, the more positive the tendency; the closer to -1, the more negative the tendency; and 0 represents neutrality.
[0091] Next, in step S420, the computer system obtains the score value of the predicted response text under the real-time evaluation score, and obtains the critical score value corresponding to the real-time evaluation score.
[0092] Continuing with the example of electricity bill query. For the above-mentioned style matching score, the computer system obtains the score value of the predicted response text under this score. The computer system can adopt a rule-based scoring system. Suppose for conciseness, a certain score is deducted for each complex clause; for professionalism, a certain score is deducted for each incorrect use of a professional term; for friendliness, a certain score is deducted for each lack of a polite expression. Through such a deduction mechanism, the specific score value of the predicted response text under the style matching score is finally obtained.
[0093] At the same time, the computer system needs to obtain the critical score value corresponding to the real-time evaluation score. To obtain this critical score value, the computer system first obtains the training database, which contains one or more binary data templates, and each binary data template contains business text and style constraint information of the business text. The computer system extracts the text representation vector of each business text. For example, the business text is converted into a vector form through text embedding technology (such as Word2Vec, etc.). Then, based on the text representation vectors of the business texts, one or more business texts are grouped to obtain s classifications (s is a positive integer greater than 1).
[0094] For example, in business texts related to electricity bill inquiries, they may be classified according to the inquiry channels (such as online official website inquiries, APP inquiries, offline business hall inquiries, etc.) or user types (residential users, commercial users, etc.). For each classification, x business texts are selected (x is a positive integer greater than 1), and text reconstruction is carried out through the power business customer service Q&A network based on the selected business texts and the corresponding style constraint information to obtain x reconstructed texts for each classification. For each reconstructed text, the computer system determines its score value under the real-time evaluation score. Assuming for the style matching score, the computer system calculates the score value of each reconstructed text according to the previously mentioned scoring system. Then, a statistical value (such as the median value) is obtained through statistics on the obtained score values, and this statistical value is used as the critical score value of the real-time evaluation score.
[0095] Finally, in step S430, the computer system compares the score value under the real-time evaluation score with the critical score value, and obtains the quality evaluation index of the predicted response text based on the comparison result.
[0096] Still taking the electricity bill inquiry as an example, if the score value of the predicted response text under the style matching score is not less than the critical score value, this indicates that the predicted response text meets a certain standard in terms of style. The computer system determines the corresponding quality evaluation index as the first quality evaluation index. For example, the first quality evaluation index can indicate that the predicted response text is qualified in terms of style and can continue to be evaluated in other aspects (such as content quality, tendency level, etc.) or directly used for subsequent network parameter adjustment. If the comparison result shows that the score value is less than the critical score value, then the computer system determines the corresponding quality evaluation index as the second quality evaluation index. For example, the second quality evaluation index may indicate that there are deficiencies in the style of the predicted response text, and key optimization needs to be carried out for the style aspect in subsequent network parameter adjustment.
[0097] During the training process of the power service customer service Q&A network, the acquisition of such quality evaluation indicators is crucial for the optimization of the entire network. For example, in the Q&A scenario of power equipment fault repair, for the predicted response text "Just tell me where it's broken", the score may be relatively low under the style matching score because it lacks professionalism and friendliness. After the computer system determines its corresponding second quality evaluation indicator, it knows that when adjusting the network parameters, it needs to adjust the network so that the generated response text is more in line with the requirements in terms of style, such as becoming more professional (e.g., "Please inform the name of the faulty equipment, the fault phenomenon, and the approximate location, etc., so that the maintenance personnel can handle it accurately"). Similarly, in the Q&A related to power service satisfaction surveys, if the predicted response text "It's okay" is below the critical score value under the content quality score and the second quality evaluation indicator is obtained, it means that the content is too brief and lacks specific service evaluation content. When adjusting the network parameters, the computer system will adjust in the direction of increasing content integrity and accuracy to improve the quality of the predicted response text.
[0098] As an implementation manner, the training binary data is any binary data template in the training database; in step S420, obtaining the score value of the predicted response text under the real-time evaluation score includes:
[0099] Step S421: Determine the target classification where the training service text is located among s classifications, and determine the selected service text in the target classification, and the target score value of the reconstructed text corresponding to the selected service text under the real-time evaluation score;
[0100] Step S422: Generate the score value of the predicted response text corresponding to the training service text under the real-time evaluation score based on the target score value.
[0101] In step S421, the computer system determines the target classification where the training service text is located among s classifications, and determines the selected service text in the target classification, and the target score value of the reconstructed text corresponding to the selected service text under the real-time evaluation score.
[0102] In the context of power business customer service Q&A, assume that the business texts in the training database are various Q&A information related to power business. The computer system first classifies them through feature analysis of the business texts. For example, based on the types of electricity consumption in power business, the business texts can be classified into different categories such as residential electricity consumption, industrial electricity consumption, commercial electricity consumption, etc.; or according to the nature of the business, they can be classified into power equipment maintenance, electricity bill inquiry, power policy consultation, etc. Here, a total of s categories are formed. Taking the electricity bill inquiry business as an example, assume that the computer system has classified the business texts according to the inquiry channels (such as online platform inquiry, offline business hall inquiry, etc.), which is a part of the s categories. When evaluating a predicted response text, the computer system needs to determine in which category the training business text corresponding to this predicted response text is, and this category is the target category. For example, a predicted response text about inquiring about electricity bills on the online platform, its corresponding training business text may be in the target category of online platform inquiry.
[0103] After determining the target category, specific business texts are selected in this target category, that is, the selected business texts. These selected business texts are representative and may be selected according to a certain sampling strategy, such as random sampling or according to certain priority rules (such as selecting the most frequently occurring business texts, etc.). For these selected business texts, the computer system uses the power business customer service Q&A network to reconstruct the text based on the selected business texts and the corresponding style constraint information. This process is similar to simulating Q&A processing on the original business texts, and the result is the reconstructed text.
[0104] For determining the target score value, the computer system can adopt various technical means. If it is to evaluate the content quality score, the computer system can set a series of evaluation criteria. For example, a certain basic score can be given for an answer that contains complete operation steps (such as opening the APP, logging in to the account, finding the inquiry entry, viewing details); additional points can be added for the detail level of each key step (such as whether it mentions that logging in to the account requires entering the correct username and password, etc.); points will be deducted if there is incorrect information in the answer. Assume that the basic score is b, the additional points for the detail level of each key step are p1, p2,..., and the deduction item for incorrect information is d. Then the target score value S can be calculated by the formula S = b + ∑ i p i - d (where i represents different additional points).
[0105] In step S422, based on the target score value, the computer system generates the score value of the predicted response text corresponding to the training business text under the real-time evaluation score.
[0106] Since the target score value of the reconstructed text corresponding to the selected service text under the real-time evaluation score has been obtained in step S421, the computer system can use this target score value and the relationship between the predicted response text and the selected service text to determine the score value of the predicted response text.
[0107] For example, the computer system can make adjustments by calculating the semantic similarity between the predicted response text and the selected service text. Suppose the computer system adopts a semantic similarity calculation method based on word vectors, such as cosine similarity. For the selected service text T1 and the predicted response text T2, they are first converted into word vector representations V1 and V2, and then their cosine similarity is calculated. If the cosine similarity is high, it indicates that the predicted response text is semantically close to the selected service text. Then, the score value of the predicted response text can be adjusted according to a certain ratio based on the target score value. Suppose this ratio is r (0 < r < 1), then the score value S of the predicted response text p can be calculated by the formula S p = S × (1 + r × (cos(θ) - 0.5)) (here 0.5 is an intermediate reference value. When cos(θ) = 0.5, the score value of the predicted response text is equal to the target score value).
[0108] As an implementation manner, in step S420, obtaining the critical score value corresponding to the real-time evaluation score includes:
[0109] Step S423: Obtain a training database. The training database contains one or more binary data templates, and each binary data template contains a service text and the style constraint information of the service text.
[0110] Step S424: Extract the text representation vectors of each service text, and group one or more service texts according to the text representation vectors of the service texts to obtain s classifications; s is a positive integer greater than 1.
[0111] Step S425: Select x service texts in each classification, and perform text reconstruction through the power business customer service question-answering network according to the selected service texts and the corresponding style constraint information to obtain x reconstructed texts for each classification; x is a positive integer greater than 1.
[0112] Step S426: Determine the score value of each reconstructed text under the real-time evaluation score, and perform statistics on the obtained score values to obtain a statistical value as the critical score value of the real-time evaluation score.
[0113] In step S423, the computer system obtains a training database. This training database is the basis for the entire evaluation process and contains one or more binary tuple data templates. In the context of power business customer service Q&A, these binary tuple data templates are crucial for constructing and optimizing the Q&A model. Each binary tuple data template contains business text and style constraint information for the business text.
[0114] In step S424, the computer system extracts the text representation vectors of each business text and groups one or more business texts based on the text representation vectors of the business texts to obtain s classifications, where s is a positive integer greater than 1.
[0115] For extracting text representation vectors, the computer system can adopt a variety of text vectorization techniques. For example, using the Bag-of-Words model or more advanced word vector models (such as Word2Vec, GloVe, etc.). Taking Word2Vec as an example, it can map each word to a low-dimensional vector space, making words with similar semantics closer in the vector space. After the computer system vectorizes each word in the business text, the text representation vector of the entire business text is obtained through a certain combination method (such as simple average or weighted average).
[0116] Suppose there is a series of power business texts, such as texts of different types regarding power equipment maintenance, electricity bill inquiries, power safety knowledge promotion, etc. After the computer system vectorizes these business texts, they are grouped according to the similarity between the text representation vectors. For example, for business texts related to power equipment maintenance, their distributions in the vector space may be relatively close because these texts may involve similar vocabulary (such as equipment names, maintenance tools, maintenance procedures, etc.), so they can be grouped into one category. Similarly, business texts related to electricity bill inquiries will be grouped into another category because they will contain similar vocabulary regarding inquiry channels, cost calculations, etc. In this way, s classifications are finally obtained.
[0117] In step S425, x business texts are selected from each classification, and text reconstruction is performed through the power business customer service Q&A network based on the selected business texts and the corresponding style constraint information to obtain x reconstructed texts for each classification, where x is a positive integer greater than 1.
[0118] When the computer system selects business texts in each classification, it can use random sampling or select them according to a certain priority. For example, in the power equipment maintenance classification, if there are some business texts marked as typical cases, these texts can be preferentially selected. Then, the selected business texts and the corresponding style constraint information are input into the power business customer service Q&A network.
[0119] The power service customer service Q&A network is a model that has been pre-trained or is in the process of being trained. It can generate new text, that is, reconstruct the text, based on the input business text and style constraint information. Taking the classification of power equipment maintenance as an example, if the selected business text is "What are the precautions for the daily maintenance of power transformers?", and the corresponding style constraint information is that the answer should be detailed, professional and well-organized, the power service customer service Q&A network can generate the following reconstructed text: "The following points should be noted for the daily maintenance of power transformers: First, regularly check the oil temperature to ensure that the oil temperature is within the normal range, generally not exceeding [specific temperature value]; Second, check the oil level. If the oil level is too low, it can affect the normal operation of the transformer; Third, check the appearance of the transformer to see if there is any oil leakage; Finally, check the wiring terminals of the transformer to ensure firm connection and avoid looseness causing potential safety hazards."
[0120] By operating in this way for each classification, x reconstructed texts for each classification can be obtained.
[0121] In step S426, the computer system determines the score value of each reconstructed text under the real-time evaluation score, and performs statistics on the obtained score values to obtain a statistical value (such as the median value) as the critical score value of the real-time evaluation score.
[0122] The process of determining the score value of each reconstructed text under the real-time evaluation score is similar to the method of evaluating the predicted response text mentioned above. Taking the content quality score as an example, if the real-time evaluation score is the content quality score, the computer system will score the reconstructed text according to the pre-set evaluation criteria for content quality. Suppose for the reconstructed text of power equipment maintenance, the evaluation criteria include whether all maintenance points are completely mentioned and whether the description of each point is accurate and detailed. If a reconstructed text completely mentions all important maintenance points and the description of each point is detailed and accurate, then it will get a higher score under the content quality score; on the contrary, if there are missing points or vague descriptions, the score will be lower.
[0123] For the x reconstructed texts in each classification, the computer system obtains their score values under the real-time evaluation score in this way. Then, the computer system conducts statistical analysis on these score values. Taking the calculation of the median value as an example, assume that the score values of the x reconstructed texts obtained in a certain classification are f1, f2, …, f xFirst, sort these fractional values in ascending (or descending) order. If x is odd, then the median is the middle fractional value; if x is even, the median is usually the average of the two middle fractional values. For example, if x = 5 and the sorted fractional values are f1 = 60, f2 = 70, f3 = 75, f4 = 80, f5 = 85, then the median is f3 = 75. This median is determined as the critical fractional value for real-time evaluation scores.
[0124] As an implementation, in step S430, based on the comparison result, quality evaluation indicators for the predicted response text are obtained, including:
[0125] Step S431: If the comparison result indicates that the fractional value is not less than the critical fractional value, then determine the corresponding quality evaluation indicator as the first quality evaluation indicator;
[0126] Step S432: If the comparison result indicates that the fractional value is less than the critical fractional value, then determine the corresponding quality evaluation indicator as the second quality evaluation indicator; where the first quality evaluation indicator and the second quality evaluation indicator are different.
[0127] In step S431, if the comparison result indicates that the fractional value is not less than the critical fractional value, then the computer system determines the corresponding quality evaluation indicator as the first quality evaluation indicator.
[0128] In an embodiment of the present invention, taking content quality evaluation as an example, assume that a predicted response text regarding power equipment maintenance is being evaluated. The computer system has obtained the fractional value of the predicted response text under the content quality score, as well as the corresponding critical fractional value. If the fractional value of the predicted response text is not less than the critical fractional value, this means that the predicted response text meets a certain standard in terms of content quality.
[0129] For example, the critical fractional value is set to 70 points, and the fractional value of the predicted response text under the content quality score is 75 points. This score indicates that the predicted response text contains relatively complete and accurate power equipment maintenance information. The computer system may adopt a technical means based on threshold judgment to determine the quality evaluation indicator. Here, since the fractional value meets the condition of not being less than the critical fractional value, the computer system determines the corresponding first quality evaluation indicator.
[0130] The first quality assessment indicator can be an identifier indicating that the predicted response text is in a qualified or good state in this assessment dimension (here it is content quality). From a technical implementation perspective, a binary value (such as 1) can be used to represent the first quality assessment indicator, and another value (such as 0) represents other situations. Or a coding method with more semantic information can also be used. For example, if a three-digit coding is adopted, the first digit indicates whether the critical score value is reached (1 means reached, 0 means not reached), and the last two digits can be used to represent more detailed levels (such as 01 means just reaching the critical score value, 10 means significantly exceeding the critical score value, etc.). Here, if 75 points significantly exceed 70 points, it can be coded as 110. This coding method helps to more carefully process the predicted response texts in different quality states during the subsequent network parameter adjustment process.
[0131] The same is true for the style matching assessment. Suppose for a power business customer service answering questions about electricity bill inquiries, the style requirement is concise, professional, and friendly. The computer system calculates the score value of the predicted response text under the style matching score. If this score value is not lower than the critical score value, it means that the predicted response text meets the requirements in terms of style. For example, the critical score value is 80 points, and the score value of the predicted response text is 85 points. The computer system determines the corresponding first quality assessment indicator. This indicator can help determine that when adjusting the network parameters, there is no need to make large adjustments to the network parameters related to style because the predicted response text has reached an acceptable level in terms of style.
[0132] In step S432, if the comparison result shows that the score value is less than the critical score value, the computer system determines that the corresponding quality assessment indicator is the second quality assessment indicator.
[0133] Continuing with the example of the content quality assessment of power equipment maintenance, if the score value of the predicted response text under the content quality score is 60 points, and the critical score value is 70 points, and the score value is less than the critical score value, this indicates that there are deficiencies in the content quality of the predicted response text. The computer system determines the corresponding second quality assessment indicator.
[0134] The second quality assessment indicator can also adopt a similar identification method to the first quality assessment indicator, but it represents the state of not meeting the standard. For example, using the binary value 0 to represent the second quality assessment indicator, or in a more complex coding method, the first digit being 0 means not reaching the critical score value, and the last two digits can be further subdivided according to the gap from the critical score value (such as 00 means a large gap from the critical score value, 01 means close to the critical score value, etc.). Here, the gap between 60 points and 70 points is large, and it can be coded as 000.
[0135] In terms of the propensity level assessment, for example, when a power business customer service answers a user's complaint about high electricity bills, the propensity level requirement is to respond positively and provide possible solutions. If the score of the predicted response text under the propensity level score is less than the critical score, the computer system determines it as the second quality assessment indicator. This means that the predicted response text does not meet the requirements in terms of the propensity level, and the answer may be relatively negative or no effective solution is provided.
[0136] This way of determining the quality assessment indicator is very crucial in the training process of the power business customer service Q&A model. For example, when the computer system adjusts the network parameters of the power business customer service Q&A network, the first quality assessment indicator and the second quality assessment indicator can be used as important bases. If the predicted response text corresponds to the first quality assessment indicator, the network parameter adjustment may focus more on optimizing other aspects or making fine-tuning; while if it is the second quality assessment indicator, the network parameter adjustment needs to focus on adjusting this assessment dimension.
[0137] As an implementation manner, when the real-time calibration link where the power business customer service Q&A network is located is the style matching calibration link, the evaluation score corresponding to the style matching calibration link is the style matching score. Based on this, in step S400, generating the backpropagation optimization error based on the quality assessment indicator, the predicted response text, and the response style constraint information includes:
[0138] Step S440: Determine the initial backpropagation optimization error in the style matching calibration link based on the predicted response text and the response style constraint information;
[0139] Step S450: Obtain the target backpropagation optimization error in the style matching stage according to the quality assessment indicator and the initial backpropagation optimization error in the style matching calibration link;
[0140] Step S460: Obtain the reconstruction error when the power business customer service Q&A network performs text reconstruction, and use the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power business customer service Q&A network in the style matching calibration link.
[0141] In step S440, based on the predicted response text and the response style constraint information, the computer system determines the initial backpropagation optimization error in the style matching calibration link. In the power business customer service Q&A scenario, style matching is the key to ensuring that the response text meets specific style requirements. For example, the answers of power business customer service may need to follow style requirements such as professional, concise, and friendly.
[0142] To determine the initial feedback optimization error, a computer system can adopt a Dual Tower Model. For the predicted response text, the computer system extracts the representation vector of the predicted response text through the Dual Tower Model. Suppose the predicted response text is "You can log in to the official website of the power company to query your electricity bill. The operation is simple. Wish you a happy life." The Dual Tower Model will convert this text into a vector representation in a specific vector space. This vector can, to a certain extent, represent various aspects of the text's semantics, syntactic structure, and style. This is the representation vector of the predicted response text.
[0143] Meanwhile, the computer system extracts the text representation vector corresponding to the response style constraint information of the training service text. The response style constraint information is, for example, "concise, professional, and friendly". The computer system constructs a standard text representation vector based on this style constraint information. This vector represents the ideal state that conforms to this style constraint.
[0144] The computer system determines the vector distance between these two text representation vectors to obtain the initial feedback optimization error in the style matching and calibration phase. The vector distance can be calculated in various ways, such as the Euclidean distance. Suppose the representation vector of the predicted response text is V p =[v p1 ,v p2 ,…,v pn , and the text representation vector corresponding to the response style constraint information is V s =[v s1 ,v s2 ,…,v sn . The Euclidean distance formula is This distance d can be used as a measure of the initial feedback optimization error. If the distance d is large, it indicates that the predicted response text is quite different from the ideal style constraint information in terms of style, and the network needs to be adjusted to a large extent; if the distance d is small, it means that in terms of style matching, it is already relatively close to the ideal state, and the adjustment range of the network can be relatively small.
[0145] Next, in step S450, based on the quality evaluation index and the initial feedback optimization error in the style matching and calibration phase, the target feedback optimization error in the style matching stage is obtained. The quality evaluation index has been obtained in the previous steps, and it reflects the overall quality status of the predicted response text.
[0146] For example, if the quality assessment index indicates that although the predicted response text is overall qualified in terms of style matching, there are still some minor flaws (such as some words not being professional enough), and the initial backpropagation optimization error shows a certain distance in the style representation vector. The computer system can adjust the initial backpropagation optimization error according to the preset rules to obtain the target backpropagation optimization error. A possible rule is that if the quality assessment index shows that the style matching is in a state close to qualified but not entirely ideal, multiply the initial backpropagation optimization error by a coefficient k slightly greater than 1 (for example, k = 1.1) to obtain the target backpropagation optimization error. The determination of this coefficient k can be based on a large amount of experimental data or empirical settings. Its role is to appropriately amplify or reduce the initial backpropagation optimization error according to the quality assessment index to more accurately reflect the degree to which the network needs to be adjusted under the current quality assessment.
[0147] Then, in step S460, the computer system obtains the reconstruction error when the power business customer service Q&A network performs text reconstruction, and takes the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power business customer service Q&A network in the style matching calibration link.
[0148] During the process of the power business customer service Q&A network performing text reconstruction, there will be a certain error. For example, when the network generates a predicted response text based on the input training business text and response style constraint information, due to reasons such as the network's parameter settings and data complexity, the generated text may deviate from the ideal text, and this deviation is the reconstruction error.
[0149] One way for the computer system to obtain the reconstruction error is to compare the differences between the predicted response text and the original training business text (under ideal style constraints). Suppose the original training business text is "You can conveniently query your electricity bill through the official website of the power company", and the predicted response text is "You can log in to the official website of the power company to query your electricity bill, and the operation is simple". Although the semantics are similar, there are some differences in expression. The computer system can analyze these differences, such as calculating the proportion of missing or changed key information in the text, etc., to quantify the reconstruction error.
[0150] Suppose the reconstruction error is e r , and the target backpropagation optimization error is e t , then the total backpropagation optimization error E = e r + e tThis total backpropagation optimization error will be used to adjust the network parameters of the power service customer service Q&A network in the style matching calibration process. For example, in the network parameter adjustment algorithm based on gradient descent, the total backpropagation optimization error will follow the reverse propagation path of the network and adjust parameters such as weights and biases in the network according to the error. If the total backpropagation optimization error is large, the adjustment amplitude of the network parameters will be large; conversely, if the total backpropagation optimization error is small, the adjustment amplitude of the network parameters will be small, thereby gradually optimizing the network, improving the performance in style matching, and making the generated predicted response text more compliant with the response style constraint information.
[0151] When the power service customer service Q&A network is in the content quality calibration process, the steps are similar. In step S400A, the computer system performs content quality assessment processing on the predicted response text through the content quality assessment network to obtain the content quality coefficient of the predicted response text. The content quality assessment network can be a neural network-based model that analyzes the predicted response text according to predefined content quality assessment criteria (such as the accuracy and integrity of information). For example, for a predicted response text about power equipment maintenance, the content quality assessment network will check whether it contains key maintenance steps, precautions, etc. If it contains more key information and is accurate, the content quality coefficient will be high; if key information is missing or there are incorrect information, the content quality coefficient will be low.
[0152] In step S400B, based on the quality assessment index and the content quality coefficient in the content quality calibration process, the target backpropagation optimization error in the content quality process is obtained. If the quality assessment index indicates that there are significant problems with the predicted response text in terms of content quality and the content quality coefficient is low, the computer system can increase the initial backpropagation optimization error (similar to the initial backpropagation optimization error in the style matching process but for content quality) to obtain the target backpropagation optimization error. For example, if the initial backpropagation optimization error is e0 and the adjustment coefficient determined according to the quality assessment index and the content quality coefficient is k' (k'>1, the specific value depends on the situation), then the target backpropagation optimization error e t ′ = k′e0.
[0153] In step S400C, the computer system obtains the reconstruction error when the power service customer service Q&A network performs text reconstruction, and determines the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power service customer service Q&A network in the content quality calibration process. The method of obtaining the reconstruction error here is similar to that in the style matching calibration process, which is determined by comparing the difference between the predicted response text and the ideal text with complete and accurate content. The final total backpropagation optimization error is used to adjust the network parameters in the content quality calibration process to optimize the network in terms of content quality, so that the generated predicted response text can contain more accurate and complete power service-related information.
[0154] When the power business customer service question-and-answer network is in the propensity level adjustment stage, in step S400a, the computer system arranges the predicted answer text by propensity level through the propensity level network to obtain the propensity level of the predicted answer text. The propensity level network can classify the predicted answer text according to predefined propensity level standards (such as positive, neutral, negative, etc.). For example, for the predicted answer text responding to the power user's complaint about the high electricity bill, if the answer contains positive factors such as understanding the user and providing solutions, the propensity level may be positive; if it is just a simple statement of facts without a positive response, the propensity level may be neutral; if the answer is impatient or evasive, the propensity level may be negative.
[0155] In step S400b, the target feedback optimization error in the tendency level link is obtained based on the quality evaluation index and the tendency level. If the quality evaluation index shows that the predicted response text does not meet the requirements in terms of tendency level (e.g., it should be positive but is actually neutral), and the current tendency level is poor, the computer system will adjust the feedback optimization error. For example, an adjustment factor k' is determined based on the quality evaluation index and the tendency level. The initial feedback optimization error is e0', and the target feedback optimization error e0' is then t ”=k”e0”.
[0156] In step S400c, the computer system obtains the reconstruction error when the electric power business customer service question and answer network performs text reconstruction, and determines the reconstruction error and the target feedback optimization error as the total feedback optimization error of the electric power business customer service question and answer network in the propensity level adjustment link. The reconstruction error is obtained in the same way as comparing the difference between the predicted answer text and the ideal text with the correct propensity level. The total feedback optimization error is used to adjust the network parameters in the propensity level adjustment link so that the predicted answer text generated by the network has the correct propensity level, for example, when responding to questions or complaints from electric power users, services can be provided with a positive attitude, thereby improving the service quality of electric power business customer service.
[0157] By accurately determining the total feedback optimization error in different adjustment links (style matching, content quality, and tendency level), and using these errors to adjust the network parameters of the power business customer service question and answer network, the network can be continuously optimized and the quality of the generated predictive answer text in various aspects (style, content, and tendency) can be improved, thereby better meeting the needs of power business customer service questions and answers and providing power users with more accurate, appropriate, and high-quality services.
[0158] As an implementation mode, step S440, based on the predicted response text and the response style constraint information, determines the initial feedback optimization error in the style matching adjustment link, including:
[0159] Step S441: Extract the predicted response text representation vector of the predicted response text through a dual-channel network, and extract the text representation vector corresponding to the response style constraint information of the training service text; the text representation vectors extracted by the dual-channel network and the text representation vector are in the same vector domain.
[0160] Step S442: Determine the vector distance between the text representation vector and the text representation vector, and obtain the initial backpropagation optimization error in the style matching adjustment phase based on the vector distance.
[0161] In step S441, the computer system extracts the predicted response text representation vector of the predicted response text through a dual-channel network, and extracts the text representation vector corresponding to the response style constraint information of the training service text; the text representation vectors extracted by the dual-channel network and the text representation vector are in the same vector domain.
[0162] The dual-channel network is a feature extraction tool. For the predicted response text, such as "You can query the electricity bill through the official website of the power company, and the operation is very convenient", the computer system inputs this text into the dual-channel network. One channel of the dual-channel network is responsible for processing the semantic information of the text, and the other channel may process the syntactic structure or other style-related information. During the processing, the network will convert the text into a vector form according to predefined rules and algorithms. Suppose a method that combines word vector embedding technology based on neural networks with a deep learning architecture is adopted. For each word in the text, the word vector embedding technology will map it to a low-dimensional vector space. For example, the word "power company" may be mapped to an n-dimensional vector [v1, v2, …, v n . Then, through the multi-layer structure of the neural network, these word vectors are combined and transformed, and finally the predicted response text representation vector V p = [v p1 , v p2 , …, v pm of the entire predicted response text is obtained. This vector can represent the semantic, syntactic structure, and style and other multi-faceted features of the predicted response text to a certain extent. At the same time, for the response style constraint information of the training service text, such as the style requirements of "concise, professional, and friendly", the computer system also constructs the corresponding text representation vector through the dual-channel network. This construction process is also based on quantifying and vector representing each element in the style constraint information. Taking "concise" as an example, it may be possible to convert it into a vector representation by counting the vocabulary, syntactic structure, etc. related to conciseness in the style constraint information. Suppose the quantitative representation of "concise" is obtained by calculating factors such as the average word length of the sentence and the number of clauses to get a vector [s1, s2, …, s k; For "professional", another vector representation [p1, p2, …, p may be obtained by counting the usage frequency of professional terms, specific industry expression methods, etc. l ; For "friendly", it may involve the usage frequency of polite expressions, etc., to obtain the vector [f1, f2, …, f m . Then combine these vectors related to the style constraint information to obtain the text representation vector V corresponding to the response style constraint information s = [v s1 , v s2 , …, v sn .
[0163] The two-channel network here ensures that the two vectors extracted (the predicted response text representation vector and the text representation vector corresponding to the response style constraint information) are in the same vector domain, which facilitates subsequent comparison and calculation. This means that they have the same dimension in the vector space, and each dimension has similar semantic meanings. For example, in this vector domain, a certain dimension of the vector may be related to the professionalism degree of the vocabulary, only the specific values in different vectors are different, reflecting the differences in professionalism between the predicted response text and the style constraint information.
[0164] In step S442, the computer system determines the vector distance between the text representation vectors, and obtains the initial backpropagation optimization error in the style matching and tuning link based on the vector distance.
[0165] The vector distance is a measure of the degree of difference between two vectors in the vector space. The computer system can adopt various methods to calculate the vector distance, and the Euclidean distance is a feasible method. Suppose the predicted response text representation vector V p = [v p1 , v p2 , …, v pm , and the text representation vector V corresponding to the response style constraint information s = [v s1 , v s2 , …, v sn (here m = n because they are in the same vector domain), the Euclidean distance formula is
[0166] For example, suppose the predicted response text representation vector V p = [1, 2, 3], and the text representation vector V corresponding to the response style constraint information s = [2, 3, 4], calculated according to the Euclidean distance formula: This is the Euclidean distance between these two vectors.
[0167] This vector distance is used as a measure of the initial backpropagation optimization error in the style matching and tuning phase. If the vector distance is large, it means that the predicted response text has a significant difference in style from the response style constraint information. In this case, in the style matching and tuning phase, the parameters of the power business customer service Q&A network need to be adjusted to a large extent to reduce this distance and make the generated predicted response text more compliant with the style requirements. Conversely, if the vector distance is small, it indicates that the predicted response text is already relatively close to the response style constraint information, and the adjustment range of the network parameters can be relatively small.
[0168] This method of determining the initial backpropagation optimization error based on vector distance is of great significance in the training of the power business customer service Q&A model. Taking the example of answering electricity bill inquiries in the power business, if the predicted response text is "You can go to the official website of the power company to check the electricity bill.", this answer is relatively colloquial and has a large gap from the required style constraint information of "concise, professional, and friendly". The vector distance calculated through the above steps will be large, and the initial backpropagation optimization error will also be large, which prompts the computer system to make significant adjustments to the network in the style matching and tuning phase, such as adjusting the parameters related to word selection and sentence structure construction in the network, so that the subsequent generated predicted response text is more compliant with the requirements in style, such as "You can check the electricity bill through the official website of the power company."
[0169] In this way, the computer system can accurately determine the initial backpropagation optimization error based on the style difference between the predicted response text and the response style constraint information, providing a basis for effectively adjusting the parameters of the power business customer service Q&A network in the subsequent style matching and tuning phase, thereby improving the style matching degree of the predicted response text generated by the power business customer service Q&A network.
[0170] As an implementation, when the real-time tuning phase of the power business customer service Q&A network is the content quality tuning phase, the evaluation score corresponding to the content quality tuning phase is the content quality score; at this time, in step S400, generating the backpropagation optimization error based on the quality evaluation index, the predicted response text, and the response style constraint information includes:
[0171] Step S400A: Perform content quality evaluation processing on the predicted response text through the content quality evaluation network to obtain the content quality coefficient of the predicted response text;
[0172] Step S400B: Obtain the target backpropagation optimization error in the content quality phase based on the quality evaluation index and the content quality coefficient of the content quality tuning phase;
[0173] Step S400C: Obtain the reconstruction error during the text reconstruction of the power service customer service Q&A network, and determine the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power service customer service Q&A network in the content quality calibration link.
[0174] In step S400A, the computer system performs a content quality assessment process on the predicted response text through the content quality assessment network to obtain the content quality coefficient of the predicted response text.
[0175] In the embodiment of the present invention, the content quality assessment network is a model structure specifically used to evaluate the content quality of the predicted response text. For example, for the predicted response text "You can query the electricity bill on the official website of the power company, but the specific operation is not very clear" regarding the electricity bill query in the power service, the content quality assessment network will evaluate it from multiple dimensions.
[0176] This network can consider factors such as the accuracy, integrity, and relevance of information. Regarding accuracy, if there is a clear operation process for querying the electricity bill on the official website of the power company, but the predicted response text says that the operation is not very clear, this affects the accuracy. In terms of integrity, a complete electricity bill query answer may need to include steps such as logging in to the account, entering the query page, and possible authentication, but these are not mentioned in this predicted response text, so the integrity is insufficient. Regarding relevance, it is an answer to the electricity bill query question and has a certain relevance, but due to the lack of key information, the relevance does not fully meet the requirements.
[0177] The computer system quantifies these factors through various algorithms and mechanisms in the content quality assessment network. Assume that a weighted calculation method is used to obtain the content quality coefficient.
[0178] In step S400B, based on the quality assessment index and the content quality coefficient in the content quality calibration link, the target backpropagation optimization error in the content quality link is obtained.
[0179] The quality assessment index is an index obtained from the overall assessment of the predicted response text in the previous step, which reflects the comprehensive situation of the predicted response text in multiple aspects (including style, content quality, tendency level, etc.). Assume that the quality assessment index shows that although the predicted response text is generally within an acceptable range in terms of content quality, there is still room for improvement, and the content quality coefficient C = 0.52.
[0180] A computer system can determine the target backpropagation optimization error according to pre-set rules. For example, if the quality assessment index indicates that the content quality is at a medium level, the computer system may set an adjustment function f(C) to calculate the target backpropagation optimization error Et. This function may be determined based on a large amount of experimental data and experience. Suppose f(C)=k(1 - C), where k is a constant determined according to the network training situation and business requirements. For example, k = 2, then Et = 2×(1 - 0.52)=0.96. This means that in the content quality link, according to the current content quality coefficient and quality assessment index, the target backpropagation optimization error is 0.96, and this error will be used to guide the adjustment of network parameters to improve the content quality of the predicted response text.
[0181] In step S400C, the computer system obtains the reconstruction error when the power business customer service Q&A network performs text reconstruction, and determines the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power business customer service Q&A network in the content quality calibration link.
[0182] During the text reconstruction process of the power business customer service Q&A network, due to the influence of various factors such as network structure, parameter settings, and data, a reconstruction error will occur. For example, for the same question about electricity bill inquiry, the original ideal answer (based on training business texts and relevant knowledge) is "You can log in to the power company's official website, enter your account number and password, and after authentication, view the electricity bill details on the bill page", while the predicted response text generated by the network is "You can query the electricity bill on the power company's official website, and the specific operation is not very clear".
[0183] The way for the computer system to obtain the reconstruction error can be to compare the difference between the predicted response text and the original ideal answer. It can be quantified from aspects such as information missing and errors. Suppose a certain deduction is given for each missing key information point, and a larger deduction is given for each incorrect information, and then these deductions are aggregated to obtain the reconstruction error Er. For example, if three key information points of logging in to the account, authentication, and viewing electricity bill details are missing, and each information point is deducted 0.1 point, with a total deduction of 0.3 points, then the reconstruction error Er = 0.3.
[0184] Then add the reconstruction error Er and the target backpropagation optimization error Et to get the total backpropagation optimization error E = Er + Et. In the above example, E = 0.3 + 0.96 = 1.26. This total backpropagation optimization error will be used to adjust the network parameters of the power business customer service Q&A network in the content quality calibration link.
[0185] As an implementation manner, when the real-time calibration link that the power service customer service Q&A network is in is the tendency level calibration link, the evaluation score corresponding to the tendency level calibration link is the tendency level; at this time, step S400, generating a feedback optimization error based on the quality evaluation index, the predicted response text, and the response style constraint information, includes:
[0186] Step S400a: Arrange the tendency levels of the predicted response text through the tendency level network to obtain the tendency level of the predicted response text;
[0187] Step S400b: Obtain the target feedback optimization error in the tendency level link according to the quality evaluation index and the tendency level;
[0188] Step S400c: Obtain the reconstruction error when the power service customer service Q&A network performs text reconstruction, and determine the reconstruction error and the target feedback optimization error as the total feedback optimization error of the power service customer service Q&A network in the tendency level calibration link.
[0189] In step S400a, the computer system arranges the tendency levels of the predicted response text through the tendency level network to obtain the tendency level of the predicted response text.
[0190] In the context of power service customer service Q&A, the tendency level reflects the characteristics of the predicted response text in terms of emotional tendency, service orientation, etc. The tendency level network is a model structure specifically designed to judge the tendency level of text. For example, for the predicted response text when a user questions the high electricity bill in the power service, the computer system inputs it into the tendency level network for analysis.
[0191] This network can determine the tendency level based on various factors such as a predefined vocabulary, grammatical structure patterns, and semantic understanding. Assume that the tendency level is divided into three levels: positive, neutral, and negative. If the predicted response text is "Understand your concern about the high electricity bill. The power company has been working hard to optimize costs, and there are some energy-saving suggestions to help you reduce your electricity bill, such as using electrical appliances reasonably, etc.", the tendency level network will identify positive factors such as understanding the user and providing solutions, and thus determine the tendency level of this predicted response text as positive.
[0192] The technical means that computer systems may adopt in the propensity rating network include classification algorithms based on machine learning. For example, using the support vector machine (SVM) algorithm, a large number of power business-related texts with pre-labeled propensity ratings are used as training data to train a model that can classify propensity ratings based on text features. These text features can be lexical features (such as positive words such as "understanding", "optimization", "help", etc., and negative words such as "unreasonable" and "too high", etc.), sentence structure features (such as using active voice may be more inclined to positive expressions, etc.) and semantic features (obtained through word vectors and semantic analysis technology).
[0193] In step S400b, a target feedback optimization error in the tendency level link is obtained according to the quality evaluation index and the tendency level.
[0194] The quality assessment index is the result of the comprehensive evaluation of the predicted answer text in the previous step, which reflects the overall situation of the predicted answer text in multiple dimensions (such as style, content quality, tendency level, etc.). Assume that the quality assessment index shows that the predicted answer text is generally qualified in terms of tendency level, but there are some areas for improvement, and the tendency level of the current predicted answer text is positive, but not very positive.
[0195] The computer system can determine the target feedback optimization error according to pre-set rules. For example, a function g(q,p) can be defined, where q represents the quality assessment indicator and p represents the propensity level. This function can be set based on a large amount of experimental data and experience. Assuming that when the quality assessment indicator shows that the propensity level is close to qualified but needs to be improved, and the propensity level is positive, the function g(q,p) may be a function that calculates the error based on the gap from the ideal positive propensity.
[0196] For example, if a numerical value is used to represent the propensity level, 1 is positive, 0 is neutral, and -1 is negative, and the propensity level of the current predicted answer text is judged to be positive, but according to the quality assessment indicators and more detailed analysis, its performance in the propensity level is equivalent to 0.6 (indicating that it is not a completely positive state). Assuming the function g(q,p)=k(1-p), where k is a constant determined according to the network training situation and business needs, such as k=1.5, then the target feedback optimization error Et=1.5×(1-0.6)=0.6. This means that in the propensity level link, according to the current quality assessment indicators and propensity level, the target feedback optimization error is 0.6, and this error will be used to guide the adjustment of network parameters to further optimize the propensity level of the predicted answer text.
[0197] In step S400c, the computer system obtains the reconstruction error when reconstructing the text of the power service customer service Q&A network, and determines the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power service customer service Q&A network in the tendency level calibration link.
[0198] The way for the computer system to obtain the reconstruction error can be to compare the difference between the predicted response text and the ideal answer. It can be quantified from aspects such as the intensity of emotional expression and the integrity of the provided solutions. Suppose a certain deduction is given for insufficient intensity of emotional expression, and a greater deduction is given for insufficient integrity of the solution, and then these deductions are aggregated to obtain the reconstruction error Er. For example, if 0.1 point is deducted for insufficient intensity of emotional expression and 0.2 point is deducted for insufficient integrity of the solution, then the reconstruction error Er = 0.3.
[0199] Then add the reconstruction error Er and the target backpropagation optimization error Et to get the total backpropagation optimization error E = Er + Et. In the above example, E = 0.3 + 0.6 = 0.9. This total backpropagation optimization error will be used to adjust the network parameters of the power service customer service Q&A network in the tendency level calibration link.
[0200] As an implementation, step S460, obtaining the reconstruction error when reconstructing the text of the power service customer service Q&A network, includes:
[0201] Step S461: Obtain the predicted perturbation cleared when the power service customer service Q&A network performs vector reduction reconstruction operation, and the adjustment fluctuation parameter used when randomly perturbing the training service text;
[0202] Step S462: Based on the error between the predicted perturbation and the adjustment fluctuation parameter, obtain the reconstruction error when the power service customer service Q&A network reconstructs the text.
[0203] In step S461, the computer system needs to obtain the predicted perturbation cleared when the power service customer service Q&A network performs vector reduction reconstruction operation, and the adjustment fluctuation parameter used when randomly perturbing the training service text. Here, the vector reduction reconstruction operation refers to the operation performed in the previous step to obtain the perturbation-removed representation vector from the perturbation representation vector, with the purpose of removing the perturbation added to the training service text to restore a representation vector closer to the original service text, and then generating the predicted response text. Then, in step S462, the computer system obtains the reconstruction error when the power service customer service Q&A network reconstructs the text based on the error between the predicted perturbation and the adjustment fluctuation parameter. To understand this process, it can be assumed that the adjustment fluctuation parameter is a vector A = (a1, a2,..., a n ) and the predicted perturbation is a vector B = (b1, b2,..., b n), where n represents the dimension of the vector. A simple technical means to calculate the error between the two can be to calculate the Euclidean distance, and its formula is This E can be used as a measure of the reconstruction error. In this way, the computer system can quantify the error brought about by the perturbation and the operation of removing the perturbation during the text reconstruction process. The calculation of this reconstruction error is very important for evaluating the performance of the power business customer service Q&A network, because it reflects the accuracy of the network in processing perturbations and reconstructing text. If the reconstruction error is large, it indicates that there is a large deviation in the network's processing of perturbations and restoring the original text features, and the network parameters may need to be further adjusted; if the reconstruction error is small, it indicates that the network performs well in this regard. In this way, the computer system can more comprehensively evaluate the performance of the power business customer service Q&A network based on the reconstruction error, and in subsequent steps, combine this reconstruction error with other errors (such as the target backpropagation optimization error) to perform more accurate parameter adjustment on the power business customer service Q&A network, thereby improving the training quality of the network and the effect of text reconstruction.
[0204] In practical applications, the computer system reasonably sets the generation rules for adjusting the fluctuation parameters according to the specific characteristics and training requirements of the power business customer service Q&A network. For example, the value range and distribution method of the adjusted fluctuation parameters can be determined according to factors such as the length and semantic complexity of the training business text.
[0205] At the same time, the calculation of the predicted perturbation is not simply direct acquisition. In the vector reduction and reconstruction operation, the computer system can adopt a series of complex algorithms and model structures to gradually remove the perturbation. These algorithms and model structures can be based on some techniques in deep learning, such as the layer structure and activation function of the neural network. For example, in a text reconstruction network based on a multi-layer perceptron (MLP), the calculation of each layer of neurons will affect the removal of the perturbation, thereby indirectly affecting the calculation of the predicted perturbation. The computer system needs to accurately capture these influencing factors in order to correctly calculate the predicted perturbation and finally obtain an accurate reconstruction error.
[0206] In addition, the calculation method of the reconstruction error is not limited to the Euclidean distance. According to different application scenarios and requirements, the computer system can also adopt other distance metrics or error calculation functions. For example, in some cases, the Manhattan distance or cosine similarity To measure the difference between the predicted perturbation and the adjusted fluctuation parameters. These different calculation methods have their own advantages and disadvantages, and the computer system needs to make a choice according to the specific situation. For example, when the dimension of the vector is high, the Euclidean distance can be affected by the curse of dimensionality, while the Manhattan distance can be more stable in this case; the cosine similarity focuses more on measuring the similarity of the vector directions and may be more applicable in some scenarios sensitive to the vector directions.
[0207] As an implementation manner, in step S300, removing the perturbation from the perturbation representation vector includes:
[0208] Step S300A: Extract the text representation vector corresponding to the response style constraint information of the training service text, and load the text representation vector into the text reconstruction network as the key vector;
[0209] Step S300B: Use the perturbation representation vector as the query vector for removing the perturbation by the text reconstruction network;
[0210] Step S300C: Query through the text reconstruction network according to the query vector under the key vector, and obtain the perturbation-removed representation vector according to the value vector.
[0211] In step S300A, the computer system extracts the text representation vector corresponding to the response style constraint information of the training service text, and loads it into the text reconstruction network as the key vector.
[0212] In step S300B, the computer system uses the perturbation representation vector as the query vector for removing the perturbation by the text reconstruction network. The perturbation representation vector is a vector obtained by randomly perturbing the training service text before.
[0213] Finally, in step S300C, the computer system queries through the text reconstruction network according to the query vector under the key vector, and obtains the perturbation-removed representation vector. In the text reconstruction network based on the attention mechanism, an interaction operation will be performed between the query vector (perturbation representation vector) and the key vector (text representation vector corresponding to the response style constraint information). Specifically, the computer system can determine the attention degree to different parts when generating the perturbation-removed representation vector by calculating the similarity score between the query vector and the key vector. A feasible technical means for calculating the similarity score is the dot product operation. Suppose the query vector is Q=(q1,q2,…,q n ), and the key vector is K=(k1,k2,…,k n ), then the similarity score S between them can be calculated by the formula It is calculated. Based on this similarity score, the text reconstruction network assigns different weights to different parts, and then obtains and combines the corresponding elements from a predefined value vector (this value vector contains the information elements required to construct the perturbation-removed representation vector), thereby obtaining the perturbation-removed representation vector. For example, if the similarity score between the query vector and the key vector is high in a certain dimension, then when constructing the perturbation-removed representation vector, more attention will be paid to the elements related to this dimension in the value vector.
[0214] In step S300C, the query-key-value operation based on the attention mechanism is the core part of the text reconstruction network. In addition to using the dot product operation to calculate the similarity score, other functions can also be used to measure the relationship between the query vector and the key vector, such as the formula S = tanh(W Q Q + W K K) in the additive attention mechanism, where W Q and W K are learnable weight matrices. After assigning weights according to the similarity score, the process of obtaining elements from the value vector to construct the perturbation-removed representation vector is not a simple direct selection either. The computer system can further process the obtained value vector elements according to the architecture and training objectives of the text reconstruction network, such as performing a non-linear transformation (through activation functions such as ReLU, Sigmoid, etc.) or performing multi-layer combination operations. For example, in a multi-layer text reconstruction network, the elements obtained from the value vector can first be processed by a ReLU activation function, and then summed or concatenated with other processed elements to finally obtain the perturbation-removed representation vector.
[0215] An embodiment of the present invention provides a computer system, as Figure 2 shown, the computer system 100 includes: a processor 101 and a memory 103. Among them, the processor 101 and the memory 103 are connected, such as connected through a bus 102. Optionally, the computer system 100 may further include a transceiver 104. It should be noted that in practical applications, the transceiver 104 is not limited to one, and the structure of the computer system 100 does not constitute a limitation to the embodiments of the present invention.
[0216] An embodiment of the present invention provides a computer system. The computer system in the embodiment of the present invention includes: one or more processors; a memory; one or more computer programs, where one or more computer programs are stored in the memory and configured to be executed by one or more processors. When the one or more programs are executed by the processor, the above method is implemented.
Claims
1. A training method for an electric power service customer service Q&A model, characterized in that, Including: Obtain training binary tuple data of the power service customer service Q&A network, where the training binary tuple data includes training service texts and response style constraint information of the training service texts; Perform random perturbation on the training service texts to obtain perturbation representation vectors of the training service texts, and obtain the number of perturbation removals relied on during the iterative calibration process of the power service customer service Q&A network; wherein, the iterative calibration process of the power service customer service Q&A network includes one or more calibration links, and the number of perturbation removals passed by different calibration links is different; Remove perturbations from the perturbation representation vectors according to the number of perturbation removals to obtain perturbation-removed representation vectors, and generate predicted response texts of the training service texts based on the perturbation-removed representation vectors; Obtain quality evaluation indicators of the predicted response texts, generate backpropagation optimization errors based on the quality evaluation indicators, the predicted response texts, and the response style constraint information, and adjust network parameters of the power service customer service Q&A network based on the backpropagation optimization errors to obtain a trained power service customer service Q&A network.
2. The method according to claim 1, characterized in that, The performing random perturbation on the training service texts to obtain perturbation representation vectors of the training service texts includes: Perform text embedding operation on the training service texts to obtain training service text representation vectors of the training service texts; Obtain adjustment fluctuation parameters, and perform random perturbation on the training service text representation vectors through the adjustment fluctuation parameters to obtain initial perturbation representation vectors of the training service texts; wherein, the adjustment fluctuation parameters are random masks, the random masks are in the same vector domain as the initial perturbation representation vectors, and the vector dimensions of the random masks and the initial perturbation representation vectors in the corresponding vector domain are the same; Perform vector encoding on the initial perturbation representation vectors to obtain perturbation representation vectors of the training service texts; The obtaining the number of perturbation removals relied on during the iterative calibration process of the power service customer service Q&A network includes: Obtain the current real-time calibration link of the power service customer service Q&A network and the total number of removals a when removing perturbations from the perturbation representation vectors; Based on the total number of removals a, determine the corresponding real-time number range during the process of adjusting network parameters of the power service customer service Q&A network in the real-time calibration link; the number ranges corresponding to the processes of adjusting network parameters of the power service customer service Q&A network in different calibration links are different; wherein, a is a positive integer greater than 0; Determine an arbitrary value in the real-time number range to obtain a selected number, and determine the selected number as the number of perturbation removals relied on during the iterative calibration process of the power service customer service Q&A network in the real-time calibration link.
3. The method according to claim 1, characterized in that, The real-time calibration link is one of the following three links: style matching calibration link, content quality calibration link, and tendency level calibration link; Among them, during the process of adjusting the network parameters of the power service customer service Q&A network in the style matching and calibration step, the corresponding number range is a numerical range greater than or equal to e and less than or equal to f, where both e and f are positive integers not greater than a; During the process of adjusting the network parameters of the power service customer service Q&A network in the content quality calibration step, the corresponding number range is a numerical range greater than f and less than or equal to g, where g is a positive integer not greater than a; During the process of adjusting the network parameters of the power service customer service Q&A network in the tendency level calibration step, the corresponding number range is a numerical range greater than f and less than or equal to h, where h is a positive integer not greater than a.
4. The method according to claim 1, wherein The removing the perturbation from the perturbation representation vector according to the number of perturbation removals to obtain a perturbation-removed representation vector includes: Performing perturbation removal on the perturbation representation vector for a corresponding number of times based on the number of perturbation removals to obtain an initial perturbation-removed representation vector; According to the random perturbation process of the training service text, performing a vector reduction and reconstruction operation on the initial perturbation-removed representation vector to obtain the perturbation-removed representation vector corresponding to the perturbation representation vector; The obtaining the quality evaluation index of the predicted response text includes: Obtaining a real-time evaluation score for content quality evaluation in the real-time calibration step; wherein, the evaluation score includes one of the following dimensions: style matching score, content quality score, and tendency level, and one calibration step corresponds to one evaluation score; Obtaining the score value of the predicted response text under the real-time evaluation score, and obtaining the critical score value corresponding to the real-time evaluation score; Comparing the score value under the real-time evaluation score with the critical score value, and obtaining the quality evaluation index of the predicted response text based on the comparison result; Among them, the obtaining the critical score value corresponding to the real-time evaluation score includes: Obtaining a training database, where the training database includes one or more binary data templates, and each binary data template includes a service text and style constraint information of the service text; Extracting the text representation vector of each service text, and grouping one or more service texts according to the text representation vector of the service text to obtain s classifications; s is a positive integer greater than 1; Selecting x service texts in each classification, and performing text reconstruction on the selected service texts and the corresponding style constraint information through the power service customer service Q&A network to obtain x reconstructed texts for each classification; x is a positive integer greater than 1; Determining the score value of each reconstructed text under the real-time evaluation score, and performing statistics on the obtained score values to obtain a statistical value as the critical score value of the real-time evaluation score.
5. The method according to claim 4, wherein The training binary data is any one binary data template in the training database; the obtaining the score value of the predicted response text under the real-time evaluation score includes: Determine the target classification where the training service text is located among the s classifications, and determine the selected service text in the target classification, as well as the target score value of the reconstructed text corresponding to the selected service text under the real-time evaluation score; Generate the score value of the predicted response text corresponding to the training service text under the real-time evaluation score based on the target score value; The quality evaluation index of the predicted response text obtained based on the comparison result includes: If the comparison result indicates that the score value is not less than the critical score value, determine the corresponding quality evaluation index as the first quality evaluation index; If the comparison result indicates that the score value is less than the critical score value, determine the corresponding quality evaluation index as the second quality evaluation index; Wherein, the first quality evaluation index and the second quality evaluation index are different.
6. The method according to claim 1, wherein When the real-time calibration link where the power service customer service Q&A network is located is the style matching calibration link, the evaluation score corresponding to the style matching calibration link is the style matching score; The generation of the backpropagation optimization error based on the quality evaluation index, the predicted response text, and the response style constraint information includes: Based on the predicted response text and the response style constraint information, determine the initial backpropagation optimization error in the style matching calibration link; According to the quality evaluation index and the initial backpropagation optimization error in the style matching calibration link, obtain the target backpropagation optimization error in the style matching stage; Obtain the reconstruction error when the power service customer service Q&A network performs text reconstruction, and use the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power service customer service Q&A network in the style matching calibration link.
7. The method according to claim 6, characterized in that, The determination of the initial backpropagation optimization error in the style matching calibration link based on the predicted response text and the response style constraint information includes: Extract the predicted response text representation vector of the predicted response text through a dual-channel network, and extract the text representation vector corresponding to the response style constraint information of the training service text; the text representation vectors extracted by the dual-channel network and the text representation vector are in the same vector domain; Determine the vector distance between the text representation vector and the text representation vector, and obtain the initial backpropagation optimization error in the style matching calibration link based on the vector distance.
8. The method according to claim 1, characterized in that, When the real-time calibration link where the power service customer service Q&A network is located is the content quality calibration link, the evaluation score corresponding to the content quality calibration link is the content quality score; the generation of the backpropagation optimization error based on the quality evaluation index, the predicted response text, and the response style constraint information includes: Perform content quality evaluation processing on the predicted response text through a content quality evaluation network to obtain the content quality coefficient of the predicted response text; According to the quality evaluation index and the content quality coefficient in the content quality calibration link, obtain the target backpropagation optimization error in the content quality link; Obtain the reconstruction error when reconstructing the text of the power service customer service Q&A network, and determine the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power service customer service Q&A network in the content quality calibration link; Alternatively, when the real-time calibration link in which the power service customer service Q&A network is located is the tendency level calibration link, the evaluation score corresponding to the tendency level calibration link is the tendency level; the generation of the backpropagation optimization error based on the quality evaluation index, the predicted response text, and the response style constraint information includes: Arrange the tendency levels of the predicted response text through a tendency level network to obtain the tendency level of the predicted response text; Obtain the target backpropagation optimization error in the tendency level link according to the quality evaluation index and the tendency level; Obtain the reconstruction error when the power service customer service Q&A network reconstructs the text, and determine the reconstruction error and the target backpropagation optimization error as the total backpropagation optimization error of the power service customer service Q&A network in the tendency level calibration link; The obtaining of the reconstruction error when the power service customer service Q&A network reconstructs the text includes: Obtain the predicted perturbations cleared during the vector reduction and reconstruction operation of the power service customer service Q&A network, and the adjustment fluctuation parameters used when randomly perturbing the training service text; Based on the error between the predicted perturbation and the adjustment fluctuation parameter, obtain the reconstruction error when the power service customer service Q&A network reconstructs the text.
9. The method according to claim 1, wherein The perturbation removal of the perturbation representation vector includes: Extract the text representation vector corresponding to the response style constraint information of the training service text, and load the text representation vector into the text reconstruction network as the key vector; Use the perturbation representation vector as the query vector for perturbation removal by the text reconstruction network; Through the text reconstruction network, query according to the query vector under the key vector, and obtain the perturbation removal representation vector according to the value vector.
10. A computer system, characterized in that, Includes: One or more processors; A memory; One or more computer programs; Wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Experimental method and system for multi-scene validity verification of large power language model
CN121329204A