Method and device for applying large model to intelligent customer service of electricity market

By applying a large language model based on deep learning in intelligent customer service, the problem of traditional intelligent customer service relying on predefined knowledge bases is solved, and more accurate and personalized user interaction response is achieved, improving the intelligent quality of power customer service.

CN119938826AInactive Publication Date: 2025-05-06INFORMATION & COMMUNICATION BRANCH STATE GRID JIBEI ELECTRIC POWER CO LTD +1

Patent Information

Application Number
CN202411862529.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional intelligent customer service relies on a predefined knowledge base for answers, resulting in the inability to accurately identify user demands, unable to provide personalized services, and difficult to meet user needs in the case of large traffic and heavy business.

Method used

A large language model trained based on deep learning transformer architecture is used as a preset session interaction model. By receiving user's power topic data, including text, voice and image data, the encoder extracts key information and understands the context, and the decoder generates interactive response data.

Benefits of technology

It achieves better interactive response effects, makes the answers more in line with user intentions, provides better intelligent customer service solutions, and improves the intelligent quality of power customer service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938826A_ABST
    Figure CN119938826A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for applying a large model to power market intelligent customer service, and relates to the technical field of power markets. According to the main technical scheme, for received power topic data, not limited to text data, voice data and image data, provided by a user, the power topic data is processed by utilizing a pre-trained preset session interaction model; the preset session interaction model is a large language model trained based on a converter architecture in deep learning, so that the advantages of the large language model are combined, and an encoder in the converter architecture is utilized to analyze the power topic data as text data, which at least comprises key information extraction and context understanding; and then, a decoder is utilized to gradually generate required output data as interactive response data for the power topic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of power market, and in particular to an application method and device for intelligent customer service in power market based on a large model. Background Art

[0002] With the continuous development of the power market and the increasing diversification of user needs, the traditional customer service model has been unable to meet the needs of users. Users not only need to obtain power information quickly and accurately, but also need personalized services and solutions. Therefore, the power industry has begun to actively explore the application of intelligent customer service to improve service quality and efficiency.

[0003] At present, traditional intelligent customer service robots usually rely on predefined knowledge bases to answer questions, but the update and expansion of the knowledge base may not be timely and comprehensive enough, resulting in the inability to accurately identify user demands. In response to the current large volume of calls and heavy seat workload, it is urgent to provide adaptive upgrades to intelligent customer service and improve the intelligent quality of power customer service. Summary of the invention

[0004] In view of this, the present application provides an application method and device based on a big model in the intelligent customer service of the power market. The main purpose is to use the big model to achieve interactive response, and use key information and context to understand such richer data information, so that the interactive effect of the answer is better and more in line with the user's intention, thereby providing a better intelligent customer service solution and improving the intelligent quality of power customer service.

[0005] In order to achieve the above objectives, this application mainly provides the following technical solutions:

[0006] In a first aspect, the present application provides an application method based on a large model in intelligent customer service in the power market, the method comprising:

[0007] Receive power topic data provided by a user, wherein the power topic data is data generated by a conversation with an intelligent customer service, and the power topic data includes at least text data, voice data, and image data;

[0008] The power topic data is processed by using a preset conversation interaction model, and interactive response data corresponding to the power topic data is output, wherein the preset conversation interaction model is a large language model trained based on a transformer architecture in deep learning; the transformer architecture includes an encoder and a decoder, the encoder is used to analyze the power topic data as text data, wherein the analysis of the text data at least includes extracting key information and understanding the context; the decoder uses the encoder to analyze the text data to gradually generate the required output data.

[0009] The second aspect of the present application provides an application device for intelligent customer service in the power market based on a large model, the device comprising:

[0010] A receiving unit, configured to receive power topic data provided by a user, wherein the power topic data is data for generating a conversation with an intelligent customer service, and the power topic data includes at least text data, voice data, and image data;

[0011] A processing unit is used to process the power topic data using a preset conversation interaction model and output interactive response data corresponding to the power topic data, wherein the preset conversation interaction model is a large language model trained based on a transformer architecture in deep learning; the transformer architecture includes an encoder and a decoder, wherein the encoder is used to analyze the power topic data as text data, wherein the analysis of the text data at least includes extracting key information and understanding the context; and the decoder uses the encoder to analyze the text data to gradually generate the required output data.

[0012] A third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the application method based on a large model in intelligent customer service in the power market as described above is implemented.

[0013] The fourth aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the application method based on the big model in the intelligent customer service of the electricity market as described above.

[0014] By means of the above technical solution, the technical solution provided by this application has at least the following advantages:

[0015] The present application provides an application method and device based on a large model in intelligent customer service in the power market. For the power topic data received from the user, which is not limited to text data, voice data and image data, the present application uses a pre-trained preset conversation interaction model to process the power topic data. Since the preset conversation interaction model is a large language model trained based on the transformer architecture in deep learning, the advantages of the large language model are combined, and the encoder in the transformer architecture is used to analyze the power topic data as text data, at least including extracting key information and context understanding, and then using a decoder to gradually generate the required output data as interactive response data to the power topic data.

[0016] Compared with the existing technology, traditional intelligent customer service relies on a predefined knowledge base to answer questions, resulting in technical problems of low quality. This application does not rely on such a predefined knowledge base, so that the answers will not be subject to such restrictions. Instead, it uses a large model to achieve interactive responses, and uses key information and context to understand such richer data information, so that the interactive effect of the answers is better and more in line with user intentions, thereby providing a better intelligent customer service solution and improving the intelligent quality of power customer service.

[0017] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0019] Figure 1 A flow chart of an application method based on a large model in intelligent customer service in the power market provided in an embodiment of the present application;

[0020] Figure 2 A flow chart of another application method based on a large model in intelligent customer service in the power market provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of a Transformer architecture exemplified in an embodiment of the present application;

[0022] Figure 4 This is a schematic diagram of an intelligent customer service example of the present application;

[0023] Figure 5 A block diagram of a composition of an application device based on a large model for intelligent customer service in the power market provided in an embodiment of the present application;

[0024] Figure 6 A block diagram of the composition of another application device for intelligent customer service in the power market based on a large model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0026] The present application embodiment provides an application method based on a large model in the power market intelligent customer service, such as Figure 1 As shown, the present application embodiment provides the following specific steps:

[0027] 101. Receive power topic data provided by a user, where the power topic data is data generated by a conversation with an intelligent customer service, and the power topic data includes at least text data, voice data, and image data.

[0028] In an embodiment of the present application, the user may be an individual user or a corporate user, and power topic data provided by the user may be received through a human-computer interaction interface. For example, it is possible but not convenient to open a conversation box, and the user may enter the power topic data in text form, or in voice form or image form. The power topic data may be, but is not limited to, asking questions, interactive chatting, etc.

[0029] The power topic data may refer to, but is not limited to, data information related to power business in the power market. Such data information has at least the following characteristics:

[0030] (1) Diversification of data sources

[0031] The data in the power market comes from all aspects of power production and consumption, such as power generation, transmission, transformation, distribution, power consumption and dispatching. Specifically, these data may include: power grid operation data: such as data generated by energy management systems, distribution network management systems, wide-area measurement management systems, etc.; equipment monitoring data: such as data generated by production management systems, power grid dispatching management systems, fault management systems, image monitoring systems, etc.; power enterprise marketing data: such as transaction electricity prices, electricity sales, electricity customers, etc., these data are mainly included in marketing business systems, customer service systems, electricity energy metering systems, electricity consumption information collection systems, etc.

[0032] (2) Complex and diverse data types

[0033] The data types in the power market are complex and diverse, including traditional structured data such as numbers and symbols, as well as a large amount of semi-structured and unstructured data such as text, images, audio, etc. For example, voice data in the customer service center, video data and image data in the equipment online monitoring system are all unstructured data. This diversity of data types increases the difficulty of data processing, but also provides a richer source of information for data analysis.

[0034] (3) Huge data size

[0035] With the popularization of smart grids and the application of Internet of Things technology, the amount of data generated in the power market has exploded. Smart grids deploy a large number of smart meters and other monitoring devices, which continuously generate a large amount of real-time data. At the same time, the production and application of energy also involve multiple systems, such as heating systems, cooling systems, gas systems, transportation systems, etc. The data interaction between these systems also further increases the scale of data.

[0036] (4) High data value

[0037] The data in the power market contains huge value. By analyzing and mining these data, power companies can understand the supply and demand of the power market, the power consumption behavior of users, the operating status of equipment and other information, so as to formulate more scientific operation strategies, improve energy efficiency and optimize the allocation of power resources. In addition, these data can also provide support for the decision-making of government departments, such as the prediction and analysis of social and economic indicators such as industrial transfer, industrial development and housing vacancy rate.

[0038] Based on the above characteristics of data information in the power market, the power topic data required by users can be constructed.

[0039] In an embodiment of the present application, after the user inputs the power topic data, the power topic data will be treated as text data by the machine, to be further identified and processed by the machine, so as to finally obtain response data that is interactively compatible with the text data and output it to the user, thereby realizing the intelligent customer service provided to the user.

[0040] 102. Use a preset conversation interaction model to process the power topic data and output interactive response data corresponding to the power topic data. The preset conversation interaction model is a large language model trained based on a transformer architecture in deep learning. The transformer architecture includes an encoder and a decoder. The encoder is used to analyze the power topic data as text data, wherein the analysis of the text data at least includes extracting key information and understanding the context. The decoder uses the encoder to analyze the text data to gradually generate the required output data.

[0041] Large Language Model (LLM), also known as big model, originates from the foundation and integrated development of natural language processing (NLP) and deep learning. Different from traditional pre-training models, LLM has more model parameters and training data. By adopting unsupervised learning pre-training and autoregressive generation, LLM can learn from massive data corpus, thereby significantly improving its ability to solve complex language environments and context perception.

[0042] The basic architecture of LLM used in the embodiment of the present application is the Transformer architecture in deep learning. The Transformer structure includes an encoder and a decoder, which work together to process and generate sequence data. The functions of the encoder and the decoder are explained below, including:

[0043] The main task of the encoder is to map the input sequence to a hidden representation of a fixed length. It analyzes the input text, understands the meaning of each element in the text, and discovers the hidden relationships between them. Specifically, the encoder's functions include:

[0044] (1) Information extraction: The encoder processes the input sequence through a multi-layer Transformer block, each of which contains a self-attention mechanism and a feedforward neural network. These components work together to extract key information from the input sequence and convert it into a vector representation.

[0045] (2) Contextual understanding: The self-attention mechanism allows the encoder to focus on other words in the input sequence when processing each word. This mechanism can capture dependencies in sequence data and help the encoder understand the contextual information of the text.

[0046] (3) Fixed-length representation: The encoder converts the input sequence into a fixed-length hidden representation that contains all the key information and contextual relationships of the input sequence. This fixed-length representation can be used as input for subsequent decoding processes or for other tasks (such as classification, regression, etc.).

[0047] The main task of the decoder is to generate the target sequence based on the output of the encoder. It uses the deep insights provided by the encoder to gradually generate the desired output. Specifically, the decoder's role includes:

[0048] (1) Sequence generation: The decoder processes the encoder output through multiple layers of Transformer blocks and gradually generates words of the target sequence. During the generation process, the decoder uses the generated words and the encoder output as input to generate the next word through self-attention and encoder-decoder attention mechanisms.

[0049] (2) Maintaining consistency: When generating the target sequence, the decoder needs to ensure that the generated text is consistent and coherent with the original text. This is mainly achieved through the encoder-decoder attention mechanism, which allows the decoder to refer to the encoded input to ensure that the generated output is consistent with the original text.

[0050] (3) Masking: During the decoding process, the decoder needs to avoid using future information. Therefore, it uses masking technology to process the input sequence to ensure that it only focuses on the generated words and the output of the encoder. This processing method helps maintain the causal order of sequence generation and ensures the coherence and consistency of the model when generating the target sequence.

[0051] In summary, the encoder and decoder play a vital role in the large model structure. The encoder is responsible for extracting the key information and contextual relationships of the input sequence and converting it into a hidden representation of a fixed length. The decoder uses the output of the encoder and the generated words to gradually generate the target sequence. Together, they constitute the core components of the Transformer architecture, providing powerful support for natural language processing and other sequence-to-sequence tasks.

[0052] As described above, the embodiment of the present application provides an application method and device based on a big model in the intelligent customer service of the electricity market. For the electricity topic data received from the user, it is not limited to text data, voice data and image data. The embodiment of the present application uses a pre-trained preset conversation interaction model to process the electricity topic data. Since the preset conversation interaction model is a large language model trained based on the transformer architecture in deep learning, the advantages of the large language model are combined, and the encoder in the transformer architecture is used to analyze the electricity topic data as text data, at least including extracting key information and context understanding, and then using the decoder to gradually generate the required output data as interactive response data to the electricity topic data.

[0053] Compared with the prior art, traditional intelligent customer service relies on a predefined knowledge base to answer questions, resulting in technical problems of low quality. The embodiment of the present application does not rely on such a predefined knowledge base, so that the answers will not be subject to such restrictions. Instead, a large model is used to achieve interactive responses, and key information and context understanding are used to enrich data information, so that the interactive effect of the answers is better and more in line with user intentions, thereby providing a better intelligent customer service solution and improving the intelligent quality of power customer service.

[0054] To explain in more detail, the present application embodiment provides another application based on a large model in the intelligent customer service of the power market, such as Figure 2 As shown, the following specific steps are provided:

[0055] 201. Receive power topic data provided by a user, where the power topic data is data generated by a conversation with an intelligent customer service, and the power topic data includes at least text data, voice data, and image data.

[0056] 202. Convert the power topic data into text data to obtain a vector corresponding to the text data.

[0057] 203. Add positional encoding to the vector to construct the input sequence corresponding to the vector.

[0058] In the Transformer architecture, the encoder’s input sequence is usually composed of a series of embedding vectors that represent words or characters in the input text.

[0059] Exemplarily, specifically, the process of generating the input sequence generally includes the following steps (1)-(4):

[0060] (1) Text preprocessing: First, the original text is preprocessed, including word segmentation, stop word removal, stemming (for languages ​​such as English), etc., to obtain a series of words or characters as input.

[0061] (2) Word Embedding: Then, each word or character is mapped to a fixed-length vector space, which is usually obtained through unsupervised learning algorithms (such as Word2Vec, GloVe, etc.) or pre-trained language models (such as BERT, GPT, etc.). This process is called word embedding, which can convert words or characters into low-dimensional dense vector representations. These vectors can capture the semantic relationship between words or characters.

[0062] (3) Positional Encoding: Since the Transformer architecture itself does not contain a loop or convolution structure, it cannot directly perceive the position information of words in the input sequence. To solve this problem, positional encoding (or position embedding) is usually added to the input embedding vector to provide additional information about the position of the word in the sequence. Positional encoding can be generated by sine and cosine functions or learned.

[0063] (4) Constructing the input sequence: Finally, the vectors containing word embeddings and positional encodings are arranged in the order of the input text to construct the encoder input sequence. This input sequence will serve as the input of the Transformer encoder for subsequent self-attention mechanism and feedforward neural network layer processing.

[0064] It should be noted that different Transformer models and tasks may make some adjustments or extensions to the construction process of the input sequence. For example, when dealing with multilingual tasks, it may be necessary to add language embedding to represent information in different languages; when dealing with text generation tasks, it may be necessary to use special start and end tags to indicate the beginning and end of the generated sequence.

[0065] 204. Determine N encoding layers of the encoder, each encoding layer includes at least a self-attention layer and a feedforward neural network layer; the self-attention layer is used to calculate the attention weights at different positions based on the input sequence, and the feedforward neural network layer is used to perform further nonlinear transformation and mapping on the output of each self-attention layer to learn complex feature representations, wherein N is a positive integer.

[0066] The encoder in the Transformer architecture is a multi-layer structure. In the Transformer architecture, the encoder part is composed of N encoder layers stacked together, each of which contains multiple sublayers, which usually include multi-head self-attention layers, feedforward neural network layers, as well as normalization layers and residual connections.

[0067] Specifically, the multi-head self-attention layer in each encoder layer allows the model to focus on other words in the input sequence when processing each word, thereby capturing dependencies in the sequence data. The feedforward neural network layer further processes and transforms the representation of each position. The normalization layer and residual connection are used to stabilize the training process and improve the generalization ability of the model.

[0068] By stacking multiple encoder layers, the Transformer architecture is able to gradually extract deep features from the input sequence and generate a hidden representation containing rich information. This hidden representation is then passed to the decoder part to generate the target sequence.

[0069] Therefore, the encoder in the Transformer architecture is a multi-layer structure, which enables the model to capture the complex features and dependencies in the input sequence, thereby achieving excellent performance in various natural language processing tasks.

[0070] 205. The input sequence is processed step by step using N encoding layers to obtain deep features corresponding to the input sequence, which are used to generate corresponding hidden representations.

[0071] In an embodiment of the present application, in each encoding layer, a self-attention layer is used to process an input sequence to obtain attention weights at different positions on the input sequence; based on inputting the attention weights at different positions on the input sequence into a feedforward neural network layer, features corresponding to the input sequence are output.

[0072] In the Transformer architecture, the encoder part is composed of N encoder layers stacked together. Each encoder layer contains a self-attention layer and a feedforward neural network layer. These two components play a vital role in the model.

[0073] The Self-Attention Layer is one of the core components of the Transformer encoder. Its main function is to allow the model to consider other positions in the sequence when processing each position of the input sequence, thereby capturing long-distance dependencies between elements. The Self-Attention Layer generates new vector representations by calculating the attention scores between the query, key, and value matrices. These matrices are usually obtained from the input embedding vector through linear transformation.

[0074] Specifically, the self-attention mechanism allows the model to calculate the attention weight based on the entire input sequence when generating the output of a certain position, thereby determining which position's information is most important to the current position. This mechanism enables the model to capture a variety of relevant information at different positions in the input sequence and dynamically adjust its contextual representation.

[0075] The Feed-Forward Neural Network Layer is a simple fully connected feed-forward network that is used to perform further nonlinear transformation and mapping on the output of each self-attention layer. It usually consists of two linear layers and an activation function (such as ReLU) to learn complex feature representations.

[0076] The main function of the feedforward neural network layer is to further process and process the output of the self-attention layer to introduce more nonlinear capabilities and improve the expressiveness of the model. Through nonlinear transformation, the feedforward neural network layer can fuse different feature information and generate richer vector representations. These vector representations are then passed to the next encoder layer or decoder part for subsequent task processing.

[0077] As shown above, in the Transformer architecture, each encoder layer in the encoder part contains a self-attention layer and a feedforward neural network layer. The self-attention layer allows the model to capture long-distance dependencies in the input sequence and dynamically adjust its contextual representation. The feedforward neural network layer further nonlinearly transforms and maps the output of the self-attention layer to introduce more nonlinear capabilities and improve the expressiveness of the model. These two components work together to enable the Transformer architecture to achieve excellent performance in various natural language processing tasks.

[0078] 206. Pass the hidden representation to the decoder for generating the target sequence.

[0079] 207. Output interactive response data corresponding to the power topic data according to the target sequence.

[0080] The embodiment of the present application utilizes 204-205 to implement processing and analysis of text data corresponding to the power topic data, in fact, in order to obtain key information and context understanding so as to enrich the data information, so that the subsequent encoder outputs response data, the interactive effect achieved is better and in line with the user's intention.

[0081] For the above steps 201-207, the present embodiment also provides a Transformer architecture diagram, such as Figure 3 As shown, and an exemplary explanation is given as follows:

[0082] The Transformer model consists of two parts: the encoder and the decoder. The encoder is the component that processes the input sequence, extracts key features through the self-attention mechanism and the feedforward network, and generates a continuous representation containing the semantic information of the source sequence, providing the necessary context information for the decoder. Figure 3 The Transformer model architecture is shown, with the encoder on the left and the decoder on the right.

[0083] The encoder first uses position encoding to extract sequence position information. The input sequence consisting of one-hot encoding is converted into d model dimensional embedding vector, and transmits the position information of each sequence element through position encoding. The length of the position encoding vector is d model , when the current word has position pos in the sequence, the following function is introduced for each element i of the position vector:

[0084]

[0085] Each dimension of the position encoding corresponds to a sinusoidal signal. The wavelength ranges from 2π to 1000·2π in a geometric series. For any fixed offset k, PEpos+k can be expressed as a linear function of PEpos, which makes it easy for the model to learn to participate by relative position.

[0086] After position encoding, the remaining encoder consists of N = 6 identical blocks stacked together, each of which contains two sub-blocks: a multi-head self-attention mechanism and a feedforward neural network. The formula for the multi-head self-attention mechanism is as follows:

[0087]

[0088] In the formula, They are d kThe query, key, and value corresponding to the original sequence x of dimension 2 as the input matrix. The self-attention mechanism enables the model to assign different attention weights to each element in the sequence when processing the input sequence, thereby capturing the relationship between different elements.

[0089] A simple, fully connected feedforward network is defined as follows:

[0090] FFN(x)=ReLU(W1x+b1)W2+b2 Formula (3);

[0091] The network is applied to each sequence element x separately, using the same set of parameters (W1, W2, b1, b2) at different positions, but different parameters across different encoder blocks.

[0092] Each sub-block is connected through a residual connection and normalizes the sum of the connection and its own output and the residual connection. The input of each sub-block is as follows: LayerNorm(x+SubLayer(x)) Formula (4).

[0093] In summary, the basic architecture of LLM is the Transformer architecture in deep learning. The Transformer structure includes an encoder and a decoder, which work together to process and generate sequence data. The encoder consists of multiple layers, each of which contains a self-attention mechanism and a feedforward neural network, responsible for converting the input text sequence into a series of high-dimensional feature representations. These feature representations capture the rich semantic information and structural relationships of the input data. Similar to the encoder, the decoder is also composed of a multi-layer structure, but is designed to generate output sequences. Each layer of the decoder receives the output of the encoder and the previously generated sequence part as input, and generates the next element of the sequence. Through the self-attention mechanism, the model can capture long-distance dependencies within the sequence. The self-attention mechanism is the core component of the Transformer architecture. It enables the model to assign different attention weights to each element in the sequence when processing the input sequence, thereby capturing the relationship between different elements, as shown in formula (1).

[0094] The calculation of the self-attention mechanism includes three steps: first, the input sequence is mapped to the key, query, and value vector space through linear transformation; second, the similarity between the query and the key is calculated through the dot product operation and normalized into a probability distribution to represent the attention weight; finally, the value vector is weighted and summed with the attention weight to obtain the final output representation. After each self-attention layer, the Transformer contains a feedforward neural network layer. Feedforward neural networks usually consist of two linear transformations and an activation function. The feedforward neural network layer is applied independently to each position in the sequence to further process and transform the data representation. Multi-head attention is an improvement on the self-attention mechanism. This mechanism can learn the characteristics of the input data from different angles by performing multiple self-attention calculations in parallel (each time is called a "head"), splicing the results, and then obtaining the final output through linear transformation. In addition, LLM contains fully connected layers, residual connections, and layer normalization to further process and optimize the data representation, enhance the learning ability and stability of the model, and avoid the gradient disappearance and gradient explosion problems. The model usually learns a general language representation on a large amount of text data through pre-training, and then adapts to specific task requirements through fine-tuning, and then uses the autoregressive property to enable LLM to continuously output the next element in the sequence.

[0095] Furthermore, for the intelligent customer service implemented as described in 201-207 above, the embodiment of the present application also provides the following application scenarios.

[0096] (1) Chatbots: Using the Transformer architecture, chatbots can better understand the context of user queries and provide more accurate answers. Through the self-attention mechanism, the chatbot can capture the key information in the user query and generate appropriate responses based on this information.

[0097] (2) Sentiment analysis: The large model of the Transformer architecture can be used to analyze customer emotions and help companies better understand customer needs. By analyzing features such as vocabulary and tone in user queries, the model can determine the user’s emotional state and provide more considerate services.

[0098] (3) Intent recognition: By identifying the user’s intent, the Transformer architecture’s large model can guide the automated system to provide the corresponding service. Intent recognition is a key link in the intelligent customer service system, which determines how the system should respond to the user’s query.

[0099] (3) Knowledge base search:

[0100] The large model of the Transformer architecture can help the system quickly retrieve relevant information from a large knowledge base to answer user questions. This enables the intelligent customer service system to provide accurate answers faster and reduce user waiting time.

[0101] (4) Multimodal input processing: The intelligent customer service system based on the Transformer architecture can also process multimodal input, such as voice, images, etc. This enables the system to understand user needs more comprehensively and provide more personalized services.

[0102] Furthermore, in some modified embodiments, in order to optimize the analysis of the Transformer structure on the text data to obtain more abundant data information related to the user's intention, the embodiments of the present application are exemplified as follows: Figure 4 Schematic diagram of intelligent customer service shown.

[0103] A knowledge fusion layer is added in each encoding layer. The data processing order of the knowledge fusion layer is before the self-attention layer. Before the self-attention layer is used to process the input sequence, the electric power patent knowledge base and the general knowledge base are connected on the knowledge fusion layer to obtain knowledge information; the input sequence is processed in combination with the knowledge information to preliminarily identify the user intention; the preliminarily identified user intention and input sequence are input into the self-attention layer for processing.

[0104] like Figure 4 In the embodiment of the present application, the electric power market knowledge base is not relied on, but the electric power patent knowledge base and the general knowledge base are connected to the knowledge fusion layer to obtain knowledge information to optimize its own large model. For example, a knowledge fusion layer is added in each coding layer, which includes: a joint reasoning layer, an electric power professional knowledge coding layer, and a general knowledge coding layer. The electric power professional knowledge coding layer is connected to the electric power patent knowledge base to obtain knowledge information, and the general knowledge coding layer is connected to the general knowledge base to obtain knowledge information, so as to feed back to the joint reasoning layer to process the input sequence and obtain relevant information for preliminary realization of the user's intention, and pass these data information to the self-attention layer and the forward feedforward layer, thereby further enriching the input sequence originally passed to these two layers, so that the ultimate goal is to output interactive response data that is more in line with the user's intention.

[0105] Furthermore, as a response to the above Figure 1 , 2 The implementation of the method shown in the embodiment of the present application provides an application device based on a large model in the power market intelligent customer service. The device embodiment corresponds to the aforementioned method embodiment. For ease of reading, the device embodiment will not repeat the details of the aforementioned method embodiment one by one, but it should be clear that the device in this embodiment can correspond to all the contents of the aforementioned method embodiment. The device is used to interact with users using intelligent customer service, specifically as follows Figure 5 As shown, the device comprises:

[0106] A receiving unit 31 is used to receive power topic data provided by a user, wherein the power topic data is data for generating a conversation with an intelligent customer service, and the power topic data includes at least text data, voice data, and image data;

[0107] The processing unit 32 is used to process the power topic data using a preset conversation interaction model and output interactive response data corresponding to the power topic data. The preset conversation interaction model is a large language model trained based on a transformer architecture in deep learning. The transformer architecture includes an encoder and a decoder. The encoder is used to analyze the power topic data as text data, wherein the analysis of the text data at least includes extracting key information and understanding the context. The decoder uses the encoder to analyze the text data to gradually generate the required output data.

[0108] Further, such as Figure 6 As shown, the processing unit 32 includes:

[0109] A first processing module 321 is used to convert the power topic data into text data to obtain a vector corresponding to the text data;

[0110] An adding module 322, configured to add a position code to the vector to construct an input sequence corresponding to the vector;

[0111] A determination module 323 is used to determine N encoding layers of the encoder, each of which includes at least a self-attention layer and a feedforward neural network layer; the self-attention layer is used to calculate the attention weights at different positions based on the input sequence, and the feedforward neural network layer is used to perform further nonlinear transformation and mapping on the output of each self-attention layer to learn complex feature representations;

[0112] A second processing module 324 is used to process the input sequence step by step using the N encoding layers to obtain deep features corresponding to the input sequence for generating corresponding hidden representations;

[0113] A first transfer module 325, for transferring the hidden representation to the decoder for generating a target sequence;

[0114] The output module 326 is used to output the interactive response data corresponding to the power topic data according to the target sequence.

[0115] Further, such as Figure 6 As shown, the second processing module 324 is specifically used for:

[0116] In each encoding layer, the input sequence is processed using the self-attention layer to obtain attention weights at different positions on the input sequence;

[0117] Based on inputting the attention weights at different positions on the input sequence into the feedforward neural network layer, feature information corresponding to the input sequence is output.

[0118] Further, such as Figure 6 As shown, the processing unit 32 also includes:

[0119] An adding module 327, configured to add a knowledge fusion layer in each encoding layer, wherein the data processing order of the knowledge fusion layer is before the self-attention layer;

[0120] A connection module 328, used to connect the electric power patent knowledge base and the general knowledge base on the knowledge fusion layer to obtain knowledge information;

[0121] The third processing module 329 is used to process the input sequence in combination with the knowledge information to preliminarily identify the user's intention;

[0122] The second transfer module 330 is used to input the initially identified user intention and the input sequence into the self-attention layer for processing.

[0123] In summary, the application device based on the big model in the intelligent customer service of the electricity market includes a processor and a memory. The above-mentioned receiving unit and processing unit are stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize the corresponding functions.

[0124] The processor contains a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the kernel parameters can be adjusted to use the large model to achieve interactive responses, and to use key information and context to understand such rich data information, so that the interactive effect of the answer is better and more in line with the user's intention, thereby providing a better intelligent customer service solution and improving the quality of intelligent power customer service.

[0125] An embodiment of the present application provides a storage medium on which a program is stored. When the program is executed by a processor, the application method based on a large model in intelligent customer service in the power market is implemented.

[0126] An embodiment of the present application provides a processor, which is used to run a program, wherein the program executes the application method based on a large model in the intelligent customer service of the power market when it is running.

[0127] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program for initializing the steps of the above-mentioned method for applying a large model to intelligent customer service in the electricity market.

[0128] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0129] In a typical configuration, the device includes one or more processors (CPU), memory and bus. The device may also include input / output interface, network interface and the like.

[0130] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip. The memory is an example of a computer-readable medium.

[0131] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0132] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0133] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0134] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. An application method based on a large model in intelligent customer service in the power market, characterized in that: The method comprises: Receive power topic data provided by a user, wherein the power topic data is data generated by a conversation with an intelligent customer service, and the power topic data includes at least text data, voice data, and image data; The power topic data is processed by using a preset conversation interaction model, and interactive response data corresponding to the power topic data is output, wherein the preset conversation interaction model is a large language model trained based on a transformer architecture in deep learning; the transformer architecture includes an encoder and a decoder, the encoder is used to analyze the power topic data as text data, wherein the analysis of the text data at least includes extracting key information and understanding the context; the decoder uses the encoder to analyze the text data to gradually generate the required output data.

2. The method according to claim 1, characterized in that The using of a preset conversation interaction model to process the power topic data and output interactive response data corresponding to the power topic data includes: Convert the power topic data into text data to obtain a vector corresponding to the text data; adding a position code to the vector to construct an input sequence corresponding to the vector; Determine N encoding layers of the encoder, each of the encoding layers at least comprising a self-attention layer and a feedforward neural network layer; the self-attention layer is used to calculate the attention weights at different positions based on the input sequence, and the feedforward neural network layer is used to perform further nonlinear transformation and mapping on the output of each self-attention layer to learn complex feature representations; Using the N encoding layers to gradually process the input sequence, obtain deep features corresponding to the input sequence, and use them to generate corresponding hidden representations; Passing the hidden representation to the decoder for generating a target sequence; According to the target sequence, interactive response data corresponding to the power topic data is output.

3. The method according to claim 2, characterized in that The step of gradually processing the input sequence using the N encoding layers to obtain deep features corresponding to the input sequence for generating corresponding hidden representations includes: In each encoding layer, the input sequence is processed using the self-attention layer to obtain attention weights at different positions on the input sequence; Based on inputting the attention weights at different positions on the input sequence into the feedforward neural network layer, feature information corresponding to the input sequence is output.

4. The method according to claim 3, characterized in that A knowledge fusion layer is added in each encoding layer, wherein the data processing order of the knowledge fusion layer is before the self-attention layer, and before the self-attention layer is used to process the input sequence, the method further includes: On the knowledge fusion layer, the electric power patent knowledge base and the general knowledge base are connected to obtain knowledge information; Processing the input sequence in combination with the knowledge information to preliminarily identify the user's intention; The user intention and the input sequence that are initially identified are input into the self-attention layer for processing.

5. An application device based on a large model in intelligent customer service in the power market, characterized in that: The device comprises: A receiving unit, configured to receive power topic data provided by a user, wherein the power topic data is data for generating a conversation with an intelligent customer service, and the power topic data includes at least text data, voice data, and image data; A processing unit is used to process the power topic data using a preset conversation interaction model and output interactive response data corresponding to the power topic data, wherein the preset conversation interaction model is a large language model trained based on a transformer architecture in deep learning; the transformer architecture includes an encoder and a decoder, wherein the encoder is used to analyze the power topic data as text data, wherein the analysis of the text data at least includes extracting key information and understanding the context; and the decoder uses the encoder to analyze the text data to gradually generate the required output data.

6. The device according to claim 5, characterized in that The processing unit comprises: A first processing module, configured to convert the power topic data into text data to obtain a vector corresponding to the text data; An adding module, used for adding position coding to the vector to construct an input sequence corresponding to the vector; A determination module, used to determine N encoding layers of the encoder, each of which includes at least a self-attention layer and a feedforward neural network layer; the self-attention layer is used to calculate the attention weights at different positions based on the input sequence, and the feedforward neural network layer is used to perform further nonlinear transformation and mapping on the output of each self-attention layer to learn complex feature representations; A second processing module, configured to process the input sequence step by step using the N encoding layers to obtain deep features corresponding to the input sequence, for generating corresponding hidden representations; A first transfer module, used for transferring the hidden representation to the decoder for generating a target sequence; The output module is used to output the interactive response data corresponding to the power topic data according to the target sequence.

7. The device according to claim 6, characterized in that The second processing module is specifically used for: In each encoding layer, the input sequence is processed using the self-attention layer to obtain attention weights at different positions on the input sequence; Based on inputting the attention weights at different positions on the input sequence into the feedforward neural network layer, feature information corresponding to the input sequence is output.

8. The device according to claim 7, characterized in that The processing unit also includes: An adding module is used to add a knowledge fusion layer in each encoding layer, and the data processing order of the knowledge fusion layer is before the self-attention layer; A connection module, used for connecting the electric power patent knowledge base and the general knowledge base on the knowledge fusion layer to obtain knowledge information; A third processing module is used to process the input sequence in combination with the knowledge information to preliminarily identify the user's intention; The second transfer module is used to input the initially identified user intention and the input sequence into the self-attention layer for processing.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the application method based on a large model in intelligent customer service in the power market as described in any one of claims 1 to 4 is implemented.

10. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the application method based on a large model in intelligent customer service in the power market as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Power equipment operation inspection cognitive large model training method and system

    CN117612189A

  • Large robot model and training method and device thereof

    CN118568504A

Cited By

  • Security reply generation method, related device, equipment and storage medium

    CN120408414A

  • Security reply generation method and related apparatus, device, and storage medium

    CN120408414B