Customer service optimization method based on computing power implementation of intelligent computing center and related device

By deploying the target GPT-Neo large language model in the intelligent computing center, key information and emotions in user consultation information are extracted, and more accurate reply text is generated, which solves the answers to unanswered questions in the existing customer service system, and improves the user experience and timeliness of reply.

CN119990148APending Publication Date: 2025-05-13DATACANVAS LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510065456.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing customer service system generates replies based on preset processing rules, which makes replies easy to answer and not ask questions, making it difficult to match the actual consulting needs of users.

Method used

By calling the target GPT-Neo large language model deployed in the intelligent computing center, key information and emotional information in the consultation information are extracted, and multiple to-choose reply texts are generated based on this information, and one text is determined as the target reply text.

Benefits of technology

It improves the accuracy of the reply, makes it better match the user's actual consulting needs and emotions, improves the user's consulting experience, and speeds up information extraction and reply generation through the abundant computing power of the intelligent computing center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990148A_ABST
    Figure CN119990148A_ABST
Patent Text Reader

Abstract

The invention provides a customer service optimization method based on computing power of an intelligent computing center and a related device. The customer service optimization method comprises the following steps: S1, receiving consultation information sent by a user through an interaction end; s2, in response to the consultation information, calling a target GPT-Neo large language model deployed in an intelligent computing center to extract key information and emotion information in the consultation information; s3, generating a plurality of to-be-selected reply texts according to the key information, determining a text from the plurality of to-be-selected reply texts as a target reply text according to the emotion information, and sending the target reply text to the interaction end; the obtained target reply text can accurately match the actual consultation demand of the user; the target reply text can be matched with the actual emotion when the user puts forward the consultation, so that the consultation experience of the user is effectively improved; and reasoning of the target GPT-Neo large language model is realized by adopting the abundant computing power of the intelligent computing center, so that the consultation experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of computing power infrastructure, and in particular, to a customer service optimization method and related devices based on the computing power of an intelligent computing center. Background Art

[0002] With the development of artificial intelligence technology and computing power technology, the concept of intelligent computing center has emerged. "Intelligent computing center" refers to the use of large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning and other scenarios) to provide the required computing power, data and algorithms. Intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0003] Currently, existing customer service systems often generate replies based on preset processing rules according to the consultation information input by the user, that is, the reply text is matched from a preset reply text set according to the keywords in the consultation information, and then the matched reply text is sent to the user. The reply text is easy to irrelevant to the question and is difficult to match the user's actual consultation needs. Summary of the invention

[0004] The embodiment of the present invention provides a customer service optimization method and related devices based on the computing power of an intelligent computing center to solve the problem that the existing customer service system often generates a reply based on the consultation information input by the user based on a preset processing rule, that is, matches the reply text from a preset reply text set according to the keywords in the consultation information, and then sends the matched reply text to the user. The reply text is easy to answer the question irrelevantly and is difficult to match the user's actual consultation needs.

[0005] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0006] In a first aspect, an embodiment of the present invention provides a customer service optimization method based on the computing power of an intelligent computing center, comprising:

[0007] Step S1: receiving consultation information sent by the user through the interactive terminal;

[0008] Step S2: In response to the consultation information, calling the target GPT-Neo large language model deployed in the intelligent computing center to extract key information and emotional information from the consultation information;

[0009] Step S3: generating a plurality of candidate reply texts according to the key information, determining a text as a target reply text from the plurality of candidate reply texts according to the emotion information, and sending the target reply text to the interaction terminal;

[0010] The target GPT-Neo large language model uses the attention mechanism Attention(Q, K, V) to extract the key information and emotional information in the consulting information. The expression of the attention mechanism Attention(Q, K, V) is as follows:

[0011]

[0012] Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key;

[0013] The target GPT-Neo large language model uses a beam search algorithm to infer the consulting information and extract the key information and the emotional information.

[0014] Optionally, the target GPT-Neo large language model is trained using a layer normalization method.

[0015] Optionally, step S2 includes:

[0016] Step S21: determining the computing resources required for the computing task of extracting the key information and the emotional information from the consulting information according to the complexity of the consulting information;

[0017] Step S22: Based on the computing resources, the target GPT-Neo large language model is called to extract key information and emotional information from the consulting information.

[0018] Optionally, the step S22 includes:

[0019] Step S221: Based on the migration of inference tasks between multiple servers of the intelligent computing center specified by the computing power resources, the target GPT-Neo large language model is called to extract the key information and the emotional information.

[0020] Optionally, the computing power resource includes a computing power cluster having at least one of the following architectures:

[0021] NVLink, high-bandwidth memory HBM, dynamic voltage and frequency scaling DVFS, hybrid cooling system of liquid cooling and air cooling, containerized management, and predictive maintenance system.

[0022] In a second aspect, an embodiment of the present invention provides a customer service optimization device based on the computing power of an intelligent computing center, including:

[0023] A receiving module, used for receiving consulting information sent by the user through the interactive terminal;

[0024] An execution module, configured to respond to the consultation information and call a target GPT-Neo large language model deployed in an intelligent computing center to extract key information and emotional information from the consultation information;

[0025] A reply module, configured to generate a plurality of reply texts to be selected according to the key information, determine a text as a target reply text from the plurality of reply texts to be selected according to the emotion information, and send the target reply text to the interaction terminal;

[0026] The target GPT-Neo large language model uses the attention mechanism Attention(Q, K, V) to extract the key information and emotional information in the consulting information. The expression of the attention mechanism Attention(Q, K, V) is as follows:

[0027]

[0028] Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key;

[0029] The target GPT-Neo large language model uses a beam search algorithm to infer the consulting information and extract the key information and the emotional information.

[0030] Optionally, the execution module is further used to determine the computing resources required for the computing task of extracting the key information and the emotional information from the consulting information according to the complexity of the consulting information;

[0031] The execution module is also used to call the target GPT-Neo large language model based on the computing power resources to extract key information and emotional information from the consulting information.

[0032] In a third aspect, an embodiment of the present invention provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the customer service optimization method based on the computing power of an intelligent computing center as described in any one of the first aspects.

[0033] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the customer service optimization method based on the computing power of an intelligent computing center as described in any one of the first aspects are implemented.

[0034] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of the customer service optimization method based on the computing power of an intelligent computing center as described in any one of the first aspects.

[0035] In an embodiment of the present invention, by calling the target GPT-Neo large language model deployed in the intelligent computing center to extract the key information and emotional information in the consulting information, multiple candidate reply texts are generated according to the key information, and a text is determined as the target reply text from the multiple candidate reply texts according to the emotional information, and the target reply text is sent to the interactive end. Compared with the existing technical solution of matching the corresponding preset reply text according to the keywords in the consulting information, the present invention uses the target GPT-Neo large language model to extract key information, and the obtained key information has higher accuracy and can better represent the user's actual consulting needs. Then, the obtained target reply text can accurately match the user's actual consulting needs. The embodiment of the present invention also calls the target GPT-Neo large language model to extract emotional information, and determines a text as the target reply text from the multiple candidate reply texts according to the emotional information, so that the target reply text can match the actual emotion of the user when making a consultation, effectively improving the user's consulting experience. In addition, by calling the target GPT-Neo large language model deployed in the intelligent computing center, the embodiment of the present invention uses the abundant computing power of the intelligent computing center to realize the reasoning of the target GPT-Neo large language model, which speeds up the extraction of key information and emotional information and ensures the high timeliness of the obtained target reply text. Users can receive the reply text more promptly, which improves the user's consultation experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0037] Figure 1 A flowchart of a customer service optimization method implemented based on the computing power of an intelligent computing center according to an embodiment of the present invention;

[0038] Figure 2 This is a principle block diagram of a customer service optimization device implemented based on the computing power of an intelligent computing center according to an embodiment of the present invention;

[0039] Figure 3 The figure is a principle block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] The terms "first", "second", etc. in the embodiments of the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "or" in the present application represents at least one of the connected objects. For example, "A or B" covers three schemes, namely, Scheme 1: including A but not including B; Scheme 2: including B but not including A; Scheme 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.

[0042] In addition, the technical features involved in different implementations of the embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0043] It should be noted that in the embodiments of the present invention, the collection, collection, update, analysis, processing, use, transmission, storage and other aspects of the user personal information involved are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information and network security.

[0044] First, the technical terms involved in the embodiments of the present invention are briefly described below.

[0045] The "computing power" described in the embodiments of the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and it mainly provides services to the society through computing power infrastructure.

[0046] The "computing power" (Computational Power, CP) described in the embodiments of the present invention refers to: it is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .

[0047] The "carrying capacity" (Network Power, NP) described in the embodiment of the present invention refers to: it is the performance of the data transmission capability of the computing power facilities, including the comprehensive capabilities of network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0048] The "Storage Power" (SP) described in the embodiments of the present invention refers to: the comprehensive capabilities of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB = 2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.

[0049] The "computing power infrastructure" described in the embodiments of the present invention is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity. It can realize centralized computing, storage, transmission and application of information, and presents characteristics such as diversity and ubiquity, intelligence and agility, security and reliability, green and low-carbon. It is of great significance to promote industrial transformation and upgrading, enable my country's scientific and technological innovation, meet people's needs for a better life, and achieve high-efficiency social governance.

[0050] The “general computing power” described in the embodiment of the present invention refers to: the computing power provided by a server based on a CPU (Central Processing Unit) chip, which is used to support basic general computing such as cloud computing and edge computing.

[0051] The "intelligent computing power" described in the embodiment of the present invention refers to: a computing platform based on GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) and other dedicated chips for various artificial intelligence innovative applications, such as natural language processing, machine vision, etc.

[0052] The "super computing power" described in the embodiments of the present invention refers to: the computing power mainly provided by high-performance computing clusters such as supercomputers. It uses the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0053] The "intelligent computing center" described in the embodiments of the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning) by using large-scale heterogeneous computing resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0054] The "computing resources" described in the embodiments of the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.

[0055] The “large language model” described in the embodiment of the present invention refers to a large-scale language model (LLM), which is a language model with a large parameter scale, designed to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0056] At present, existing customer service systems often generate replies according to the consultation information input by the user based on preset processing rules, that is, the reply text is matched from the preset reply text set according to the keywords in the consultation information, and then the matched reply text is sent to the user. For example, there is "Hello" in the consultation information, and the reply text matched by the keyword "Hello" from the preset reply text set is "Hello, how can I help you?". As long as the keyword "Hello" exists in the consultation information, the corresponding reply text is "Hello, how can I help you?". It can only reply mechanically, and cannot generate a reply file according to the user's consultation focus (in this example, that is, the user's actual consultation needs contained in other information other than "Hello" in the consultation information). The reply text is easy to answer irrelevant questions and is difficult to match the user's actual consultation needs.

[0057] For another example, a complete consultation message is "Hello, where is my delivery?" In the context of this consultation message, the user uses the colloquial word "things" to refer to "express delivery". The keyword "things" cannot match the reply text that "express delivery" can match from the preset reply text set; therefore, the reply text will be the text that matches "Hello", "Hello, how can I help you?" The reply text is likely to be irrelevant and difficult to match the user's actual consultation needs.

[0058] The embodiment of the present invention provides a customer service optimization method based on the computing power of an intelligent computing center, see Figure 1 As shown, Figure 1 The flowchart of the customer service optimization method implemented based on the computing power of the intelligent computing center according to an embodiment of the present invention is as follows:

[0059] Step S1: receiving consultation information sent by the user through the interactive terminal;

[0060] Step S2: In response to the consultation information, the target GPT-Neo large language model deployed in the intelligent computing center is called to extract key information and emotional information from the consultation information;

[0061] Step S3: generating a plurality of candidate reply texts according to the key information, determining a text as a target reply text from the plurality of candidate reply texts according to the emotional information, and sending the target reply text to the interaction terminal;

[0062] Among them, the target GPT-Neo large language model uses the attention mechanism Attention(Q,K,V) to extract key information and emotional information from the consulting information. The expression of the attention mechanism Attention(Q,K,V) is as follows:

[0063]

[0064] Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key;

[0065] The target GPT-Neo large language model uses the beam search algorithm to infer consulting information and extract key information and emotional information.

[0066] In the embodiment of the present invention, customer service refers to an artificial intelligence customer service system used to respond to user (customer) inquiries.

[0067] It should be noted that the GPT-Neo large language model is an open source language model developed by EleutherAI, which aims to provide similar functions to OpenAI's GPT-3. GPT-Neo is a model based on the Transformer architecture, mainly used for natural language processing (NLP) tasks such as text generation, question answering, translation, etc.

[0068] The target GPT-Neo large language model in the embodiment of the present invention is an improved model generated from the open source GPT-Neo large language model, at least improving the attention mechanism and reasoning algorithm. The expression of the attention mechanism Attention(Q, K, V) used by the target GPT-Neo large language model is as follows:

[0069]

[0070] Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key, d k Used to scale the dot product results to prevent gradient instability caused by excessive values.

[0071] The attention mechanism described by the above expression, Attention(Q,K,V), calculates the similarity between the query and the key (scaled dot product) and then weighted sums the values ​​to obtain the attention output. This attention mechanism effectively improves the processing speed and output quality of the model.

[0072] The attention mechanism Attention(Q,K,V) can significantly reduce the computational complexity of long sequence processing. Specifically, the computational complexity is reduced in the following ways:

[0073] Scaled Dot-Product Attention: Introducing a scaling factor when calculating the similarity between the query and the key It avoids the problem of excessively large values ​​caused by long sequence input, improves numerical stability, and thus optimizes computational efficiency.

[0074] Matrix operation optimization: Use GPU to accelerate matrix operations (for example, QK T), enabling large-scale computing tasks to be executed quickly.

[0075] Sparse attention mechanism: By sparsely focusing on the similarity matrix, only the key parts of the input sequence are focused on, which significantly reduces the computational cost (for example, from O(n 2 )) is reduced to O(nlogn) or lower).

[0076] The attention mechanism Attention(Q, K, V) can also better focus on the relevant parts of the input data (specifically, the consulting information in the embodiment of the present invention), thereby improving the processing speed and accuracy. Specifically, the input data can be better focused in the following ways:

[0077] Softmax weighted summation: The similarity between the query and the key is normalized by the softmax function, retaining the important parts (high weights) and ignoring the secondary information (low weights), so that the model pays more attention to the key parts of the input sequence.

[0078] Dynamic routing mechanism: Combined with the weight distribution of the attention mechanism, it dynamically selects the processing focus to adaptively allocate computing resources.

[0079] More efficient attention heads (Multi-Head Attention): Through the parallel multi-head attention mechanism, the model can analyze input data from multiple angles and improve the ability to capture relevant information.

[0080] Beam Search is a heuristic search algorithm widely used in natural language processing (NLP), machine translation, speech recognition and other fields, especially in generation tasks. It is an improved breadth-first search method that aims to find a near-optimal solution in the search space while controlling the complexity of the search.

[0081] In an embodiment of the present invention, through the attention mechanism Attention (Q, K, V) of the target GPT-Neo large language model and the inference algorithm of the target GPT-Neo large language model including the beam search algorithm, the core content of user needs can be accurately extracted from long texts, greatly reducing computational redundancy; ensuring that the generated replies focus on the key points of user input, thereby improving the pertinence and accuracy of the responses; combined with the computing power of the intelligent computing center, high-throughput processing is achieved through efficient reasoning optimized by the attention mechanism.

[0082] Understandably, in reality, the consulting information that users input into the customer service system is often long text information with a lot of redundant information. The target GPT-Neo large language model uses the attention mechanism Attention(Q, K, V) to quickly locate key content (i.e. key information) in the user's long text input (i.e. consulting information). For example, when a user inputs "Why hasn't my order been shipped yet? It's been two weeks since I placed the order", the attention mechanism will automatically focus on keywords such as "order", "shipping", and "two weeks", while ignoring other irrelevant descriptions. In addition, the attention mechanism Attention(Q, K, V) gives priority to the core emotional words expressed by the user. For example, from the frequent appearance of "why" and "not shipped yet" in the consulting information, the question is classified as "logistics inquiry", and the user's emotions are analyzed as "anxiety or dissatisfaction" (i.e. emotional information), laying the foundation for subsequent reply generation.

[0083] In an embodiment of the present invention, multiple reply texts to be selected are generated according to key information. Specifically, a beam search algorithm combined with an attention mechanism is used to generate multiple reply texts to be selected, that is, according to the core semantics of the user's consultation information, a corresponding solution or reply scheme is determined, and multiple reply texts to be selected are generated. For example, from the frequent occurrence of "why" and "not shipped yet" in the consultation information, the problem is classified as "logistics query (i.e. key information)", and then "the courier company has been urged to deliver as soon as possible", "sorry, I understand your anxious mood, and the courier company has been urged to deliver as soon as possible", "sorry, I understand your anxious mood, and the courier of your order is currently located in XXX ("XXX" refers to a place name, generated according to the data of the logistics company), and it is very close to the delivery place. Please wait patiently", "sorry, I understand your anxious mood, we are working hard to solve it, and we have reported it to the superiors" and other multiple reply texts to be selected.

[0084] It should be noted that determining a text as the target reply text from multiple candidate reply texts based on the emotion information can be to determine the text that best matches the emotion information from multiple candidate reply texts. For example, if the emotion is anxious, the user should be comforted with actual data, and the target reply text is "Sorry, I understand your anxiety. I found that the courier for your order is currently located in XXX ("XXX" refers to a place name, generated based on the actual data of the logistics company), and is very close to the delivery location. Please wait patiently." For example, if the emotion is stable, the reply should be made with a short sentence that expresses the attitude of taking the problem seriously, and the target reply text is "The courier company has been urged to deliver the package as soon as possible."

[0085] According to the emotion information, a text is determined as the target response text from multiple candidate response texts. Alternatively, according to a preset emotion-response word relationship mapping table, the text containing the most response words corresponding to the current emotion is determined from multiple candidate response texts as the target response text. For example, the emotion information is anxiety. According to the preset emotion-response word relationship mapping table, the response words corresponding to anxiety are honorifics or soothing sentences such as "you", "we are working hard to solve it", and "we have reported to the superiors". Further, among the multiple candidate response texts, the text containing the most honorifics or soothing sentences such as "you", "we are working hard to solve it", and "we have reported to the superiors" is selected as the target response text.

[0086] The adjustments mentioned in the patent are mainly about the fine-tuning and optimization of the depth and width of the large language model (LLM). These adjustments are aimed at finding an ideal balance between performance and computing requirements to ensure that the model can run efficiently and stably in tasks of different scales.

[0087] In the embodiment of the present invention, the target GPT-Neo large language model includes at least one of the following improvements:

[0088] 1) Optimization of depth and width:

[0089] Depth adjustment: By increasing or decreasing the number of layers in the model, its performance on complex tasks can be enhanced while avoiding over-computation when the tasks are not complex.

[0090] Width adjustment: By adjusting the number of neurons in each layer, the model can more accurately match the requirements of specific tasks, thereby achieving a balance between performance and resource utilization.

[0091] 2) Detailed design of model architecture:

[0092] The model structure has been optimized to make it more applicable in a variety of application scenarios. Regardless of the size of the task, the model can meet the accuracy requirements while ensuring computational efficiency through optimized depth and width.

[0093] 3) Integration of assistive technologies:

[0094] Layer Normalization: It speeds up the model training process and reduces the problem of gradient vanishing during training. Its function is to standardize the output of each layer to make the training more stable.

[0095] Residual Connection: This allows information to flow more smoothly between different layers. This connection method effectively avoids the gradient vanishing problem common in deep networks and accelerates the convergence process.

[0096] In some optional embodiments, the above 1) 2) 3) are combined to enhance the convergence speed and overall stability of the target GPT-Neo large language model. In actual tests, the target GPT-Neo large language model has shown higher adaptability and better energy efficiency when facing tasks of different scales.

[0097] In an embodiment of the present invention, by calling the target GPT-Neo large language model deployed in the intelligent computing center to extract the key information and emotional information in the consulting information, multiple candidate reply texts are generated according to the key information, and a text is determined as the target reply text from the multiple candidate reply texts according to the emotional information, and the target reply text is sent to the interactive end. Compared with the existing technical solution of matching the corresponding preset reply text according to the keywords in the consulting information, the present invention uses the target GPT-Neo large language model to extract key information, and the obtained key information has higher accuracy and can better represent the user's actual consulting needs. Then, the obtained target reply text can accurately match the user's actual consulting needs. The embodiment of the present invention also calls the target GPT-Neo large language model to extract emotional information, and determines a text as the target reply text from the multiple candidate reply texts according to the emotional information, so that the target reply text can match the actual emotion of the user when making a consultation, effectively improving the user's consulting experience. In addition, by calling the target GPT-Neo large language model deployed in the intelligent computing center, the embodiment of the present invention uses the abundant computing power of the intelligent computing center to realize the reasoning of the target GPT-Neo large language model, which speeds up the extraction of key information and emotional information and ensures the high timeliness of the obtained target reply text. Users can receive the reply text more promptly, which improves the user's consultation experience.

[0098] In some embodiments of the present invention, optionally, the target GPT-Neo large language model is trained using a layer normalization method.

[0099] Layer Normalization is a normalization technique used in deep learning models, mainly used to improve the stability of the training process and accelerate convergence. The layer normalization method provides a more stable training process and better model performance by independently normalizing the features of each sample.

[0100] In some embodiments of the present invention, optionally, step S2 includes:

[0101] Step S21: determining the computing resources required for the computing task of extracting key information and emotional information from the consulting information according to the complexity of the consulting information;

[0102] Step S22: Based on computing resources, the target GPT-Neo large language model is called to extract key information and emotional information from the consulting information.

[0103] In an embodiment of the present invention, each time a user initiates a consultation and executes step S2 in response to the consultation information, the computing power resources required for the computing task of extracting key information and emotional information from the consultation information are determined according to the complexity of the consultation information; the target GPT-Neo large language model is called to extract the key information and emotional information in the consultation information according to the computing power resources, thereby realizing dynamic adjustment of computing power resources with the complexity of information processing contained in the consultation information, which can effectively improve the utilization efficiency of computing power resources.

[0104] In some embodiments of the present invention, optionally, step S22 includes:

[0105] Step S221: Based on the migration of reasoning tasks between multiple servers in the intelligent computing center specified by computing resources, the target GPT-Neo large language model is called to extract key information and emotional information.

[0106] In the embodiment of the present invention, the intelligent computing center can be a computing center deployed in the cloud, and the target GPT-Neo large language model is the local serverless cloud computing large language model ServerlessLLM. Under the ServerlessLLM architecture, the intelligent computing center allocates computing resources on demand, which can adapt to dynamic workloads more flexibly and efficiently, and is particularly suitable for processing irregular and highly concurrent reasoning requests.

[0107] In a serverless architecture, LLM reasoning tasks usually require a large amount of KV cache. The existing migration method requires transferring the entire KV cache (Key-Value Cache, which stores intermediate data generated during the model reasoning process to accelerate subsequent reasoning steps. The KV cache is usually very large. If it is transferred during the migration process, it will increase the network load) from the source server to the target server, causing a large amount of network load.

[0108] Based on this, the types of reasoning task migration adopted in the embodiment of the present invention include: lightweight migration based on Token and multi-round iterative migration process.

[0109] Token-based lightweight migration: During the real-time migration process, the source server only transmits the generated tokens (small data volume, about 10-100KB) instead of the complete KV cache (usually 1-10GB) to reduce network transmission overhead. After the target server receives the token data, it calls a specific cache synchronization algorithm to incrementally update the state of the KV cache to achieve smooth migration of the task. In addition, pipeline transmission and asynchronous cache synchronization are used to further reduce the load pressure and network latency of the target server, thereby ensuring the continuity of the reasoning task.

[0110] Multi-round iterative migration process: Inference requests can be seamlessly transferred between multiple GPU servers, and the system rebuilds the cache on the target server and continues the inference operation. This approach reduces latency during the migration process and ensures real-time responsiveness.

[0111] The specific process of reasoning task migration in the embodiment of the present invention includes:

[0112] Step 1: The load scheduler sends a model loading request to the target server, instructing it to load the model required for the inference task on the GPU. For example, load model A onto the GPU of the target server. If the model already exists on the target server, skip this step.

[0113] Step 2: After loading is complete, the scheduler sends a migration request to the source server, informing it of the address of the target server.

[0114] Step 3: After receiving the migration request, the source server starts marking the "migrating" state and sends a "resume request" to the target server. This request contains intermediate tokens (including input tokens and generated output tokens; the input token refers to the input content of the reasoning task passed from the original server to the target server (for example, the text or query entered by the user). This is the starting point of the reasoning task and is used by the target server to continue the reasoning operation; the output token refers to the partial result of the reasoning task that the original server has completed (for example, the generated partial text or answer). These output tokens are the output of the original server and will also serve as the input of the target server to help the target server continue reasoning from the intermediate state). If the reasoning is completed, it returns directly to the scheduler and skips the subsequent steps.

[0115] Step 4: The target server rebuilds the KV cache based on the received token (reconstruction means that the target server recalculates the intermediate state (KV cache) required for the reasoning task based on the received token, so as to continue the subsequent tasks from the reasoning progress of the source server) to ensure that reasoning can continue from this state. The reconstruction speed of the KV cache is usually much faster than the time required to generate the same number of new tokens.

[0116] Step 5: After the recovery request is completed, the source server stops reasoning, returns a response containing all tokens to the scheduler (including the target server rebuilding the KV cache through the received tokens in step 4), and marks it as "migrated".

[0117] Step 6: The scheduler completes the migration, unloads the model from the source server, and starts loading other tasks (e.g., Model B) on the source server.

[0118] Step 7: The request router checks the "migrate" flag in the response. If it has been migrated, the routing table is updated to replace the source server with the target server so that the inference task can continue.

[0119] In addition to the migration of inference tasks, the ServerlessLLM architecture in the embodiment of the present invention also includes at least one of the following designs: fast multi-level checkpoint loading and a scheduling strategy for optimizing startup time.

[0120] Fast multi-level checkpoint loading:

[0121] The ServerlessLLM architecture in the embodiment of the present invention proposes a multi-level checkpoint loading mechanism to improve the model loading efficiency, making full use of the multi-level storage bandwidth and capacity of the GPU server and reducing the loading time from storage to GPU memory. The mechanism includes the following key technologies:

[0122] Load-optimized checkpoint format: To reduce cold start delays, the present invention adopts a block-based load-optimized checkpoint format. The parameters of each model are read sequentially in blocks (blocks are divided according to the actual needs of the user) (that is, the model parameters are managed by blocks, and only the part required for the current reasoning task is loaded each time, instead of loading the entire model at one time to avoid resource waste), avoiding delays caused by random I / O operations. This format ensures that checkpoint data can be efficiently addressed in memory and allocated to the GPU. Checkpoint: A state file saved during model training and reasoning, including the model's weights and parameter configuration, used for rapid recovery after model interruption or to allocate model parameters in a distributed system. Checkpoint files are generally large, so efficient loading of checkpoints is critical to serverless reasoning performance.

[0123] Multi-level loading subsystem: The system implements pipeline loading based on memory pool and multi-threading, gradually loading the model from local SSD, DRAM and other multi-level storage to GPU (gradual loading, that is, only the generated token data (10-100KB) is transferred during migration, instead of the complete KV cache (1-10GB), which greatly reduces the network load and transmission delay), and reduces data transmission delay through direct I / O operations. Experiments show that this mechanism significantly improves the loading efficiency of LLM models of different sizes, which is 3.6 to 8.2 times higher than the traditional loading method. Multi-level storage (architecture): refers to a hierarchical storage structure containing different storage media, including DRAM, NVMe SSD, etc. This architecture can efficiently store and read data between different storage devices to improve system performance. Direct I / O: A data transmission method that bypasses the operating system cache and directly transfers data from the storage device to the memory or GPU to improve transmission efficiency.

[0124] The combination of a load-optimized checkpoint format and a multi-level loading subsystem enables efficient execution of inference tasks with limited computing and network resources.

[0125] Model scheduling strategy for startup time optimization:

[0126] The ServerlessLLM scheduler schedules the model based on the server's storage hierarchy and the locality of the checkpoints, accurately estimates the loading time and migration time of different servers, and selects the server with the smallest latency to perform the inference task:

[0127] Model loading time estimation: The scheduler estimates the loading time by calculating factors such as queue waiting time, model size, storage bandwidth, etc. For multi-level storage (such as SSD and DRAM) in the model loading path, the scheduler prioritizes the level with the lowest bandwidth to ensure estimation accuracy.

[0128] Migration time estimation and dynamic scheduling: For inference tasks that need to be migrated, the system estimates the time to rebuild the KV cache based on the token generation situation, and dynamically selects the server with the lowest latency for allocation. During peak load periods, the system can realize dynamic migration of requests and latency-minimizing scheduling, improving overall resource utilization.

[0129] Through three innovative designs, namely, inference task migration, fast multi-level checkpoint loading, and scheduling strategies that optimize startup time, ServerlessLLM effectively improves the loading speed, real-time migration capability, and scheduling efficiency of LLM inference, solving the bottleneck problems of high latency and resource utilization in existing serverless inference systems. In actual tests, the response speed of ServerlessLLM has increased by 10 to 200 times, significantly reducing the startup time of large-scale inference tasks, and is suitable for LLM services that require fast response.

[0130] In some embodiments of the present invention, optionally, the computing power resources include a computing power cluster having at least one of the following architectures:

[0131] NVLink, high-bandwidth memory HBM, dynamic voltage and frequency scaling DVFS, hybrid cooling system of liquid cooling and air cooling, containerized management, and predictive maintenance system.

[0132] A) High-speed NVLink: The cluster uses NVLink technology, a high-bandwidth, low-latency interconnection method that can achieve a throughput of up to 300GB per second. NVLink transmits data in parallel through multiple channels, significantly reducing communication latency between nodes, and is suitable for application scenarios such as deep learning and large-scale parallel computing.

[0133] B) High Bandwidth Memory (HBM): High Bandwidth Memory Module (HBM) is used to accelerate data access, providing higher bandwidth than traditional DDR memory (up to 1TB / s). HBM's stacked design and layout close to the computing unit reduce data transmission latency, significantly improving overall performance when processing large data sets.

[0134] C) Dynamic Voltage and Frequency Scaling (DVFS): DVFS technology uses real-time load monitoring to dynamically adjust system power consumption. The specific algorithm includes a regulation method based on P-control (proportional control):

[0135] Set the target frequency f target , current frequency f current , the load change rate is ΔP, and ΔP affects the reference value of frequency adjustment.

[0136] The adjustment algorithm can be expressed as: new =f current +K p ·(P target -P current )

[0137] Among them, f new For frequency, DVFS (Dynamic Voltage and Frequency Scaling) calculates the final result of the adjusted system frequency. currentThe frequency at which the current system is running, indicating the working frequency in the current state. K p P is the proportional gain coefficient, which is used to adjust the control system response rate parameter and reflects the system characteristics. target is the target power, the ideal power consumption target for system design. current It is the current power, which monitors the actual power consumption of the system in real time.

[0138] The algorithm ensures that the system reduces the frequency when the load is low to reduce power consumption, and increases the frequency when the load is high to meet computing needs.

[0139] D) Hybrid cooling system of liquid cooling and air cooling:

[0140] Liquid cooling systems use coolant to flow through high-heat-density components (such as GPUs and CPUs) to remove heat, and work with efficient radiators and fan systems to ensure that the flowing coolant can reduce the temperature in the shortest possible time. Specific steps: Design a closed-loop liquid cooling system to ensure smooth circulation of coolant. Arrange temperature sensors near key heat sources to monitor temperature changes in real time. The control system automatically adjusts the flow rate of the liquid pump and the fan speed based on feedback from the temperature sensor to maintain the system temperature within a safe range.

[0141] E) The cluster uses Docker or Kubernetes for container management. The specific implementation methods include:

[0142] Use Docker to package applications into containers to ensure their environmental consistency. Use Kubernetes to orchestrate containers and automatically manage container deployment, expansion, and load balancing. Implement CI / CD processes to automate the building, testing, and deployment of containerized applications. Monitor container operation status and dynamically adjust resource allocation to meet different workload requirements. In addition, use Helm to manage Kubernetes application configurations to simplify the deployment process of complex applications.

[0143] F) Predictive Maintenance System:

[0144] The integrated predictive maintenance system uses machine learning algorithms to analyze equipment operation data to predict potential failures. Specific algorithm: Random Forest model is used for failure prediction.

[0145] Implementation steps: Collect system sensor data (such as temperature, humidity, workload, etc.). Label historical fault data (give each piece of data a clear label, such as "normal" or "fault", or further mark the specific fault type), and train a random forest model. Regularly update the model with real-time data and adjust the prediction results. Set a threshold, and when the predicted probability of failure exceeds the threshold, trigger an early warning and schedule maintenance.

[0146] The output of the model is the failure probability P(fault), and the predictive maintenance algorithm can be expressed as:

[0147]

[0148] Here, z is the linear combination of the input features calculated by the random forest model.

[0149] The optimization of the target GPT-Neo large language model is described below with reference to specific embodiments.

[0150] The target GPT-Neo large language model also integrates an intelligent caching mechanism that allows the reuse of previously calculated results, avoiding redundant calculations and significantly reducing latency during reasoning. This caching mechanism is particularly useful in real-time applications, especially when repeated queries or processing of similar inputs are required, which can significantly reduce response time.

[0151] In order to efficiently process larger data sets and models, the target GPT-Neo large language model uses model parallel technology to distribute model parameters on multiple GPUs. This parallelization method achieves effective expansion without sacrificing speed and accuracy, ensuring smooth processing of a large number of complex tasks.

[0152] In addition, the load balancing algorithm is optimized to evenly distribute inference tasks to available computing resources, thereby maximizing throughput and minimizing latency, ensuring optimal resource utilization. The inference cluster supporting the target GPT-Neo large language model also incorporates energy efficiency management technologies such as dynamic voltage and frequency scaling (DVFS), which adjusts power consumption according to real-time workloads, achieving significant energy savings.

[0153] At the hardware level, the redesigned cluster architecture introduces high-bandwidth memory (HBM) modules and advanced interconnect technology. These hardware optimizations reduce latency in data transmission, improve overall throughput, and enable the inference cluster to efficiently handle high-demand tasks. In addition, a customized cooling solution ensures stable operating temperature, thereby improving system reliability and efficiency.

[0154] These improvements make the reasoning process of the target GPT-Neo large language model more efficient, scalable, and environmentally sustainable, meeting the market demand for high-performance, low-energy AI solutions. These optimizations provide a strong foundation for processing complex natural language tasks and set a new standard for the reasoning efficiency of large-scale language models.

[0155] In order to evaluate the performance and energy efficiency of the redesigned cluster (the computing power cluster architecture of the intelligent computing center includes: NVLink, high-bandwidth memory HBM, dynamic voltage and frequency adjustment DVFS, a hybrid cooling system of liquid cooling and air cooling, containerized management, and predictive maintenance system) and the target GPT-Neo large language model, a series of key indicators were used. These performance indicators include throughput (the number of input sequences processed per second), latency (the time required to generate an output sequence), the accuracy of model predictions, the scalability of the system (performance improvement obtained by increasing computing resources), and the robustness of the system (performance under different loads and input complexities). Through these indicators, the performance of the new system under different conditions was comprehensively evaluated, verifying its improvements in performance and stability.

[0156] Energy efficiency indicators include the ratio of total data center power consumption to computing resource power consumption (PUE, Power Usage Effectiveness, which means energy efficiency index in Chinese and is used to measure the energy efficiency of data centers. The lower the PUE, the higher the energy efficiency), energy consumption per inference, and carbon footprint (total greenhouse gas emissions). Experimental results show that the optimized cluster is significantly better than the baseline configuration in both throughput and latency. For example, when processing five input sequences simultaneously, the new cluster can process 2,500 sequences per second, while the baseline cluster can only reach 1,500 sequences. In addition, the average latency of the new cluster is 50 milliseconds, significantly lower than the 80 milliseconds of the baseline cluster, showing a significant improvement in response speed.

[0157] In the GLUE benchmark, the target GPT-Neo large language model showed an accuracy improvement of about 5% on a standard dataset, further demonstrating its optimization in inference efficiency and accuracy. In terms of energy efficiency, the new cluster configuration has a PUE of 1.2, which is higher than the baseline cluster's 1.5. The redesigned cluster consumes 22 joules of energy per inference when processing five input sequences simultaneously, while the baseline cluster requires 42 joules. In addition, carbon footprint analysis shows that the new cluster's greenhouse gas emissions are reduced by about 33% compared to the baseline configuration, showing a significant improvement in environmental sustainability.

[0158] Through these improvements, the new cluster and the target GPT-Neo large language model have achieved comprehensive improvements in throughput, latency, accuracy, energy consumption, and carbon emissions. These optimizations not only improve the adaptability of the model in practical applications, but also significantly reduce environmental impact in terms of energy efficiency, providing an effective reference for the sustainable development of large-scale AI systems.

[0159] The embodiment of the present invention provides a customer service optimization device based on the computing power of an intelligent computing center, see Figure 2 As shown, Figure 2The schematic diagram of the customer service optimization device implemented based on the computing power of the intelligent computing center according to an embodiment of the present invention is as follows:

[0160] The receiving module 21 is used to receive the consultation information sent by the user through the interactive terminal;

[0161] An execution module 22 is used to call a target GPT-Neo large language model deployed in an intelligent computing center to extract key information and emotional information from the consultation information in response to the consultation information;

[0162] A reply module 23, configured to generate a plurality of candidate reply texts according to the key information, determine a text as a target reply text from the plurality of candidate reply texts according to the emotion information, and send the target reply text to the interaction terminal;

[0163] The target GPT-Neo large language model uses the attention mechanism Attention(Q, K, V) to extract the key information and emotional information in the consulting information. The expression of the attention mechanism Attention(Q, K, V) is as follows:

[0164]

[0165] Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key;

[0166] The target GPT-Neo large language model uses a beam search algorithm to infer the consulting information and extract the key information and the emotional information.

[0167] In some embodiments of the present invention, optionally, the target GPT-Neo large language model is trained using a layer normalization method.

[0168] In some embodiments of the present invention, optionally, the execution module 22 is further used to determine the computing resources required for the computing task of extracting the key information and the emotional information from the consulting information according to the complexity of the consulting information;

[0169] The execution module 22 is also used to call the target GPT-Neo large language model based on the computing power resources to extract key information and emotional information from the consulting information.

[0170] In some embodiments of the present invention, optionally, the execution module 22 is also used to call the target GPT-Neo large language model to extract the key information and the emotional information based on the migration of reasoning tasks between multiple servers of the intelligent computing center specified by the computing power resources.

[0171] In some embodiments of the present invention, optionally, the computing power resources include a computing power cluster having at least one of the following architectures:

[0172] NVLink, high-bandwidth memory HBM, dynamic voltage and frequency scaling DVFS, hybrid cooling system of liquid cooling and air cooling, containerized management, and predictive maintenance system.

[0173] The customer service optimization device based on the computing power of the intelligent computing center provided in the embodiment of the present application can achieve Figure 1 The various processes implemented by the method embodiment and achieving the same technical effect are not described here to avoid repetition.

[0174] The embodiment of the present invention provides an electronic device 30, see Figure 3 As shown, Figure 3 This is a principle block diagram of an electronic device 30 according to an embodiment of the present invention, including a processor 31, a memory 32, and a program or instruction stored in the memory 32 and executable on the processor 31. When the program or instruction is executed by the processor, any step of the customer service optimization method based on the computing power of an intelligent computing center according to the present invention is implemented.

[0175] An embodiment of the present invention provides a readable storage medium, on which programs or instructions are stored. When the programs or instructions are executed by a processor, the various processes of an embodiment of a customer service optimization method based on the computing power of an intelligent computing center, such as any one of the above, are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.

[0176] The readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc. In some examples, the readable storage medium may be a non-transitory readable storage medium.

[0177] An embodiment of the present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the various processes of any of the above-mentioned customer service optimization method embodiments based on the computing power of an intelligent computing center are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.

[0178] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. A customer service optimization method based on the computing power of an intelligent computing center, characterized in that: include: Step S1: receiving consultation information sent by the user through the interactive terminal; Step S2: In response to the consultation information, calling the target GPT-Neo large language model deployed in the intelligent computing center to extract key information and emotional information from the consultation information; Step S3: generating a plurality of candidate reply texts according to the key information, determining a text as a target reply text from the plurality of candidate reply texts according to the emotion information, and sending the target reply text to the interaction terminal; The target GPT-Neo large language model uses the attention mechanism Attention(Q, K, V) to extract the key information and emotional information in the consulting information. The expression of the attention mechanism Attention(Q, K, V) is as follows: Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key; The target GPT-Neo large language model uses a beam search algorithm to infer the consulting information and extract the key information and the emotional information.

2. The customer service optimization method based on the computing power of the intelligent computing center according to claim 1 is characterized in that: The target GPT-Neo large language model is trained using a layer normalization method.

3. The customer service optimization method based on the computing power of the intelligent computing center according to claim 1 is characterized in that: The step S2 comprises: Step S21: determining the computing resources required for the computing task of extracting the key information and the emotional information from the consulting information according to the complexity of the consulting information; Step S22: Based on the computing resources, the target GPT-Neo large language model is called to extract key information and emotional information from the consulting information.

4. The customer service optimization method based on the computing power of the intelligent computing center according to claim 3 is characterized in that: The step S22 comprises: Step S221: Based on the migration of inference tasks between multiple servers of the intelligent computing center specified by the computing power resources, the target GPT-Neo large language model is called to extract the key information and the emotional information.

5. The customer service optimization method based on the computing power of the intelligent computing center according to claim 3 is characterized in that: The computing resources include a computing cluster having at least one of the following architectures: NVLink, high-bandwidth memory HBM, dynamic voltage and frequency scaling DVFS, hybrid cooling system of liquid cooling and air cooling, containerized management, and predictive maintenance system.

6. A customer service optimization device based on the computing power of an intelligent computing center, characterized in that: include: A receiving module, used for receiving consulting information sent by the user through the interactive terminal; An execution module, configured to respond to the consultation information and call a target GPT-Neo large language model deployed in an intelligent computing center to extract key information and emotional information from the consultation information; A reply module, configured to generate a plurality of reply texts to be selected according to the key information, determine a text as a target reply text from the plurality of reply texts to be selected according to the emotion information, and send the target reply text to the interaction terminal; The target GPT-Neo large language model uses the attention mechanism Attention(Q, K, V) to extract the key information and emotional information in the consulting information. The expression of the attention mechanism Attention(Q, K, V) is as follows: Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key; The target GPT-Neo large language model uses a beam search algorithm to infer the consulting information and extract the key information and the emotional information.

7. The customer service optimization device based on the computing power of the intelligent computing center according to claim 6 is characterized in that: The execution module is further used to determine the computing resources required for the computing task of extracting the key information and the emotional information from the consulting information according to the complexity of the consulting information; The execution module is also used to call the target GPT-Neo large language model based on the computing power resources to extract key information and emotional information from the consulting information.

8. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps in the customer service optimization method based on the computing power of an intelligent computing center as described in any one of claims 1 to 5.

9. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps in the customer service optimization method based on the computing power of the intelligent computing center as described in any one of claims 1 to 5 are implemented.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the customer service optimization method based on the computing power of an intelligent computing center as described in any one of claims 1 to 5.

Citation Information

Cited By

  • Method and device for intelligent computing center to consume computing power by reasoning Serverless

    CN120494103A