Text generation method and electronic equipment

By introducing a text evaluation model into the MoE large language model, dynamically adjusting the number of activated experts, solving the problem of resource mismatch when existing models deal with complex or simple prompt information, and achieving more efficient and high-quality text generation services.

CN120012727AActive Publication Date: 2025-05-16BEIJING FEISHU TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510479915.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

When the existing MoE-based large language model processes prompt information, the number of activated experts is fixed, which makes the model not flexible enough and difficult to adapt to the processing needs of complex or simple prompt information, resulting in resource mismatch and inefficiency.

Method used

By introducing a text evaluation model, the number of activated experts in the MoE big model is dynamically adjusted according to the complexity score of the prompt information, thereby realizing dynamic control and resource optimization of the model.

Benefits of technology

It realizes dynamic adjustment of the number of experts based on the complexity of the prompt information, avoids resource mismatch, improves resource utilization, reduces hardware and energy costs, and provides more efficient and high-quality text generation services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012727A_ABST
    Figure CN120012727A_ABST
Patent Text Reader

Abstract

The invention relates to a text generation method and electronic equipment. The method comprises the steps that prompt information is input into a MoE large model, vector representation corresponding to the prompt information is determined, and the MoE large model has the maximum expert number determined through training; a score value of the prompt information is determined based on vector representation by using a text evaluation model, and the score value of the prompt information is used for representing a quantitative performance index of the complexity of the prompt information; based on the score value of the prompt information, the number of experts to be activated in the MoE large model is determined, and the number of the experts to be activated is not larger than the maximum number of the experts; and generating an output corresponding to the prompt information based on the activated expert of the MoE large model. In this way, the MoE large model can dynamically adjust the number of activated experts according to the complexity of the cue word, the output accuracy is improved, and the hardware cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of computers, and more particularly to a text generation method and an electronic device. Background Art

[0002] Generative large language models are deep learning models that can generate natural language text. Such models are based on complex neural network architectures, such as transformers, and are pre-trained on large amounts of text data to learn the statistical properties and patterns of language. The core capability of generative large language models is to generate coherent, logical, and grammatically correct text. This means that the model can output completely new text content, such as articles, conversations, poems, etc., rather than just copying or restating the content in the training data.

[0003] In large language models, Mixture of Experts (MoE) can be integrated into the transformation layer to reduce the amount of computation by activating a smaller number of experts. However, the number of activated experts in the current solution is fixed, which makes the model inflexible. Summary of the invention

[0004] According to an exemplary embodiment of the present disclosure, a method, apparatus, electronic device, computer-readable storage medium, and computer program product for text generation are provided. The number of experts to be activated in the MoE large model can be determined based on the complexity of the prompt information, thereby enabling dynamic control and dynamic use of the model.

[0005] In a first aspect of the present disclosure, there is provided an information processing method, comprising: obtaining prompt information; inputting the prompt information into a MoE large model, and determining a vector representation corresponding to the prompt information, wherein the MoE large model has a maximum number of experts determined through training; further comprising using a text evaluation model to determine a score value of the prompt information based on the vector representation, wherein the score value of the prompt information is a quantitative performance indicator for characterizing the complexity of the prompt information; determining the number of experts to be activated in the MoE large model based on the score value of the prompt information, wherein the number of experts to be activated is not greater than the maximum number of experts; and generating an output corresponding to the prompt information based on the activated experts of the MoE large model.

[0006] In a second aspect of the present disclosure, an electronic device is provided, comprising: at least one processing unit; and at least one memory, wherein the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, and when the instructions are executed by the at least one processing unit, the electronic device executes the method described in the first aspect of the present disclosure.

[0007] In a third aspect of the present disclosure, a text generation device is provided, including: an acquisition unit configured to acquire prompt information; a first determination unit configured to input the prompt information into a MoE large model, and determine a vector representation corresponding to the prompt information, wherein the MoE large model has a maximum number of experts determined through training; a second determination unit configured to determine a score value of the prompt information based on the vector representation using a text evaluation model, wherein the score value of the prompt information is a quantitative performance indicator for characterizing the complexity of the prompt information; a third determination unit configured to determine the number of experts to be activated in the MoE large model based on the score value of the prompt information, wherein the number of experts to be activated is not greater than the maximum number of experts; and a generation unit configured to generate an output corresponding to the prompt information based on the activated experts of the MoE large model.

[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, which has machine-executable instructions stored thereon, and when the machine-executable instructions are executed by a device, the device performs the method described in the first aspect of the present disclosure.

[0009] In a fifth aspect of the present disclosure, a computer program product is provided, comprising computer executable instructions, wherein the computer executable instructions implement the method described according to the first aspect of the present disclosure when executed by a processor.

[0010] In a sixth aspect of the present disclosure, an electronic device is provided, comprising: a processing circuit configured to execute the method described according to the first aspect of the present disclosure.

[0011] The invention summary is provided to introduce a series of concepts in a simplified form, which will be further described in the specific embodiments below. The invention summary is not intended to identify the key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein: Figure 1 A schematic diagram showing a system according to some embodiments of the present disclosure; Figure 2 A schematic flow chart of a method for text generation according to some embodiments of the present disclosure is shown; Figure 3 A schematic diagram showing a text generation process according to some embodiments of the present disclosure; Figure 4A block diagram showing an example apparatus according to some embodiments of the present disclosure; and Figure 5 A block diagram of an example device that may be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0014] In the embodiments of the present disclosure, the MoE large model can be referred to as a MoE large language model (LLM), a large model based on the MoE architecture, a large language model based on the MoE, a generative large language model, a large language model based on the transformer architecture, a large language model, a large model, etc., and the present disclosure is not limited thereto. The large model in the embodiments of the present disclosure can be applied to a variety of application scenarios, such as writing articles, generating dialogues, creating poems, generating content for social media / news, etc., generating code / annotations / documents, generating teaching materials, assisting language learning, etc.

[0015] Large language models can deeply understand and process natural language, and complete various complex tasks such as text generation and knowledge question answering based on neural networks. Large language models can be based on transformer architectures, which can effectively handle long-distance dependencies through self-attention mechanisms and have excellent performance in parallel computing. Large language models can include multiple transformation layers, for example, a transformation layer can be an encoder layer or a decoder layer. Transformation layers can include self-attention layers, feedforward neural network (FNN) layers, and layer normalization (LN) layers.

[0016] The self-attention layer allows the model to establish associations between words or phrases at different positions and capture long-distance dependencies. The FNN layer further nonlinearly transforms the output of the self-attention layer. The LN layer can normalize between all activations of each sample. For each sample, the LN layer can calculate the mean and standard deviation of all activation values ​​of the sample, and then use these statistics to normalize each activation value to have unit mean and unit variance.

[0017] In practical applications, in order to reduce the cost of large language models during training and reasoning, the MoE architecture is introduced.

[0018] However, the traditional large language model based on MoE has not completely solved the cost problem. MoE can only activate a fixed number of expert networks through the gating function, and cannot dynamically adjust the number of activated experts according to the difficulty of the prompt information. The processing ability for complex prompt information is limited, making it difficult to generate effective responses; and for simple prompt information, it will cause excessive resource consumption.

[0019] In order to solve the above problems and potential other problems, an embodiment of the present disclosure provides a method for text generation. In an embodiment of the present disclosure, a text evaluation model can be used to evaluate the complexity of the prompt information input to the MoE large model, thereby dynamically adjusting the number of experts activated by the MoE large model. The text evaluation model can score the complexity of the vector representation corresponding to the prompt information, and based on the score value, the number of experts to be activated in the MoE large model can be determined.

[0020] In this way, the complexity of the prompt information is scored with the help of the text evaluation model, and the number of experts that need to be activated by the MoE large model is dynamically adjusted. For complex prompt information, the number of experts can be increased to generate high-quality responses; for simple prompt information, the MoE large model can reduce the number of activated experts. In this way, the resource mismatch problem caused by a fixed number of activated experts is avoided, resource utilization is effectively improved, hardware and energy costs are reduced, and users are provided with better quality and more efficient text generation services.

[0021] Figure 1 A schematic diagram of a system 100 according to some embodiments of the present disclosure is shown. The system 100 includes a MoE big model 102 and an evaluation system 104, wherein the evaluation system 104 includes a context manager 106, a text evaluator 108 (or simply an evaluator), and a system dynamic scheduler 110 (or simply a scheduler).

[0022] like Figure 1As shown, the MoE large model 102 and the evaluation system 104 can receive and transmit information to each other and work together to complete related tasks. The MoE large model 102 is a deep learning model architecture for natural language processing, which is composed of multiple sub-models, namely experts. Each expert is essentially a feedforward neural network (FFN) with specific parameter settings and calculation structures. After specific training, specialized learning can be performed for different data features or tasks. For example, some experts are good at processing semantic understanding, while some experts are more proficient in grammatical analysis. Through specialized division of labor, the MoE large model 102 can have stronger generalization capabilities, can face more diverse input information, and reduce the limitations that a single model may have when processing complex tasks. Furthermore, the MoE large model 102 also includes a gating function. The gating function is a special mechanism whose input is the current input data and whose output is a weight vector for each expert, reflecting the importance of each expert in processing the current input data, forming a probability distribution. Based on this probability distribution, the model can decide which experts to activate to process the input data.

[0023] For example, assume that there are N experts, and it can be expressed as: . Exemplarily, the gating function can be expressed as: (1) In formula (1), s(x) represents the scoring function, s i (x) Indicates i Experts f i (x) 's rating.

[0024] Exemplarily, the output of MoE can be expressed as: (2) Sparse MoE is a way to implement MoE. For a MoE with N expert groups, only some of the most important expert models can be activated. For example, for N=16, based on the output of the gating function, K=4 experts can be activated to achieve sparsity and reduce the amount of calculation. This method can be called Top-K-MoE, such as Top-4-MoE. For example, the output of Top-K-MoE can be expressed as: (3) For example, in the MoE large model, the encoding layer or decoding layer may include a self-attention layer, a Top-K-MoE layer, and an LN layer. Or it can be understood that the FFN layer in the large language model is replaced by a MoE layer.

[0025] The configuration of the MoE large model, including the number of layers, vector dimensions, and the MoE parameter K, is fixed before the model training process and cannot be changed after the training is completed. These configurations directly determine the computing cost consumed by the large model, including hardware cost and energy cost.

[0026] In the embodiment of the present disclosure, an evaluation system 104 is introduced. By working in cooperation with the evaluation system 104, the number of activated experts can be dynamically adjusted according to the complexity of the prompt information input into the MoE large model 102.

[0027] In some embodiments of the present disclosure, after receiving the prompt information, the MoE big model 102 can process the prompt information and convert it into a vector representation. It can be understood that the prompt information input by the user is mostly presented in the form of natural language. The prompt information input by the user to the big model can be a piece of text used to guide the model to generate a specific output. For example, the prompt information can be a question, a sentence fragment, an instruction, or any other form of text. The big model can predict the next content based on the prompt information. For example, in a task of generating text, if the prompt information is "Today's weather", the big model may generate words such as "sunny" or "cloudy" to complete the output sentence.

[0028] Exemplarily, the prompt information may include a word sequence, such as a plurality of word tokens.

[0029] The MoE large model 102 can encode the semantics, syntax and other features of the prompt information and convert them into vector representations. Exemplarily, the vector representation output by the intermediate layer of the decoder of the MoE large model 102 can be stored in the context manager 106. For example, the context manager 106 can store the intermediate hierarchical features in the vector buffer. For example, the decoder outputs of the 16th layer and the 32nd layer of the 32-layer decoder of the MoE large model 102 can be stored in the buffer, where the output of the 16th layer and the output of the 32nd layer are both 4096-dimensional vectors.

[0030] Furthermore, the text evaluator 108 can calculate the difficulty and complexity of the prompt information for the MoE large model 102 based on the vector representation of the prompt information, and express the complexity with a score value. Specifically, the text evaluator 108 can process the vector representation through a convolutional neural network (CNN) or a recurrent neural network (RNN).

[0031] Specifically, the text evaluator 108 may use the decoder output of the vector buffer to predict the difficulty and complexity of the prompt word to the MoE large model 102.

[0032] The system dynamic scheduler 110 can be implemented as a configurable scheduling function. In some embodiments of the present disclosure, the developer of the MoE large model 102 can freely define multiple thresholds for the complexity of the prompt information, in order to adapt to different application scenarios and task requirements. Different scenarios have different requirements for the accuracy and efficiency of prompt information processing. As a quantitative standard, the threshold can accurately define the complexity interval and provide a basis for resource allocation. Exemplarily, the system dynamic scheduler 110 can determine the number of experts to be activated based on multiple thresholds. Specifically, based on the score value output by the text evaluator 108, the minimum first-level threshold higher than the score value can be screened out from the preset threshold, so as to clarify the threshold level of the prompt information and achieve an accurate match between complexity and resource allocation.

[0033] Further, based on the result of the system dynamic scheduler 110, the MoE large model 102 activates a corresponding number of experts and generates an output. At the same time, in the embodiment of the present disclosure, the threshold can also be adjusted according to the working state of the server. When the server is highly loaded, the threshold is increased to reduce the number of activated experts and ensure system stability; when the server is under low load, the threshold is lowered to make full use of resources and improve processing efficiency.

[0034] Figure 2 A schematic flow chart of a method 200 for text generation according to some embodiments of the present disclosure is shown. In box 202, prompt information is obtained. In box 204, the prompt information is input into the MoE large model, and a vector representation corresponding to the prompt information is determined, wherein the MoE large model has a maximum number of experts determined through training. In box 206, a score value of the prompt information is determined based on the vector representation using a text evaluation model, wherein the score value of the prompt information is a quantitative performance indicator for characterizing the complexity of the prompt information. In box 208, based on the score value of the prompt information, the number of experts to be activated in the MoE large model is determined, wherein the number of experts to be activated is not greater than the maximum number of experts. In box 210, an output corresponding to the prompt information is generated based on the activated experts of the MoE large model.

[0035] Exemplarily, the vector representation corresponding to the prompt information includes: hierarchical features corresponding to the prompt information output by the middle layer of the MoE large model. As mentioned above, it can be the hierarchical features output by the middle layer of the decoder of the MoE large model, and these hierarchical features can be saved in the vector buffer by the context manager.

[0036] Exemplarily, the text evaluation model can be understood as follows Figure 1 The evaluation system 104 or text evaluator 108 shown in FIG.

[0037] In some embodiments of the present disclosure, a mode in which a text evaluation model and a MoE large model work together is adopted to process prompt information more efficiently and energy-savingly. The text evaluation model can evaluate the complexity of the prompt information for the MoE large model, and present the evaluation results in the form of a score value. It can be understood that the model cannot directly understand the prompt information in natural language, so the MoE large model will pre-process the input prompt information, output the corresponding hierarchical features from the middle layer, and convert the word sequence in the prompt information into a vector representation. Specifically, when the prompt information is input, it will undergo a tokenization process, which will divide the prompt information into a series of word units, and then map it with the vocabulary and convert it into a corresponding vector representation, also called embedding.

[0038] Exemplarily, the text evaluation model may use a quantitative performance indicator to represent the complexity of the prompt information. For example, the obtained score value may be any quantitative performance indicator, such as the accuracy, recall rate, cross entropy, etc. of sequence prediction.

[0039] In some embodiments, the text evaluation model can be based on CNN, which extracts features through convolutional layers and pooling layers and is suitable for discovering local patterns in sequences. For example, the convolution operation of the convolutional layer can be expressed as: (4) In formula (4) f is the input sequence, g is the convolution kernel (or filter), t is the position subscript in the sequence.

[0040] In CNN, each convolution operation may be followed by a nonlinear activation function, for example, represented as: (5) The pooling layer in CNN is used to reduce the dimension of data. For example, the maximum pooling operation can be expressed as: (6) In formula (6) k Represents the size of the pooling window. The output layer in a CNN can include only one neuron and no activation function is used.

[0041] In some embodiments, the text evaluation model can be based on RNN, which processes sequence data through recursive connections and can be suitable for processing context dependencies in sequences. Exemplarily, the hidden state update in RNN can be expressed as: (7) In formula (7), h tIs the first t The hierarchical features of word units, W h and W x is the weight matrix, b h is the bias term, x t is the time step t Input.

[0042] In RNN, the output of the model ( y t ) can be determined based on the hidden state, for example, as: (8) In formula (8), W y represents the output weight matrix, b y Represents the output bias term.

[0043] In an embodiment of the present disclosure, a neural network training system may be used to train a text evaluation model. Exemplarily, the training system may include back propagation, stochastic gradient descent, parallel training, and the like.

[0044] Exemplarily, the loss function can be calculated by the model in the forward propagation. Then the loss function is differentiated with respect to the model parameters to obtain the gradient. The gradient is then back-propagated back to the model to update the model parameters. Exemplarily, in the process of stochastic gradient descent, a suitable learning rate can be selected according to the task and model, and then the model parameters are iteratively updated according to the gradient in the iterative process. Exemplarily, in the parallel training process, the training data can be divided into multiple batches, and each batch can be trained on a GPU. If the model is too large to be completed on a single GPU, the model can be divided into multiple parts to be trained separately on multiple GPUs.

[0045] Optionally, during the model training process, a batch normalization layer can be added to the neural network model to speed up the training. Optionally, during the model training process, a regularization term can be added to the loss function to prevent overfitting.

[0046] Optionally, during the model training process, forward propagation, back propagation, stochastic gradient descent and other processes may be repeated until a preset stop condition is reached. In this way, a trained text evaluation model can be obtained and better adapted to specific tasks.

[0047] In some embodiments, multiple thresholds may be determined, for example, multiple threshold intervals formed by multiple thresholds correspond to multiple numbers of candidate experts in a one-to-one manner, wherein the difference between the number of multiple threshold intervals and the number of multiple thresholds is 1. In some embodiments, in response to a first threshold interval of the score value in the multiple threshold intervals, the number of candidate experts corresponding to the first threshold interval is determined as the number of experts to be activated.

[0048] In some embodiments of the present disclosure, in order to process prompt information more accurately and determine the complexity score of prompt information, multiple thresholds for the complexity of prompt information for large models can be defined based on usage scenarios, customer needs and other factors. The threshold intervals formed by these thresholds correspond one-to-one to the number of multiple candidate experts, where the difference between the number of multiple threshold intervals and the number of multiple thresholds is one. The score value is compared with multiple thresholds, and the smallest first-level threshold higher than the score value is selected. According to the threshold interval where the score value is located, the corresponding number of candidate experts, that is, the number of experts to be activated, is determined. For example, assuming that the multiple thresholds include threshold 1 and threshold 2, and threshold 1 is less than threshold 2, then the multiple threshold intervals are: less than threshold 1, threshold 1 to threshold 2, and greater than threshold 2.

[0049] It can be understood that each of the multiple candidate expert numbers is less than the maximum number of experts, and the maximum number of experts is determined by the MoE large model during the training process. In addition, different usage scenarios and customer needs have different requirements for the model to process prompt information. For example, in a real-time interactive scenario, customers may pay more attention to response speed. At this time, a relatively low threshold can be set to allow the model to activate fewer experts when the prompt information is less complex, so as to speed up the processing speed; in scenarios where the accuracy of the results is extremely high, a higher threshold can be set to ensure that enough experts can participate in the processing of complex prompt information.

[0050] Taking a large MoE model with N=16 experts and a maximum of K=4 experts activated at the same time as an example, 4-1=3 thresholds from small to large can be configured in advance. For example, 0.5, 0.8, 0.9, the corresponding threshold intervals are 4, respectively (0-0.5), (0.5-0.8), (0.8-0.9), (>0.9), and each threshold interval corresponds to a different number of candidate experts (such as 1, 2, 3, 4 respectively). After the text evaluator predicts the complexity and difficulty of the prompt information and gives a score value, the system dynamic scheduler 110 can perform threshold judgment, select the minimum first-level threshold value higher than the score value, and thus determine the threshold interval where the score value is located. For example, for a prompt word with a score value of 0.47, 0.5>0.47, select the first-level threshold 0.5, the score value of this prompt word falls in the threshold interval of (0-0.5), and the number of experts to be activated can be 1 accordingly. For a prompt word with a degree value of 0.87, because 0.8<0.87<0.9, the third-level threshold of 0.9 is selected, and the score value of the prompt word is in the threshold interval of (0.8-0.9), and the number of experts to be activated can be 3. For prompt words that exceed all three-level thresholds, the scheduler selects the fourth-level threshold interval, and the number of experts to be activated can be 4. In this way, the MoE large model can accurately allocate expert resources according to the actual complexity of the prompt information, which not only ensures the processing effect but also avoids the waste of resources.

[0051] In some embodiments of the present disclosure, when determining multiple thresholds of the MoE large model, the working state of the server running the model can be comprehensively considered. The working state of the server changes dynamically, and the threshold can be adjusted dynamically according to the working state of the server to ensure the work efficiency and quality of the MoE large model to generate output according to the prompt information. The working state of the server may include one or more key indicators. The number of tasks processed simultaneously by the server (such as the number of prompt words) can reflect its current load pressure. If too many tasks are processed at the same time, the processing capacity of the server will be challenged, which may lead to a slower response speed or even a freeze. The utilization rate of the graphics processing unit (GPU), the memory occupancy rate of the server, and the running speed of the server are also important indicators. Among them, the GPU plays a key role in the calculation of the large model, and high utilization may mean that the server is in a high-intensity computing state.

[0052] Based on the working status information of the server, multiple thresholds can be determined more scientifically. When the server is in a busy state, for example, more than a certain number (such as 20) of prompt messages are processed at the same time, or the GPU utilization rate exceeds a certain proportion (such as 50%), or the memory occupancy rate is too high and the running speed is too slow, the system dynamic scheduler 110 can determine a higher threshold, such as increasing the previously set thresholds at each level by a predetermined value, such as the original thresholds of 0.5, 0.8, and 0.9 are increased by 0.05 at the same time, becoming 0.55, 0.85, and 0.95. By raising the threshold, each prompt message is responded to by fewer activated experts, thereby reducing the burden on the server, ensuring the stable operation of the system, and maintaining processing efficiency to a certain extent. In this way, the threshold is dynamically adjusted according to the real-time working status of the server, which can achieve the optimal configuration of computing resources and improve the overall performance of the MoE large model.

[0053] In some embodiments of the present disclosure, the system has two working modes, exemplarily referred to as the first working mode (also referred to as mode A) and the second working mode (also referred to as mode B), which can provide more flexible solutions for different application requirements. In the first working mode, after the prompt information is input into the MoE large model, a score value will be output for the prompt information as a whole, the number of experts to be activated will be determined, and the experts activated based on the MoE large model will generate a complete output at one time according to the prompt information. Specifically, the evaluator and the scheduler will select the corresponding k-th level threshold according to the complexity of the prompt information, and then the MoE large model will activate k experts for the overall reply, that is, the input of the generation model. When generating the output, these activated experts will combine their respective professional capabilities, conduct in-depth analysis and processing of the prompt information, and finally generate complete output content. This working mode reply content is complete and coherent, which can meet the user's demand for overall information. By activating an appropriate number of experts, the accuracy and professionalism of the reply can be guaranteed, and the content quality can be improved. It is suitable for scenarios that require complete and comprehensive replies, such as writing articles, reports, and detailed answers to questions. After the user enters the prompt word, the system can give a complete answer covering all the key points at one time. It can be understood that in the first working mode, the number of activated experts is determined once and for all based on the complexity of the prompt information.

[0054] In the second working mode, the text evaluation model can be used to determine the current score value based on the vector representation of the prompt information and the vector representation of multiple word units in the generated output. Further, based on the current score value, the number of experts to be activated in the MoE large model is determined to generate the next word unit of the multiple word units. Specifically, the system determines the corresponding vector representation based on the prompt information and the last generated preset number of word units in the output that has been generated in the current state of the MoE large model (assuming that it is the first t word units in the generated output). The text evaluator can determine a score value based on the current vector representation, and based on the comparison of the score value with the preset threshold, the number of experts to be activated to generate the next word unit (i.e., the t+1th word unit) can be determined. The MoE large model can generate the next word unit based on these activated experts. This process is repeated until the complete text output is generated. For example, the last t words generated can be selected by a fixed-size window, and the number of experts to be activated for generating the next word can be determined based on the vector representation of the words in the window.

[0055] It is understandable that the preset number of word units that have been generated provide important contextual information for predicting the next word unit. This working mode can make dynamic predictions based on the previous content generated in real time, adjust the reply strategy in time, has strong real-time performance, and can dynamically adjust the reply according to the context, providing a more natural and smooth interactive experience. By gradually generating words, it can better adapt to different contexts and user needs, and improve the flexibility and pertinence of replies. It is suitable for scenarios that require real-time interaction, automatic completion or text generation, such as chatbot conversations, intelligent writing assistance, etc. Subsequent content can be dynamically generated based on existing replies to make conversations or text generation smoother. It is understandable that in the second working mode, the number of experts to be activated is re-determined for the next word to be output. That is to say, in the process of determining the complete output, the number of experts to be activated is dynamically updated.

[0056] Figure 3 FIG. 3 is a schematic flow chart of a text generation process 300 according to some embodiments of the present disclosure. Figure 3 As shown, at the start of the entire system operation, the MoE large model can be loaded. For example, the trained MoE large model has determined the maximum number of experts, represented by K, such as K=4 or other values.

[0057] All parameters and functions of the MoE model are loaded into the GPU or CPU of the server. These parameters are the weights, biases and other data learned by the model during the training process, which determine the model's processing mode and ability to process input data; functions cover the program code for the model to perform operations such as calculations and conversions, such as self-attention mechanism functions and feedforward neural network functions. It can be understood that these parameters and functions are the basic configuration for the normal operation of the MoE model. After the MoE model is successfully loaded, the system receives requests through the network portal. These requests contain prompt information entered by the user. The user expects the MoE model to generate a reply text based on these prompt information to answer his own questions. It can be understood that the difficulty of the prompt information is different, and the number of experts to be activated is also different. By introducing the context manager, text evaluator and system dynamic scheduler to work with the MoE model, the number of activated experts can be dynamically adjusted according to the difficulty of the prompt information. For complex prompt information, more experts need to be activated to process it. For simple prompt information, the number of activated experts is reduced to avoid wasting resources.

[0058] Exemplarily, after receiving the request, the MoE large model processes the prompt information and outputs hierarchical features from the intermediate layer. The context manager stores these intermediate features in the vector buffer. The text evaluator calculates the complexity of the prompt information based on the vector representation and represents it with a score value.

[0059] Furthermore, the dynamic scheduler can read the configured multi-level thresholds and adjust the thresholds according to the working status of the server of the current MoE large model. Specifically, the dynamic scheduler will adjust these thresholds according to the hardware conditions such as the GPU utilization rate and memory occupancy rate of the current server and the number of requests. Then, the threshold with the smallest level greater than the score given by the text evaluator is selected.

[0060] like Figure 3As shown in the figure, the system has two working modes. In working mode A, after determining the score value according to the prompt information, the dynamic scheduler determines the number of experts to be activated according to the score value. The activated experts conduct in-depth analysis of the prompt information and generate a complete response to the prompt information at one time based on their respective professional capabilities. This mode is suitable for scenarios with high requirements for response completeness and accuracy, such as document generation and professional question answering. In working mode B, the MoE large model calls the context manager again when determining the next word to reply. The context manager determines the vector representation of multiple words and prompt information in the currently generated output. The text evaluator recalculates the score value, and the dynamic scheduler recalculates the number of experts to be activated to predict the next word. This cycle repeats until a complete response is generated. This mode is suitable for real-time interaction scenarios, such as chatbot dialogues, and can dynamically generate subsequent content based on the previous text, making the response more real-time and natural, and improving the user experience.

[0061] It is understandable that the embodiments of the present disclosure do not limit the training method of the MoE large model. For example, the MoE large model can be obtained through training through back propagation, stochastic gradient descent, etc. Exemplarily, the trained MoE large model is a Top-K-MoE large model with a determined K value, where the K value is the aforementioned maximum number of experts.

[0062] It should be understood that in the embodiments of the present disclosure, "first", "second", "third", etc. are only used to indicate that multiple objects may be different, but at the same time do not exclude that two objects are the same, and should not be interpreted as any limitation on the embodiments of the present disclosure.

[0063] It should also be understood that the methods, situations, categories and divisions of the embodiments in the present disclosure are only for the convenience of description and should not constitute special limitations. The features in various methods, categories, situations and embodiments can be combined with each other when it is logical.

[0064] It should also be understood that the above content is only to help those skilled in the art better understand the embodiments of the present disclosure, rather than to limit the scope of the embodiments of the present disclosure. Those skilled in the art may make various modifications, changes or combinations based on the above content. Such modifications, changes or combinations are also within the scope of the embodiments of the present disclosure.

[0065] It should also be understood that the description of the above content focuses on emphasizing the differences between the various embodiments, and the same or similar points can be referenced or borrowed from each other. For the sake of brevity, they will not be repeated here.

[0066] Figure 44 shows a schematic block diagram of an example device 400 according to some embodiments of the present disclosure. The device 400 may be implemented in software, hardware, or a combination of both. Figure 4 As shown, the apparatus 400 includes an acquiring unit 402 , a first determining unit 404 , a second determining unit 406 , a third determining unit 408 and a generating unit 410 .

[0067] The acquisition unit 402 is configured to acquire prompt information. The first determination unit 404 is configured to input the prompt information into the MoE large model, and determine the vector representation corresponding to the prompt information, wherein the MoE large model has a maximum number of experts determined through training. The second determination unit 406 is configured to use a text evaluation model to determine a score value of the prompt information based on the vector representation, wherein the score value of the prompt information is a quantitative performance indicator for characterizing the complexity of the prompt information. The third determination unit 408 is configured to determine the number of experts to be activated in the MoE large model based on the score value of the prompt information, wherein the number of experts to be activated is not greater than the maximum number of experts. The generation unit 410 is configured to generate an output corresponding to the prompt information based on the activated experts of the MoE large model.

[0068] In some embodiments, the third determination unit 408 is configured to: determine multiple thresholds, multiple threshold intervals consisting of the multiple thresholds correspond one-to-one to multiple numbers of candidate experts, wherein the difference between the number of the multiple threshold intervals and the number of the multiple thresholds is one; and in response to a first threshold interval of the score value in the multiple threshold intervals, determine the number of candidate experts corresponding to the first threshold interval as the number of experts to be activated.

[0069] In some embodiments, the third determining unit 408 is configured to: obtain the working status of the server running the MoE large model; and determine a plurality of thresholds based on the working status of the server.

[0070] In some embodiments, the working status of the server includes at least one of the following: the number of tasks processed simultaneously by the server, the utilization rate of a GPU of the server, the memory occupancy rate of the server, or the running speed of the server.

[0071] In some embodiments, each of the plurality of candidate numbers of experts is less than the maximum number of experts.

[0072] In some embodiments, the vector representation corresponding to the prompt information includes: hierarchical features corresponding to the prompt information output by the middle layer of the MoE large model.

[0073] In some embodiments, the generation unit 410 may be configured to generate a complete output based on the activated experts of the MoE large model.

[0074] In some embodiments, the generating unit 410 is configured to generate a next word-gram of the plurality of word-grams based on the activated experts of the MoE large model and the plurality of word-grams in the generated output.

[0075] In some embodiments, the second determining unit 406 is further configured to: determine the score value based on the vector representation of the prompt information and the vector representation of the multiple word units in the generated output using the text evaluation model. Optionally, the multiple word units are the last preset number of word units generated in the generated output.

[0076] In some embodiments, the text evaluation model adopts a CNN or RNN structure.

[0077] Figure 4 The device 400 can be used to achieve the above combination Figures 1 to 3 For the sake of brevity, the process will not be described in detail here.

[0078] The division of modules or units in the embodiments of the present disclosure is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in the disclosed embodiments may be integrated into one unit, or may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0079] Figure 5 1 shows a block diagram of an example device 500 that can be used to implement embodiments of the present disclosure. It should be understood that Figure 5 The device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. Figures 1 to 3 The process described.

[0080] like Figure 5 As shown, device 500 is in the form of a general-purpose computing device. Components of computing device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be an actual or virtual processor and is capable of performing various processes according to a program stored in memory 520. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to increase the parallel processing capabilities of computing device 500.

[0081] The computing device 500 typically includes a plurality of computer storage media. Such media may be any available media accessible to the computing device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 may be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 may be a removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which may be capable of being used to store information and / or data (e.g., training data for training) and may be accessed within the computing device 500.

[0082] The computing device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 5 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules that are configured to perform various methods or actions of various implementations of the present disclosure.

[0083] The communication unit 540 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Therefore, the computing device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0084] Input device 550 may be one or more input devices, such as a mouse, keyboard, tracking ball, etc. Output device 560 may be one or more output devices, such as a display, a speaker, a printer, etc. Computing device 500 may also communicate with one or more external devices (not shown) through communication unit 540 as needed, such as storage devices, display devices, etc., communicate with one or more devices that allow users to interact with computing device 500, or communicate with any device that allows computing device 500 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0085] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0086] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0087] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0088] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0089] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0090] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A text generation method, comprising: Get prompt information; Inputting the prompt information into a mixture of experts (MoE) large model to determine a vector representation corresponding to the prompt information, wherein the MoE large model has a maximum number of experts determined through training; Determining a score value of the prompt information based on the vector representation using a text evaluation model, wherein the score value of the prompt information is a quantitative performance indicator for characterizing the complexity of the prompt information; Determining the number of experts to be activated in the MoE large model based on the score value of the prompt information, wherein the number of experts to be activated is not greater than the maximum number of experts; as well as Based on the activated expert of the MoE large model, an output corresponding to the prompt information is generated.

2. The method of claim 1, wherein determining the number of experts to be activated in the MoE large model based on the score value comprises: Determine a plurality of thresholds, wherein a plurality of threshold intervals formed by the plurality of thresholds correspond one-to-one to a plurality of numbers of candidate experts, wherein a difference between the number of the plurality of threshold intervals and the number of the plurality of thresholds is one; and In response to the score value being in a first threshold interval among the plurality of threshold intervals, the number of candidate experts corresponding to the first threshold interval is determined as the number of experts to be activated.

3. The method of claim 2, wherein determining the plurality of thresholds comprises: Obtaining the working status of the server running the MoE large model; as well as The plurality of thresholds are determined based on the working status of the server.

4. The method according to claim 3, wherein the working status of the server comprises at least one of the following: The number of tasks processed simultaneously by the server, the utilization rate of a graphics processing unit (GPU) of the server, the memory occupancy rate of the server, or the operating speed of the server. The method according to claim 2 , wherein each of the plurality of candidate numbers of experts is smaller than the maximum number of experts.

6. The method according to claim 1, wherein the vector representation corresponding to the prompt information comprises: The hierarchical features corresponding to the prompt information output by the middle layer of the MoE large model.

7. The method of claim 1 , wherein generating the output comprises: Based on the activated experts of the MoE large model, a complete output is generated.

8. The method of claim 1 , wherein generating the output comprises: Based on the activated expert of the MoE large model and a plurality of word-grams in the generated output, a next word-gram of the plurality of word-grams is generated.

9. The method of claim 8, wherein determining the score value comprises: The text evaluation model is used to determine the score value based on the vector representation of the prompt information and the vector representation of a plurality of word units in the generated output.

10. The method according to claim 8, wherein the plurality of word tokens are a preset number of word tokens generated last in the generated output.

11. The method according to claim 1, wherein the text evaluation model adopts a convolutional neural network (CNN) or a recurrent neural network (RNN) structure.

12. An electronic device comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 11.

13. A text generation device, comprising: An acquisition unit, configured to acquire prompt information; A first determining unit is configured to input the prompt information into a mixture of experts MoE large model to determine a vector representation corresponding to the prompt information, wherein the MoE large model has a maximum number of experts determined through training; A second determining unit is configured to determine a score value of the prompt information based on the vector representation using a text evaluation model, wherein the score value of the prompt information is a quantitative performance indicator for characterizing the complexity of the prompt information; a third determining unit configured to determine the number of experts to be activated in the MoE large model based on the score value of the prompt information, wherein the number of experts to be activated is not greater than the maximum number of experts; as well as A generating unit is configured to generate an output corresponding to the prompt information based on the activated expert of the MoE large model.

14. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Unified anomaly detection model and detection method based on hybrid expert system

    CN117911328A

  • Prompt generation method and device, electronic equipment and computer readable storage medium

    CN119226482A

  • Dynamic efficient routing method and device oriented to hybrid expert large model

    CN119514638A

  • Data analysis method and device, computer equipment, readable storage medium and program product

    CN119670742A

  • Mixture of experts models with sparsified weights

    US20230316042A1