Temperature sampling method and electronic device

By classifying and extracting features from the input sequences of a pre-trained language model, and combining a temperature decision-maker and a pre-defined feature table, the problem of temperature parameter adjustment in mixed task instructions for large-scale language models is solved, achieving a dynamic balance between text diversity and accuracy, and reducing computational complexity and deployment costs.

CN121597834BActive Publication Date: 2026-05-19INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing large-scale language models cannot dynamically adjust temperature parameters when processing mixed task instructions, making it difficult to balance text diversity and accuracy. Dynamic temperature sampling technology suffers from high computational complexity, limited temperature control accuracy, complex structure, and high training and deployment costs.

Method used

By classifying the input sequences of a pre-trained language model, extracting local feature data and text classification results, and adjusting temperature parameters using a temperature decision maker and a pre-defined feature table matching method, the calculation is simplified and the control accuracy and interpretability are improved.

Benefits of technology

It achieves a dynamic balance between text diversity and accuracy under different task requirements, reduces computational complexity and training deployment costs, and improves the accuracy and interpretability of temperature control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597834B_ABST
    Figure CN121597834B_ABST
Patent Text Reader

Abstract

The application discloses a temperature sampling method and electronic equipment, and relates to the field of natural language generation, which comprises the following steps: classifying an input sequence of a pre-trained language model to obtain a text classification result, so as to accurately identify the task type corresponding to the current input and avoid the defect that a fixed temperature cannot adapt to multiple types of tasks; when the language model is inferred, local feature data related to the output quality of the pre-trained language model is extracted from the probability distribution of the current time step, so as to reduce the calculation complexity; the temperature decision maker is used to find a preset feature table to determine the temperature parameter, so that the operation is simplified, the additional calculation overhead is reduced, the searching logic of the preset feature table is clear, the temperature decision process has good interpretability, and finally the probability distribution of the current time step is adjusted based on the determined temperature parameter, so that the dynamic balance between the text diversity and accuracy under different task requirements is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language generation, and in particular to a temperature sampling method and electronic device. Background Technology

[0002] The development of artificial intelligence is driving the pursuit of higher quality in natural language generation, requiring a balance between syntax, semantics, diversity, and factual consistency, and its widespread application across multiple fields. Temperature is a key parameter for controlling output diversity: low temperatures adapt to factual tasks, while high temperatures meet creative needs. However, mainstream large-scale language models, when handling mixed task instructions, use fixed temperature values ​​for generation, which cannot be dynamically adjusted, making it difficult to balance text diversity and accuracy, thus becoming a bottleneck in technological development.

[0003] While dynamic temperature sampling technologies attempt to overcome bottlenecks, each has its limitations. Entropy-based Dynamic Temperature Sampling (EDT) adjusts temperature linearly using text entropy values, but the computational complexity of the vocabulary probability distribution impacts real-time performance; furthermore, entropy, as a global quantity, struggles to capture distribution details, limiting temperature control accuracy. Adaptive decoding methods based on Latent Preference Optimization (LPO), which rely on additional neural networks to predict temperature, suffer from complex structures, high training and deployment costs, and a lack of interpretability in their temperature control logic, hindering technology implementation and optimization. Summary of the Invention

[0004] This application provides a temperature sampling method and electronic device to at least solve the problems of high computational complexity, limited temperature control accuracy, complex structure, high training and deployment costs, and poor interpretability in related dynamic temperature sampling technologies.

[0005] This application provides a temperature sampling method, comprising: classifying the input sequence of a pre-trained language model to obtain a text classification result; extracting local feature data from the probability distribution of the current time step of the pre-trained language model when the pre-trained language model performs inference based on the input sequence, wherein the local feature data is related to the output quality of the pre-trained language model; inputting the local feature data and the text classification result into a temperature decision unit to obtain a temperature parameter; wherein the temperature parameter is a preset temperature value in a preset feature table that matches the local feature data and the text classification result; and adjusting the probability distribution of the current time step based on the temperature parameter.

[0006] This application also provides a temperature sampling device, including:

[0007] The classification module is used to classify the input sequences of the pre-trained language model and obtain the text classification results;

[0008] The feature extraction module is used to extract local feature data from the probability distribution of the current time step of the pre-trained language model when the pre-trained language model is performing inference based on the input sequence. The local feature data is related to the output quality of the pre-trained language model.

[0009] The temperature decision module is used to input local feature data and text classification results into the temperature decision unit to obtain temperature parameters; the temperature parameters are preset temperature values ​​in the preset feature table that match the local feature data and text classification results.

[0010] The adjustment module is used to adjust the probability distribution of the current time step based on the temperature parameter.

[0011] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the temperature sampling method described above.

[0012] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described temperature sampling method.

[0013] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described temperature sampling method.

[0014] This application first classifies the input sequences of a pre-trained language model to obtain text classification results, accurately identifying the task type corresponding to the current input sequence. This provides a task-level basis for subsequent targeted adjustments to temperature parameters, avoiding the limitation of fixed temperatures being unsuitable for multiple task types. During inference in the pre-trained language model, local feature data related to the output quality of the pre-trained language model is extracted from the probability distribution of the current time step. This eliminates the need to calculate the probability distribution of the entire vocabulary, reducing computational complexity. Simultaneously, local feature data captures the internal details of the probability distribution, improving temperature control accuracy. A temperature decision-maker is used to determine temperature parameters by searching a pre-defined feature table, simplifying operations and reducing additional computational overhead. The search logic of the pre-defined feature table is clear, making the temperature decision-making process highly interpretable and facilitating technology implementation and subsequent optimization. Finally, by adjusting the probability distribution of the current time step based on the determined temperature parameters, a dynamic balance between text diversity and accuracy under different task requirements is achieved. Attached Figure Description

[0015] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A schematic diagram of the specific hardware architecture on which the execution of a temperature sampling method provided in this application depends;

[0017] Figure 2 A schematic flowchart of a temperature sampling method provided in an embodiment of this application;

[0018] Figure 3 A schematic flowchart of another temperature sampling method provided in an embodiment of this application;

[0019] Figure 4 This is a schematic diagram of the structure of a temperature sampling device provided in an embodiment of this application;

[0020] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0022] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0023] To more clearly illustrate the embodiments of this application, the technical terms used in the embodiments will be briefly introduced below:

[0024] Large Language Models (LLMs) are a class of deep learning models based on the Transformer architecture. They have billions to trillions of parameters, are self-supervised pre-trained on massive amounts of unlabeled text data (such as predicting the next word in a text), and possess deep understanding of human natural language, high-quality generation of text that conforms to semantic logic, and also have the ability to reason, store and transfer knowledge.

[0025] Transformer is a neural network architecture built entirely on a multi-head self-attention mechanism. It is designed specifically for processing sequence data. Through an encoder-decoder structure, combined with modules such as word embedding, position encoding, and feedforward networks, it achieves efficient capture of global dependencies in sequences and supports a deep learning infrastructure for parallel computing.

[0026] A token is the smallest semantic unit obtained by an artificial intelligence model after preprocessing natural language text. It is the basic data unit for the model to understand and generate text.

[0027] In the context of large language models, when a language model generates the next token, it outputs a probability distribution across the entire vocabulary. The entropy of this distribution directly measures the model's uncertainty about the next word.

[0028] The SiLU activation function (Sigmoid Linear Unit, SiLU) is a non-linear activation function that combines the Sigmoid function with a linear transformation. Its core function is to introduce non-linear feature mapping into neural networks, thereby enhancing the model's ability to fit complex data.

[0029] A multilayer perceptron (MLP) is a feedforward neural network consisting of an input layer, one or more hidden layers (fully connected layers), and an output layer. The neurons in each layer are connected by weights to transmit signals. By introducing nonlinear transformations through activation functions, it can learn complex nonlinear mapping relationships of input data and is one of the most basic and general neural network models in deep learning.

[0030] Tongyi Qianwen-7B (Qwen-7B) is an open-source large language model with 7 billion parameters based on the Transformer architecture. It is the basic version of the Tongyi Qianwen (Qwen) large model series. It is trained on a massive pre-training dataset (covering web text, books, code, etc., with a scale of over 2.2 trillion tokens) and has the capabilities of multilingual understanding and generation, code processing, etc.

[0031] Knowledge distillation (KD) is a model compression and knowledge transfer technique that uses a "teacher-student" paradigm to transfer the knowledge (including output probability distribution, intermediate layer features, etc.) contained in a complex, high-performance teacher model to a lightweight, efficient student model. This allows the student model to retain the performance and generalization ability of the teacher model as much as possible while significantly reducing the number of parameters and computational costs. The core of this technique is to use soft-objective-assisted training.

[0032] The loss function built on KL divergence loss (KL Loss) is used to quantify the degree of difference between two probability distributions. Its value is non-negative, and the smaller the value, the closer the two distributions are. Its core function is to guide the model to learn the characteristics of the target probability distribution.

[0033] Prompts are structured textual information used to guide AI models to understand task intent and output target results. They are the core medium for user-model interaction and can be used in the form of instructions, dialogue context, task examples, constraints, etc., enabling models to complete specific tasks such as text generation, logical reasoning, and information extraction in zero-sample, few-sample, or full-sample scenarios.

[0034] Logits refer to the original output vector of an artificial intelligence model before it is normalized by the Softmax or Sigmoid function. Each element of the vector represents the log probability that a sample belongs to a certain class. Mathematically, Logits = log(p / (1-p)), where p is the normalized probability.

[0035] Order of V (O(V)) is used to indicate that the time or space complexity of an algorithm is linearly related to the size of the variable V.

[0036] The Softmax function is a non-linear activation function used to transform the unnormalized raw vector (such as logits) output by a model into a normalized probability distribution. Its core function is to use exponential transformation and normalization to ensure that each element in the output vector corresponds to the probability of a sample belonging to a certain class, and that the sum of the probabilities of all classes is 1.

[0037] A probability distribution is a set of probabilities that a model has for the next possible token. These probabilities have the property that they sum to 1 and are used to quantify the confidence level of the model in predicting different tokens.

[0038] With the rapid development of artificial intelligence technology, the requirements for output quality in natural language generation tasks are increasing. Not only must grammatical correctness and semantic fluency be guaranteed, but diversity and factual consistency must also be achieved. Its applications have broadly covered multiple fields such as machine translation, text summarization, and dialogue systems. In the language model decoding process, the temperature parameter is a core hyperparameter controlling output diversity: lower temperatures (close to 0) make the output more deterministic, suitable for tasks with high requirements for factual accuracy; higher temperatures (close to 1 or higher) can enhance the randomness and creativity of the output, meeting the needs of creative text generation.

[0039] However, current mainstream large-scale language models generally use a single fixed temperature value throughout the generation process of all examples and all tokens when dealing with general instructions that include multiple task types (such as those involving both factual and creative needs). This makes it difficult to flexibly adjust the temperature parameter according to the dynamic changes in task requirements, resulting in a difficulty in achieving a dynamic balance between text diversity and accuracy. This has become a key bottleneck restricting further improvement in the quality of natural language generation.

[0040] To overcome these bottlenecks, dynamic temperature sampling technology has emerged, with typical solutions including entropy-based dynamic temperature sampling (EDT) and latent preference-optimized adaptive decoding methods. However, both methods have significant drawbacks. For the EDT method, its core mechanism is to determine the degree of uncertainty based on the entropy value of the generated text and dynamically adjust the temperature using a linear function: increasing the temperature to enhance diversity when the entropy value is low, and decreasing the temperature to ensure accuracy when the entropy value is high. However, this method requires logarithmic operations and summation of the probability distribution of the vocabulary (typically covering tens to hundreds of thousands of dimensions) during entropy calculation, resulting in a computational complexity of O(V) (where V is the vocabulary size). This leads to significant performance degradation in low-latency, high-concurrency scenarios of real-time decoding. Furthermore, entropy, as a global statistic, cannot capture the detailed information within the probability distribution, easily losing key local features affecting temperature adjustment, thus limiting the accuracy of temperature control.

[0041] For the adaptive decoding method based on Latent Preference Optimization (LPO), temperature is predicted by adding a neural network. This module adopts a 3-layer MLP structure (2048 hidden dimensions), SiLU activation function, and LPO training method. It needs to generate multiple candidate temperature values ​​first, and then combine the hidden state information to select the temperature and adjust the probability distribution. However, this scheme has the problems of complex structure and high training and deployment costs. The high-dimensional design of the MLP hidden layers and the additional preference optimization training process increase the model training cost and deployment complexity; moreover, the temperature control logic of the adaptive module lacks clear interpretability, making it difficult to trace the decision basis for temperature adjustment, which is not conducive to the implementation of the technology and subsequent optimization.

[0042] It is evident that the aforementioned dynamic temperature sampling technologies either suffer from complex real-time performance due to the complexity of vocabulary probability distribution calculations, or have control accuracy limited by the global characteristics of entropy values, or suffer from complex structures, high training and deployment costs, and a lack of interpretability in temperature control logic due to the addition of neural networks. As a result, it is difficult to balance real-time performance, control accuracy, deployment feasibility, and interpretability.

[0043] This application addresses the issues of complex vocabulary probability distribution calculation and insufficient real-time performance in the EDT method. It abandons the complex real-time calculation of global quantities such as entropy and instead adopts a pre-set feature table matching mechanism. The temperature decision-maker does not need to perform real-time statistics of the global probability distribution or complex calculations. It can quickly determine the temperature parameters through the mapping relationship between local feature data, text classification results and pre-set temperature values. The core operation is simplified to feature extraction and table lookup, which reduces additional computational overhead and improves the adaptability to real-time scenarios.

[0044] To address the shortcomings of entropy as a global quantity in capturing distribution details and limiting the precision of temperature control, this application, on the one hand, extracts local feature data directly related to output quality from the probability distribution of the current time step of the pre-trained language model, accurately capturing the probability distribution details of the current generation step and avoiding the loss of details caused by global quantities; on the other hand, it combines the text classification results of the input sequence to make the temperature parameter adapt to the characteristics of both the global text scene and the local generation step, achieving more precise temperature control and making the correlation between temperature adjustment and model output quality stronger.

[0045] To address the issues of complex structure and high training and deployment costs associated with LPO methods, this application eliminates the need for additional neural networks. The temperature decision-maker achieves temperature matching based on a pre-defined feature table, which can be calibrated offline with a small amount of validation data. This eliminates the need for large-scale training and does not increase the structural redundancy of the original model. During deployment, only the pre-defined feature table needs to be loaded, reducing the resource consumption for training and deployment and improving the feasibility of technology implementation.

[0046] Meanwhile, to address the issue of the lack of interpretability in temperature control logic, the temperature parameter generation in this application relies on an explicit feature-temperature mapping relationship. The construction logic of the preset feature table is traceable, rather than a black-box reasoning, making the temperature control logic transparent and interpretable, which facilitates subsequent technical optimization and debugging.

[0047] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0048] The specific application environment architecture or specific hardware architecture on which the temperature sampling method depends is described here.

[0049] like Figure 1 The diagram shows the specific hardware architecture upon which the temperature sampling method relies. The specific architecture and the function of each component are as follows:

[0050] The core component can use a graphics processing unit (GPU) to handle the main inference process of the pre-trained language model. This type of hardware has high parallel computing capabilities, which can support the massive parameter calculation of the pre-trained language model, while providing a low-latency tensor calculation interface to meet real-time inference requirements.

[0051] The auxiliary computing unit can be a general-purpose central processing unit (CPU) responsible for executing the entire process of lightweight signal processing and temperature decision-making: including extracting local feature data such as Top-2 Margin, Top-5 Mean, and Top-1 from the probability distribution output by the pre-trained language model; driving the inference of the text classifier; performing discretization and table lookup operations of the temperature decision-maker; and finally temperature scaling, Softmax normalization, and sampling generation.

[0052] In addition, the hardware architecture needs to be equipped with high-speed memory and storage media. Memory is used to cache the intermediate results (Logits, hidden layer states), probability distribution data and preset feature tables of the pre-trained language model, avoiding increased latency caused by frequent reading and writing of storage media; storage media is used to store the weights of the pre-trained LLM model, text classifier parameters and feature tables, ensuring fast loading at startup.

[0053] The high-load computation of the pre-trained language model is handled by the graphics processing unit (GPU) acceleration hardware, while the lightweight temperature sampling process is completed by a general-purpose CPU. This eliminates the need for additional complex accelerator cards, reducing hardware costs and ensuring low-latency linkage between all stages.

[0054] Embodiments of this application provide a temperature sampling method, such as... Figure 2 As shown, the method includes the following steps:

[0055] S201. Classify the input sequence of the pre-trained language model to obtain the text classification result.

[0056] In some embodiments, a pre-trained text classifier is used to classify the input sequences of the pre-trained language model, and the category corresponding to the dimension with the highest probability value is selected as the text classification result. The text classification results include: factual text, creative text, and mixed-type text.

[0057] The pre-trained text classifier is constructed based on a feature extraction model with frozen parameters (such as Qwen 7B) and a single-layer linear classification head. The single-layer linear classification head is obtained through supervised training based on training data and a preset loss function. The training data generates soft labels in the form of a three-dimensional probability distribution through a teacher model.

[0058] The preset loss function can be KL divergence loss based on knowledge distillation. A soft label in the form of a three-dimensional probability distribution is generated for each prompt word using a teacher model.

[0059] For example, suppose the prompt for a training data item is "explaining the basic principles of quantum entanglement", and the corresponding soft label content is similar to [0.85, 0.05, 0.10]. This means that the probability of this training data item being factual text is 85%, the probability of being mixed text is 5%, and the probability of being creative text is 10%.

[0060] The pre-trained text classifier achieves efficient training and fast inference while ensuring classification accuracy through a minimal single-layer network structure, effectively balancing performance and computational cost.

[0061] The training process of the aforementioned single-layer linear classifier head includes: first, constructing an instruction sample library, which includes factual text, creative text, and mixed-type text; and generating soft labels for each sample through a teacher model. Then, based on the pre-trained loss function, calculating the loss value between the probability distribution output by the single-layer linear classifier head based on the instruction sample library and the soft labels, and then updating the parameters of the single-layer linear classifier head according to the loss value until the loss value converges, thus completing the training of the single-layer linear classifier head.

[0062] In the specific implementation of the above embodiments, the architecture of the text classifier is first designed, which can adopt a minimal structure of a pre-trained feature extraction model and a single-layer linear classification head. The basic model is the Qwen 7B large language model, whose parameters are completely frozen and used only as a feature extractor to avoid resource consumption caused by additional training. Utilizing the general semantic understanding capability of Qwen 7B learned from massive amounts of text, deep semantic features of the user input text are extracted, and the output is a hidden layer feature vector with a dimension of 4096 (corresponding to the hidden layer dimension of Qwen 7B). The classification head is designed as a single-layer linear network with an input dimension consistent with the hidden layer dimension of Qwen 7B (4096) and an output dimension of 3 (corresponding to the three categories of factual, creative, and mixed types, respectively). The network structure has no redundancy, and the computational complexity during inference is only O(H*C) (H=4096, C=3), which is suitable for real-time classification requirements.

[0063] The second step is the preparation and annotation of training data. A diverse sample library of 10,000-50,000 instructions can be constructed, covering three types of text scenarios: factual texts, which include objective knowledge questions and answers, information statements, etc.; creative texts, which cover story creation, creative descriptions, etc.; and mixed-type texts, which combine features of both. The annotation process uses a teacher model to generate soft labels, utilizing Qwen 7B as the teacher model to output soft labels in the form of a three-dimensional probability distribution for each sample. Compared to traditional hard labels (such as [1,0,0]), soft labels can more delicately capture the fuzzy boundaries of text classification results, improving the classification accuracy of mixed-type texts.

[0064] Next, the text classifier training and optimization are performed. The training objective uses the KL divergence loss function based on knowledge distillation, with the soft labels generated by the teacher model serving as the supervision signal to guide the single-layer linear classifier head to learn the discriminative features of the text classification results. KL divergence measures the difference between the output probability distribution of the single-layer linear classifier head and the soft label distribution of the teacher model. By minimizing this difference, the lightweight single-layer linear classifier head can quickly learn the classification logic of complex models. During training, only the parameters of the single-layer linear classifier head are updated, while the parameters of the Qwen 7B feature extraction layer are frozen to reduce the number of training parameters and shorten the training cycle. At the same time, an Adaptive Moment Estimation (Adam) optimizer is used, with appropriate learning rate and batch size set. The classification accuracy is monitored on the validation set (which can be divided into 1:9 ratios from the instruction sample library). Training is stopped when the accuracy does not improve for several consecutive rounds to avoid overfitting.

[0065] Finally, the classification execution occurs during the inference phase. Upon receiving user input text, the text is preprocessed, such as unifying the text encoding format and truncating excessively long texts to the maximum length supported by Qwen 7B. Then, the Qwen 7B with frozen parameters is input to extract the 4096-dimensional hidden layer feature vector from its last layer output. This feature vector is then input into a trained single-layer linear classification head to obtain a three-dimensional probability distribution. Based on the maximum likelihood principle, the category corresponding to the dimension with the highest probability value is selected as the final classification result. If the highest probability corresponds to the first dimension, it is classified as factual text; the second dimension as creative text; and the third dimension as mixed-type text. The entire inference process can be completed in milliseconds, meeting the real-time requirements of dynamic temperature sampling.

[0066] In the above embodiments, the text classifier is based on the Qwen 7B model (parameters are completely frozen, only feature extraction is performed). The classification layer adopts a single-layer linear structure without complex deep networks, which greatly simplifies the model structure and reduces computational and training costs. For training, KL divergence loss based on knowledge distillation is used. A teacher model generates three-dimensional probability distribution soft labels for diverse instruction samples, which can fully utilize the knowledge of the teacher model and improve the classifier's ability to distinguish text classification results. At the same time, the massive and comprehensive training samples also ensure classification accuracy.

[0067] S202. When the pre-trained language model performs inference based on the input sequence, extract local feature data from the probability distribution of the current time step of the pre-trained language model.

[0068] Local feature data is used to reflect the real-time state of the pre-trained language model output and is strongly correlated with the generation quality. Local feature data includes at least one of the following: optimal probability value (Top1), word probability difference (e.g., Top-2 Margin), and word probability mean (e.g., Top-5 Mean).

[0069] The optimal probability value is the highest probability output by the pre-trained language model, reflecting the confidence of the pre-trained language model in the currently generated word; the lexical probability difference is the difference between the probabilities of the pre-trained language model outputs ranking in the top M positions, used to quantify the deterministic differences in the outputs of the pre-trained language model; the mean lexical probability is the average of the probabilities of the pre-trained language model outputs ranking in the top N positions, used to quantify the diversity of the outputs of the pre-trained language model; M is less than N. For example, M=2, N=5.

[0070] In some embodiments, when the pre-trained language model performs inference based on the input sequence, a heap sort algorithm is used to sort the probability distribution at the current time step and select the top N probability values. The optimal probability value ranked first is determined from the top N probability values; the word probability difference is calculated by subtracting the second-ranked probability value from the optimal probability value ranked first; and the arithmetic mean of the top N probability values ​​is calculated to determine the word probability mean.

[0071] In the process of extracting local feature data, firstly, after the pre-trained language model outputs logits, it is converted into a normalized token probability distribution (satisfying the property that the sum of probabilities is 1) using the Softmax function. This is the basic data source for feature extraction. To reduce computational overhead, a heap sort algorithm is used to sort all tokens in the probability distribution in descending order of their probability values, selecting the top N tokens (Top-N). This eliminates the need to traverse the entire vocabulary, simplifying the sorting complexity from O(VlogV) to O(V+KlogK) (where V is the vocabulary size and K=5), thus adapting to real-time inference requirements.

[0072] Secondly, for the selected Top-N tokens, local feature data strongly correlated with generation quality are calculated: First, the Top1 probability value (p1) is directly extracted from the probability of the token ranked first after sorting, without additional calculation, intuitively reflecting the model's confidence in the current generated word; second, the Top-2 Margin is obtained by subtracting the probability value of the second-ranked token (p1-p2) from the Top1 probability value, quantifying the difference in certainty between the pre-trained language model and the optimal token; third, the Top-5 Mean is calculated as the arithmetic mean of the probabilities of the top 5 tokens, reflecting the concentration of high-probability candidate tokens. The calculation of the three features revolves around local high-probability tokens, with a complexity of only O(5), avoiding the high computational cost of global statistics.

[0073] The entire extraction process is seamlessly integrated with the language model inference process, and is completed synchronously within each token generation cycle without adding any inference delay. Furthermore, the features are directly related to the confidence, determinism, and diversity of the model generation, providing efficient and reliable real-time data support for the subsequent accurate temperature adaptation.

[0074] The extraction of local feature data in the above embodiments enables a multi-dimensional and accurate characterization of the generated state of the pre-trained language model; the synergistic effect of local feature data can avoid judgment bias caused by a single feature, making the control of model generation quality more targeted.

[0075] S203. Input the local feature data and text classification results into the temperature decision-maker to obtain the temperature parameters.

[0076] The temperature decision-maker discretizes continuous spatial values ​​(local feature data and text classification results), then uses the discretized results to look up a preset feature table to determine the corresponding temperature parameter. The preset feature table stores the mapping relationship between all feature combinations and temperature; its horizontal dimension represents the discretized local feature data, and its vertical dimension represents the text classification results. The temperature parameter is a preset temperature value in the preset feature table that matches the local feature data and the text classification results.

[0077] For example, when the Top-2 Margin, Top-5 Mean, and Top-1 probability values ​​are high, the pre-trained language model is confident and uses a lower temperature; when only the Margin is low, the pre-trained language model hesitates and appropriately increases the temperature; when all parameters are low and the text is factual, the temperature is slightly increased and controlled within a low range.

[0078] In some embodiments, the temperature decision-maker first discretizes the local feature data, converting it into discrete labels. Then, it retrieves unit data that matches the text classification result from a preset feature table, and then retrieves the temperature parameter corresponding to the discrete label from the unit data.

[0079] Local feature data, originally continuous values, are transformed into discrete labels to compress the decision space. During decision-making, text classification results are first matched to lock in the unit data, and then discrete labels are matched to obtain the temperature, eliminating the need for additional computation and reducing the latency of temperature decisions. Discretization and lookup table mechanisms ensure that temperature decisions are fully interpretable, reducing optimization costs. Text classification results are the core identifier of the scene, prioritizing matching to corresponding unit data to ensure that the basic direction of temperature parameters aligns with scene requirements. Discrete labels correspond to the real-time state generated by the model, further refining temperature selection within the unit data, allowing temperature adjustments to meet both the macroscopic needs of the scene and the microscopic states of the generation process.

[0080] In the process of converting local feature data into discrete labels, discrete labels are divided according to confidence level for the optimal probability value; discrete labels are divided according to deterministic difference for the word probability difference; and discrete labels are divided according to diversity for the word probability mean.

[0081] For the optimal probability value, dividing according to confidence level can transform continuous probability values ​​into intuitive deterministic labels, avoiding interference from irrelevant dimensions; for the word probability difference, dividing according to deterministic difference can accurately quantify the model's selection tendency between the optimal and second-best candidate words; for the word probability mean, dividing according to diversity can be associated with the subsequent temperature adjustment direction. This design allows discrete labels to get rid of the ambiguity of generalized classification and become a precise carrier for conveying the generated state of the pre-trained language model.

[0082] Specifically, for local feature data, it is necessary to formulate refined discretization intervals by combining the model generation rules and the requirements of the classification scenario:

[0083] For the Top1 probability value (ranging from 0 to 1), it can be divided into three intervals: "high", "medium", and "low" based on the confidence level of the pre-trained language model. A high value indicates that the pre-trained language model is highly certain about the generation of the current token, while a low value indicates that the probability is highly uncertain.

[0084] For the Top-2 Margin (range 0-1), it can be divided into three intervals: "large", "medium", and "small" according to the deterministic difference. A large value indicates that the model has no dispute about the selection of the optimal token, while a small value indicates that the pre-trained language model hesitates between "outputting the optimal token or the top 2 candidate word".

[0085] For the Top-5 Mean (ranging from 0 to 0.8, since the sum of the probabilities of the top 5 does not exceed 1), it can be divided into three intervals: "high", "medium" and "low" according to the concentration of high-probability candidates. A high value indicates that the pre-trained language model tends to favor a few high-probability tokens (low diversity), while a low value indicates that the candidate distribution is scattered (high diversity).

[0086] As shown in Table 1:

[0087] Table 1

[0088]

[0089] In actual execution, after receiving three consecutive local feature values, the temperature decision-maker automatically matches the corresponding discretization intervals and converts the consecutive local feature values ​​into discrete labels of "high / medium / low". For example, the Top1 probability value of 0.92 corresponds to "high", the Top-2 Margin of 0.15 corresponds to "low", and the Top-5 Mean of 0.45 corresponds to "medium".

[0090] Next, a preset feature table is constructed. The horizontal dimension of the feature table is the combination of discretized local feature data, such as Top1 probability value × Top-2 Margin × Top-5 Mean, for a total of 3 × 3 × 3 = 27 basic combinations; the vertical dimension is the text classification result (factual, creative, mixed type). The two are cross-referenced to form 27 × 3 = 81 feature combination scenarios, and each scenario corresponds to a preset temperature value.

[0091] For example, when the combination of local feature data is "Top1 high + Margin high + Mean high" and the text classification result is factual, it indicates that the output of the pre-trained language model is highly deterministic and its accuracy must be guaranteed. Therefore, the temperature parameter is set to a preset temperature value of 0.1. Table 2 shows a partial view of the preset feature table.

[0092] Table 2

[0093]

[0094] For example, when the combination is "Top1 Medium + Margin Low + Mean Medium" and the text classification result is creative, it indicates that the pre-trained language model has some uncertainty and needs to improve diversity; therefore, the temperature parameter is set to the preset value of 0.9. When the combination is "Top1 Low + Margin Medium + Mean Low" and the text classification result is mixed, it indicates that the pre-trained language model is hesitant and needs to balance certainty and diversity; therefore, the temperature parameter is set to the preset value of 0.6. After the feature table is constructed, it is stored as structured data (such as JSON or Excel format) to ensure that the temperature decision-maker can quickly read and match.

[0095] When the temperature decision-maker receives the discretized feature labels (such as Top1 "high", Margin "high", Mean "high", and text classification result "factual"), it first filters out the rows that are consistent with the current text classification result, then matches the feature combination in the row, finds the unique corresponding cell, and extracts the preset temperature value stored in it as the final temperature parameter.

[0096] For example, if the user input is factual text, the local feature data, after discretization, is "Top1 high (0.95), Margin high (0.6), Mean high (0.7)". The temperature decision-maker first locates the "factual text" line, then matches the combination of "Top1 high + Margin high + Mean high", and finally obtains the preset temperature value of 0.2 from the preset feature table as the temperature parameter. The whole process does not require complex calculations and can be completed by simply looking up the table. The time consumption is negligible and it is fully adapted to the real-time requirements of token-level temperature adjustment.

[0097] Targeted segmentation makes temperature adjustments more precise, improving the balance between accuracy and diversity. Discrete labels for different features directly correspond to explicit temperature control logic. A high optimal probability value corresponds to a lower temperature to solidify accuracy, while a low optimal probability value corresponds to an appropriate increase in temperature to mitigate uncertainty. Small word probability differences correspond to increased temperature to increase candidate choices, while large word probability differences correspond to lower temperature to lock in the optimal result. Low average word probability corresponds to decreased temperature to avoid logical confusion, while high average word probability corresponds to increased temperature to enrich expression. This direct binding of labels to control logic eliminates the need for additional feature interpretation in temperature decisions, reducing decision bias and ensuring that the generated results not only match scenario requirements but also dynamically respond to the real-time state generated by the model, resulting in a more stable and controllable balance.

[0098] S204. Adjust the probability distribution of the current time step based on the temperature parameter.

[0099] In some embodiments, when performing step S204, the log odds of the pre-trained language model are first obtained, and then the value of each dimension in the log odds is divided by the temperature parameter to obtain the scaling result. The scaling result is then normalized to obtain the probability distribution of the current time step.

[0100] The temperature parameter T output by the temperature decision unit t Substitute into the standard sampling formula P t `=Softmax(Logits t / T t Specifically, the first step is to obtain the logits of the raw output of the pre-trained language model. Within each token generation cycle, the pre-trained language model encodes the current input sequence (including generated tokens and user prompts), and calculates the logits using modules such as multi-head attention and feedforward networks. The final output is a logits vector (i.e., unnormalized log probabilities) with a dimension equal to the vocabulary size (V). Since logits may have the risk of numerical overflow, a simple preprocessing is required: the maximum value is subtracted from the entire logits vector to avoid subsequent division by the temperature parameter T. t Numerical anomalies may occur later, but the relative differences between the dimensions remain unchanged, ensuring that the adjustment logic of the probability distribution is not affected.

[0101] Secondly, it is based on the temperature parameter T. t Logits scaling operation. Obtain the current time step T from the temperature decision maker. t Then, the value of each dimension in the preprocessed Logits vector is divided by T. t This allows for adjustment of the smoothness of the probability distribution. When T t When the value is less than 1, the difference in Logits after scaling will be amplified. For example, when Logits is [3,1], T t After scaling to [6,2] by 0.5, the proportion of high-probability tokens will further increase after subsequent Softmax normalization, resulting in a more certain generation result; when T... t When the value is greater than 1, the difference in Logits after scaling will decrease. For example, when Logits is [3,1], T t =2 scaling results in [1.5, 0.5], and after Softmax, the probability of each token is more even, improving the diversity of generation; when T t When =1, the scaled Logits are consistent with the original, equivalent to regular sampling. The entire scaling process is an element-wise division operation with a complexity of only O(V) (V is the vocabulary size), and can be completed quickly through hardware acceleration, adapting to real-time inference requirements.

[0102] Next comes the normalization calculation using the Softmax function, which transforms the scaled Logits into a valid probability distribution. The Softmax function is then applied to the scaled Logits vector to ensure that the output P... t The vector contains all elements with values ​​between 0 and 1, and their sum is 1, which conforms to the basic characteristics of a probability distribution.

[0103] In the above embodiments, the log-odds ratio, as the unnormalized raw output of the pre-trained language model, reflects the relative importance of each token. By dividing by the temperature parameter dimension by dimension, the steepness of the probability distribution can be precisely adjusted. Subsequent normalization ensures that the scaled result meets the basic characteristics of the probability distribution, avoiding numerical anomalies. This ensures that the adjusted probability distribution not only conforms to the temperature control intention but can also be directly used for the next word sampling. The temperature scaling of the log-odds ratio is essentially an element-wise division operation with a complexity of only O(V) (V is the vocabulary size). The normalization process can also be efficiently parallelized through the optimized interface of the deep learning framework. Neither has complex matrix operations or iterative processes, and the computation time per time step is negligible. Compared to schemes that rely on complex neural networks to adjust the probability distribution, this process avoids additional model inference overhead and can be seamlessly integrated into token-level dynamic temperature sampling. It achieves real-time control of the probability distribution without slowing down the generation speed, ensuring model response performance in high-concurrency, low-latency scenarios.

[0104] This method can also be based on the normalized probability distribution P t The generation of the next word is adapted to the temperature parameters. The appropriate sampling strategy is selected based on task requirements: if prioritizing generation accuracy is necessary (e.g., for factual questions and answers), a "greedy sampling" approach can be used, directly selecting P. t The token with the highest probability in the middle is used as the next generated word; if a balance between accuracy and diversity (such as mixed text types) is needed, "Top-K sampling" can be used to first select the P... t The top K tokens with the highest probability (e.g., K=50) are selected, and their probabilities are renormalized before random sampling. For finer control over diversity (e.g., creative writing), "Nucleus sampling (Top-p sampling)" can be used, accumulating and selecting P tokens. t Tokens are sorted from high to low probability until the cumulative probability reaches a preset threshold p (e.g., p=0.9), and then random sampling is performed within that subset.

[0105] Regardless of the strategy chosen, the sampling process is based on P. t The provided probability weights ensure that the selection of generated words is related to the temperature parameter T. t Adaptable, low T t Corresponding to high deterministic sampling, high T tWith highly random sampling, the final generated token will be concatenated to the input sequence and enter the temperature adjustment and sampling process of the next inference cycle, forming a complete dynamic generation closed loop.

[0106] In summary, the temperature sampling method provided in this application optimizes the sorting complexity from O(VlogV) to O(V+NlogN) by using heap sort to filter the top N probability values. Combined with the low-complexity calculation of local features such as Top-2 Margin, Top-5 Mean, and optimal probability value, the feature extraction overhead is reduced. The text classifier adopts a pre-trained model with frozen parameters and a single-layer linear head architecture, resulting in extremely low training and inference costs. The computational complexity of the classification stage is only O(H*C). Temperature decision is achieved through feature discretization and lookup of a preset feature table, with the decision time approaching O(1). With the efficient operation of logarithmic probability scaling and normalization, the additional overhead of the entire process is negligible, perfectly adapting to high-concurrency, low-latency real-time inference scenarios and supporting token-level dynamic temperature adjustment requirements.

[0107] In terms of generation quality, the collaborative decision-making of multi-dimensional local features and text classification results enables precise characterization of the generation state: the optimal probability value reflects the model confidence, the word probability difference quantifies the selection of certainty, and the word probability mean measures diversity. These three factors, combined with the text type (factual / creative / mixed type), upgrade temperature adjustment from a single dimension to a dual-driven approach of scene and state. By scaling and normalizing the log probability using temperature parameters, the steepness of the probability distribution can be flexibly adjusted. Low temperatures enhance accuracy, while high temperatures improve diversity. Combined with the temperature matching mechanism of the two-layer retrieval, this effectively balances the core needs of different generation tasks, reduces factual errors and logical confusion, and avoids problems of repetitive expression and lack of creativity.

[0108] like Figure 3 As shown, Figure 3 This is a flowchart illustrating another temperature sampling method provided in an embodiment of this application.

[0109] First, the process starts with the input sequence. The pre-trained language model encodes and infers the input sequence, outputting the hidden layer state (Ht0) and log odds (Logits). The hidden layer state serves as the feature input for the text classifier; the log odds are used for subsequent probability distribution calculations.

[0110] The text classifier categorizes hidden layer states into factual, creative, or mixed-type texts, and this classification result is a key basis for temperature decision-making. The log odds are normalized using the Softmax function to transform into a probability distribution, and then the Top-2 Margin, Top-5 Mean, and Top-1 local feature data are extracted from this distribution. These features directly reflect the confidence and uncertainty of the model in generating the current token and are one of the core signals for temperature adjustment.

[0111] Subsequently, the local feature data and the text classification results are input into the temperature decision-maker. The temperature decision-maker obtains the temperature parameters through discretized features and table lookup matching. Then, the temperature parameters are substituted into the formula, and the original log-probability is scaled before recalculating the Softmax to obtain the adjusted probability distribution. Low temperatures amplify the proportion of high-probability tokens to increase certainty, while high temperatures make the probability distribution more uniform to increase diversity.

[0112] Finally, based on the adjusted probability distribution, a new token is selected through a sampling strategy and fed back to the pre-trained language model as input for the next round of inference. This process is repeated until a complete sequence is generated. Dynamic temperature adaptation during the generation process of the pre-trained language model is achieved, ensuring generation quality while reducing overhead through lightweight design.

[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0114] like Figure 4 As shown, embodiments of this application also provide a temperature sampling device, which includes:

[0115] The classification module 401 is used to classify the input sequence of the pre-trained language model to obtain the text classification result;

[0116] The feature extraction module 402 is used to extract local feature data from the probability distribution of the current time step of the pre-trained language model when the pre-trained language model performs inference based on the input sequence. The local feature data is related to the output quality of the pre-trained language model.

[0117] The temperature decision module 403 is used to input local feature data and text classification results into the temperature decision unit to obtain temperature parameters; the temperature parameters are preset temperature values ​​in the preset feature table that match the local feature data and text classification results.

[0118] The adjustment module 404 is used to adjust the probability distribution of the current time step based on the temperature parameter.

[0119] As an optional implementation provided in this application, the classification module 401 is used to: classify the input sequence using a pre-trained text classifier, and select the category corresponding to the dimension with the highest probability value as the text classification result. The text classification result includes factual, creative, and mixed types. The pre-trained text classifier is constructed based on a feature extraction model with frozen parameters and a single-layer linear classification head. The single-layer linear classification head is obtained through supervised training based on training data and a preset loss function. The training data generates soft labels in the form of a three-dimensional probability distribution through a teacher model.

[0120] As an optional implementation provided in this application, the device further includes a training module for training a single-layer linear classifier head, comprising: constructing an instruction sample library and generating soft labels for each sample through a teacher model to obtain training data; the instruction sample library includes factual text, creative text, and mixed-type text; calculating the loss value between the probability distribution output by the single-layer linear classifier head based on the instruction sample library and the soft labels according to a preset loss function; and updating the parameters of the single-layer linear classifier head according to the loss value until the loss value converges.

[0121] As an optional implementation provided in this application, the local feature data includes at least one of the following: optimal probability value, word probability difference, and word probability mean; wherein, the optimal probability value is the highest probability value output by the pre-trained language model, used to reflect the confidence of the pre-trained language model in the currently generated word; the word probability difference is the difference between the probabilities of the pre-trained language model output ranking in the top M positions, and the word probability difference is used to quantify the deterministic differences in the output of the pre-trained language model; the word probability mean is the average of the probabilities of the pre-trained language model output ranking in the top N positions, and the word probability mean is used to quantify the diversity of the output of the pre-trained language model; M is less than N.

[0122] As an optional implementation provided in this application embodiment, the feature extraction module 402 is used to: sort the probability distribution of the current time step using a heap sort algorithm when the pre-trained language model performs inference based on the input sequence, and filter out the top N probability values; determine the optimal probability value ranked first among the top N probability values; calculate the word probability difference by subtracting the second-ranked probability value from the optimal probability value ranked first; and calculate the arithmetic mean of the top N probability values ​​to determine the word probability mean.

[0123] As an optional implementation provided in this application, the temperature decision module 403 is used to: convert local feature data into discrete labels through a temperature decision unit; retrieve unit data that matches the text classification result from a preset feature table; and retrieve temperature parameters corresponding to the discrete labels from the unit data.

[0124] As an optional implementation provided in this application embodiment, the temperature decision module 403, during the process of converting local feature data into discrete labels through the temperature decision unit, is used to: classify discrete labels according to confidence level for the optimal probability value; classify discrete labels according to deterministic difference for the word probability difference; and classify discrete labels according to diversity for the word probability mean. As an optional implementation provided in this application embodiment, the horizontal dimension of the preset feature table is the discretized local feature data, and the vertical dimension is the text classification result.

[0125] As an optional implementation provided in this application, the adjustment module 404 is used to: obtain the log odds of the pre-trained language model; divide the values ​​of each dimension in the log odds by the temperature parameter to obtain the scaling result; and normalize the scaling result to obtain the probability distribution.

[0126] For a description of the features in the embodiment corresponding to the temperature sampling device, please refer to the relevant description of the embodiment corresponding to the temperature sampling method, which will not be repeated here.

[0127] like Figure 5 As shown, embodiments of this application also provide an electronic device, including a memory 501 and a processor 502, wherein the memory 501 stores a computer program and the processor 502 is configured to run the computer program to perform the steps in any of the temperature sampling method embodiments described above.

[0128] Embodiments of this application also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described temperature sampling method embodiments when running.

[0129] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0130] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the temperature sampling method embodiments described above.

[0131] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described temperature sampling method embodiments.

[0132] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0133] The temperature sampling method and electronic device provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A temperature sampling method, characterized in that, include: The input sequence of the pre-trained language model is classified to obtain the text classification result; The text classification results include factual text, creative text, and mixed-type text; When the pre-trained language model performs inference based on the input sequence, local feature data is extracted from the probability distribution of the current time step of the pre-trained language model. The local feature data is related to the output quality of the pre-trained language model and includes at least one of the following: the optimal probability value reflecting the model's confidence in the currently generated word, the word probability difference quantifying the deterministic differences in the model output, and the word probability mean quantifying the diversity of the model output. The local feature data and the text classification result are input into the temperature decision-maker to obtain temperature parameters; the temperature parameters are preset temperature values ​​in the preset feature table that match the local feature data and the text classification result. The probability distribution of the current time step is adjusted based on the temperature parameter.

2. The method according to claim 1, characterized in that, The process of classifying the input sequences of the pre-trained language model to obtain text classification results includes: The input sequence is classified using a pre-trained text classifier, and the category corresponding to the dimension with the highest probability value is selected as the text classification result. The pre-trained text classifier is constructed based on a feature extraction model with frozen parameters and a single-layer linear classification head. The single-layer linear classification head is obtained through supervised training based on training data and a preset loss function. The training data generates soft labels in the form of a three-dimensional probability distribution through a teacher model.

3. The method according to claim 2, characterized in that, The training process of the single-layer linear classification head includes: A sample instruction library is constructed, and soft labels for each sample are generated using the teacher model to obtain the training data; the sample instruction library includes factual text, creative text, and mixed-type text. The loss value between the probability distribution output by the single-layer linear classification head based on the instruction sample library and the soft label is calculated according to the preset loss function. The parameters of the single-layer linear classifier are updated based on the loss value until the loss value converges.

4. The method according to claim 1, characterized in that, The optimal probability value is the highest probability value output by the pre-trained language model; The word probability difference is the difference between the probabilities of the pre-trained language model outputting a rank in the top M positions. The mean probability of a word is the average of the probabilities of the pre-trained language model outputting a word that ranks in the top N positions; M is less than N.

5. The method according to claim 4, characterized in that, When the pre-trained language model performs inference based on the input sequence, extracting local feature data from the probability distribution of the current time step of the pre-trained language model includes: When the pre-trained language model performs inference based on the input sequence, a heap sort algorithm is used to sort the probability distribution of the current time step and select the top N probability values. For the probability values ​​of the top N, determine the optimal probability value that ranks first. The probability difference of the word is calculated by subtracting the probability value of the second-ranked word from the first-ranked optimal probability value. Calculate the arithmetic mean of the probability values ​​of the top N probabilities to determine the mean probability of the term.

6. The method according to claim 4, characterized in that, The step of inputting the local feature data and the text classification result into the temperature decision-maker to obtain temperature parameters includes: The local feature data is converted into discrete labels using the temperature decision-maker. Retrieve unit data that matches the text classification result from the preset feature table; Retrieve the temperature parameter corresponding to the discrete label from the unit data.

7. The method according to claim 6, characterized in that, The process of converting the local feature data into discrete labels using the temperature decision-maker includes: For the optimal probability value, the discrete labels are divided according to the confidence level; Based on the word probability difference, the discrete labels are divided according to the deterministic difference; Based on the mean probability of the lexical units, the discrete labels are divided according to diversity.

8. The method according to claim 1, characterized in that, The horizontal dimension of the preset feature table is the discretized local feature data, and the vertical dimension is the text classification result.

9. The method according to claim 1, characterized in that, The step of adjusting the probability distribution of the current time step based on the temperature parameter includes: Obtain the log odds of the pre-trained language model; The scaling result is obtained by dividing the value of each dimension of the logarithmic probability by the temperature parameter. The probability distribution is obtained by normalizing the scaling result.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the temperature sampling method as described in any one of claims 1 to 9 when executing the computer program.