Inference acceleration method for sentiment classification task, storage medium and device

By introducing a cache bypass and a token filter in the sentiment classification task, and utilizing GPU memory to cache highly saliency token embeddings, the problem of excessively long inference time in deep learning models is solved, thereby improving inference speed and hit rate.

CN120996200AActive Publication Date: 2025-11-21SUZHOU NEW HOPE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511167184.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-21
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

Deep learning models take too long to infer in sentiment classification tasks, resulting in low efficiency.

Method used

We employ a BERT classification model combined with a cache bypass and a token filter. We discard low-salience tokens and retain high-salience tokens by calculating the salience score of the token embedding. We also cache the representation vector and label pairs in GPU memory and use dot product to calculate similarity to accelerate the inference process.

Benefits of technology

It improves the inference speed of sentiment classification tasks, enhances the hit rate and accuracy of cache lookup, and reduces inference time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996200A_ABST
    Figure CN120996200A_ABST
Patent Text Reader

Abstract

The invention relates to a reasoning acceleration method for an emotion classification task, a storage medium and a device. The method mainly comprises a BERT classification model and further comprises a cache bypass, a Token filter is arranged in the cache bypass, and the method is based on approximate cache, that is, the closer the two inputs are embedded in a high-dimensional space, the higher the possibility that the two inputs belong to the same category is. According to the method, a trained Token filter (Token Filter) is adopted to predict the significance score of each Token in the input, and the low-score Token is filtered out according to the significance score of each Token. And a dot product method is adopted, so that the cache search speed in the GPU memory is improved, the hit rate and the accuracy are improved, and the interpretability of the similarity in an application scene is enhanced. By introducing the novel similarity caching mechanism, the problem that the reasoning time is too long in the sentiment classification task is solved, the reasoning speed can be effectively increased, and the reasoning time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of emotion classification task technology, and in particular to inference acceleration methods, storage media and apparatus for emotion classification tasks. Background Technology

[0002] Sentiment classification is a fundamental task in Natural Language Processing (NLP) that aims to identify and classify opinions or sentiments expressed in textual data. It typically involves categorizing text into predefined sentiment categories, such as positive, negative, or neutral. Applications of sentiment analysis span multiple fields, including social media monitoring, customer feedback assessment, and product review analysis.

[0003] Deep learning (DL) methods have demonstrated superior performance in sentiment classification. However, the large number of parameters in these models often leads to excessively long inference times. Summary of the Invention

[0004] Based on this, a method for accelerating inference in sentiment classification tasks is provided. This method helps to improve the inference speed of sentiment classification tasks, that is, to shorten the inference time.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: An inference acceleration method for sentiment classification tasks includes a BERT classification model, a cache bypass with a token filter, and the following steps: S100. Represent the input text as a format of size [size missing]. n × h matrix X, X = [ x i ] i∈[n] ,in x i For the first i Embedding of individual tokens, S200. Input matrix X into the Token filter and calculate the saliency score (Saliency(xi)) for each Token embedding. Based on the saliency score (Saliency(xi), discard q proportions of Token embeddings with lower saliency scores (Saliency(xi)), where q ∈ [0, 1). Retain the remaining (1 − q) proportion of Token embeddings as a representation matrix of size (1 − q)n × h. , ,in The matrix is ​​formed by summing along the embeddings of the (1 −q) × n token indices before the significance score (Saliency(xi)). Merge them into a new vector x of size h, that is ,in Let x be the j-th value of the embedding xi, therefore, x represents the embedding of the corresponding input text. S300: A batch of entries is stored in the cache, representing vector-label pairs. ,in For the first k Embedding of cache entries, y k For the corresponding tags, K The input is calculated using the dot product to determine the cache size. x Similarity to all cached entries, i.e. Choose the most similar m There are 1 sample, denoted as _ . Statistics belong to the same category y Number of embeddings If the maximum value among all categories is max y ( d y Greater than If α ∈ (0, 1], it is considered a hit, and the label result is returned directly. y If max y ( d y ) ≤ If the cache misses, the input representation X obtained from the embedding layer is input into the BERT classification model, and the classification process of the BERT classification model begins.

[0006] In one embodiment, the cache is stored in GPU memory.

[0007] In one embodiment, in step S100, the input text is first tokenized using a tokenizer. n Each token is used to generate a token embedding representation, position embedding, and segment embedding through an embedding layer. Each embedding has a dimension of [missing information]. h .

[0008] In one embodiment, the method for calculating the salience score of each token embedding in step S200 includes: , where vi and xi represent the output log odds and embedding of the i-th token, respectively, and the operator ⊙ represents element-wise multiplication.

[0009] In one embodiment, in step S200, the discarding q Embedsions with lower proportion significance scores (Saliency(xi)) include: sorting each token embedding according to its significance score (Saliency(xi)) from lowest to highest. q The proportion of token embeddings is discarded.

[0010] In one embodiment, the token filter includes an MLP network to predict the saliency score of each token embedding.

[0011] In one embodiment, during training, the MLP network is updated by minimizing the mean squared error loss between the output of the MLP network and the significance value calculated through backpropagation.

[0012] In one embodiment, the input text includes multiple sentences.

[0013] A computer storage medium storing at least one executable instruction that causes a processor to perform operations corresponding to the inference acceleration method for the emotion classification task.

[0014] A computer device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other through the communication bus. The memory stores at least one executable instruction, which causes the processor to perform an operation corresponding to the inference acceleration method for the emotion classification task.

[0015] The beneficial effects of this application are as follows: This application proposes a GPU in-memory caching mechanism specifically designed for sentiment classification tasks. Sentiment classification involves classifying text into different sentiments, and models are trained to perform this classification. The method in this application is based on approximate caching, meaning that the closer the embeddings of two inputs are in a high-dimensional space, the higher the probability that they belong to the same category. This application employs a trained token filter to predict the saliency score of each token in the input and filters out low-scoring tokens accordingly. A dot product method is used, which not only improves the speed of in-memory cache lookup on the GPU, thereby increasing hit rate and accuracy, but also enhances the interpretability of similarity in application scenarios. By introducing this novel similarity caching mechanism, the aim is to address the problem of excessively long inference time in sentiment classification tasks, effectively improving inference speed and reducing inference time. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating an embodiment of the reasoning acceleration method for an emotion classification task according to this application.

[0017] Figure 2 This is a schematic diagram of the token composition of the input text in an embodiment of this application. Detailed Implementation

[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0019] like Figure 1 As shown, compared with the traditional BERT classification task process, the method of this application introduces caching and token filtering into the process. The complete reasoning process of the method of this application is as follows: S100. The input text is first tokenized using a tokenizer. n Each token is used to generate a token embedding representation, position embedding, and segment embedding through an embedding layer. Each embedding has a dimension of [missing information]. h Next, the input text will be represented as a value of [size missing]. n × h matrix X, X = [ x i ] i∈[n] ,in x i For the first i Embedding of individual tokens.

[0020] S200. In the cache bypass, input matrix X into the token filter, calculate the saliency score (Saliency(xi)) for each token embedding, and discard q proportions of token embeddings with lower saliency scores (Saliency(xi)) based on the saliency scores (Saliency(xi)), where q ∈ [0, 1). Retain the remaining (1 − q) proportion of token embeddings as a representation matrix of size (1 − q)n × h. , ,in The matrix is ​​formed by summing along the embeddings for the (1 − q) × n token indices before the significance score (Saliency(xi)). Merge them into a new vector x of size h, that is ,in Let x be the j-th value of the embedding xi, therefore, x represents the embedding of the corresponding input text.

[0021] S300: A batch of entries is stored in the cache, representing vector-label pairs. ,in For the first k The embedding of each cached entry, which corresponds to the "embedding of input text" at the end of S200, is the embedding of the cached sample after processing by step S200. y k For the corresponding tags, K Given the cache size, the similarity between the input x (the new vector x in S200) and all cache entries is calculated using the dot product, i.e. This allows for full utilization of CUDA's optimizations for matrix multiplication, resulting in a significant speedup. Selecting the most similar... m m samples (ranked by similarity from highest to lowest, selecting the top m), denoted as [m samples]. Statistics belong to the same category y Number of embeddings If the maximum value among all categories is max y ( d y Greater than If α ∈ (0, 1], then it is considered a hit, and the label result is returned directly. y If max y ( d y ) ≤ If the cache misses, the input representation X (matrix X) obtained from the embedding layer is input into the remaining BERT pipeline, and the normal BERT classification process begins.

[0022] It's important to note that the Token filter and cache need to be initialized before the inference process begins. The Token filter is a pre-trained MLP network that approximates the token using a saliency function. The cache is created by randomly selecting tokens from the training set. K The system initializes the data points by calculating their representation matrices, with the specific proportion of sample initialization determined by user-defined system parameters. Considering that GPUs typically handle the majority of computation in sentiment classification tasks due to their faster processing speed compared to CPUs, and that data exchange latency between CPUs and GPUs is significant, the cache is stored in GPU memory instead of common external storage such as hard drives. This accelerates the inference process.

[0023] The method in this application is inspired by the unique challenges of sentiment classification tasks. Unlike general classification tasks, sentiment analysis typically relies on a small number of key tokens (such as adjectives and adverbs) carrying significant sentiment cues. These tokens often dominate the sentiment of a sentence, making it possible to discard less relevant context without affecting classification accuracy. In sentiment classification tasks, the input is first tokenized, divided into multiple tokens, each represented by a high-dimensional embedding, thus the input is represented as a size... n × h Storing such inputs in a cache is costly, and measuring the similarity between two input matrices remains challenging because caching systems must balance the effectiveness of similarity metrics with search speed. This application's method significantly speeds up cache lookup by aggregating the embeddings of all tokens into a single vector as a representation of the sentence, transforming cache lookup from matrix operations to vector computation. However, this method raises a new question: can the overall semantics of the sentence be adequately represented by the summation of the semantics of individual tokens? Here, this application uses a scenario as an example. Consider a binary sentiment analysis task, labeled as positive and negative. There are two inputs: 1) "I think the movie is quite interesting." 2) "I think the movie is quite boring." Clearly, these two inputs have opposite semantics, but differ by only one word, with the rest being identical. If we superimpose the token embeddings of these inputs, their distinguishability might be low, similar to comparing two polynomials that differ by only one term. However, if all other terms in the polynomial are zero, their distinguishability can be maximized.

[0024] The method in this application is based on this principle: the information in the input undoubtedly has varying importance to the current task, and not all information has equal weight. For example, in the above example, "interesting" and "boring" are clearly the most critical words. By selectively identifying such key information from the input, it is hoped to enhance the discriminative ability of different categories of samples in the cache, thereby improving the cache hit rate and hit accuracy.

[0025] Saliency-Based Filter Design: This application employs gradient-based saliency as a measure of token importance, aiming to retain tokens with high saliency scores and discard others. Saliency is a gradient-based interpretability method used to detect the importance of input features. In text classification scenarios, the first derivative of the output logits with respect to the input embedding is calculated to obtain the sensitivity of each class to each dimension of the embedding. This sensitivity is then multiplied element-wise with the corresponding value in the embedding to obtain the influence of each dimension of the embedding on each class. Since each token is represented by an embedding, obtaining the saliency score for each token requires L2 normalization or averaging of the result value. In the method of this application, the saliency score of each token is defined as: , Where vi and xi represent the output log odds and embedding of the i-th token, respectively, and the operator ⊙ represents element-wise multiplication.

[0026] Considering the computational cost of calculating saliency values, it is impractical to compute gradients via backpropagation during inference. Therefore, a Multilayer Perceptron (MLP) is employed to predict the saliency score for each token. This MLP consists of two feedforward layers with a ReLU activation function in between. During training, the MLP is updated by minimizing the mean squared error (MSE) loss between the MLP output and the saliency value computed via backpropagation. Although the learning capacity of a single MLP is limited, it is sufficient for filtering purposes because only the relative score ranking of each token is needed, rather than the exact saliency score computed via backpropagation.

[0027] Furthermore, to discuss the discriminative nature of samples in the cache, we assume a binary sentiment classification scenario. The length of the input sequence is denoted as n, representing the number of tokens in each input. Based on the significance score, tokens are divided into three categories: positive, negative, and neutral tokens. It should be noted that both positive and negative tokens have high significance scores, indicating discriminable sentiment. As shown in Figure 2, for each input, the token filter discards q × n tokens, where 0 ≤ q < 1. On average, among the discarded q × n tokens, there are a × q × n neutral tokens and b × q × n positive or negative tokens, where both a and b belong to the interval [0, 1) and a + b = 1.

[0028] Since similarity is calculated by summing the token embeddings and using a dot product, the similarity between two inputs A and B can be expressed as the sum of the similarities for each token pair: , in, Let represent the similarity between the i-th token of A and the j-th token of B. In other words, using the dot product to measure similarity is essentially summing the similarity of all token pairs between the two sentences. To ensure that differences in sample length do not affect the similarity calculation, we normalize the number of retained tokens for all samples through token filtering. Specifically, in the experiment, the parameter n (sequence length) is set uniformly for all samples, and the proportion q of discarded tokens is fixed, so the number of retained tokens (1−q)n is the same for all samples. This ensures the fairness of the similarity calculation because the retained token embeddings can be directly compared regardless of the original input length. The normalized similarity between two inputs A and B is then calculated. Defined as the difference between the original similarity score S and its minimum value divided by the range of the original similarity scores: , Here, min(S) and max(S) represent the minimum and maximum values ​​of all calculated pairwise similarity scores, respectively. This normalization process facilitates comparison of similarity scores between different sample pairs without being affected by their absolute magnitude. Using normalized similarity to represent discriminativeness ensures that the discriminative measure is not overly influenced by the absolute value of the similarity score, but focuses on the relative similarity between samples, which is more relevant for distinguishing different categories.

[0029] Naturally, the discriminative power between positive and negative samples can be represented by the difference in normalized similarity: , Here, A1 and B represent samples from different categories, while A1 and A2 represent samples from the same category, signifying that their token distribution reflects the overall average distribution. Clearly, as long as the cache size isn't too small, the two samples with the highest similarity must be from the same category, and the proportion of positive (or negative) tokens within them will inevitably be higher than the average. Since the similarity between positive (or negative) tokens is higher than the similarity between other token combinations, the similarity score will be higher. The average similarity between positive (or negative) tokens is denoted as... The average similarity between a neutral token and other tokens is denoted as . The average similarity between positive and negative tokens is denoted as Obviously, .

[0030] set up , and For the filtered samples, since some tokens were discarded, their contribution to the final input similarity was also removed. Therefore, we have: , , because There exists N > M. Assuming the token filter accurately removes the average number of neutral tokens, i.e., all neutral tokens in A1, A2, and B are filtered out, then b = 0. For these three samples, we have: , Therefore, the discriminative power of the filtered samples for: .

[0031] Therefore, if the filter can filter out an appropriate number of tokens, it can improve the discriminative power of samples in the cache.

[0032] Discriminability Defined as the normalized difference between intra-class and inter-class similarity, it directly reflects the clustering behavior of samples in the cache. When discriminancy is high, intra-class samples exhibit higher similarity, forming denser clusters, while inter-class samples are more dispersed. This enhanced segregation allows the cache to more effectively retrieve the correct label based on the nearest neighbor cache entry of the input sample, thereby improving hit accuracy and overall cache efficiency.

[0033] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for accelerating inference in a sentiment classification task, characterized in that, The system includes the BERT classification model, a cache bypass with a token filter, and the following steps: S100. Represent the input text as a format of size [size missing]. n × h matrix X,X = [ x i ] i∈[n] ,in x i For the first i Embedding of individual tokens, S200. Input matrix X into the Token filter and calculate the saliency score (Saliency(xi)) for each Token embedding. Based on the saliency score (Saliency(xi)), discard q proportions of Token embeddings with lower saliency scores (Saliency(xi)), where q ∈ [0, 1). Retain the remaining (1 − q) proportion of Token embeddings as a representation matrix of size (1 − q)n × h. , ,in The matrix is ​​formed by summing along the embeddings for the (1 − q) × n token indices before the significance score (Saliency(xi)). Merge them into a new vector x of size h, that is ,in Let x be the j-th value of the embedding xi, therefore, x represents the embedding of the corresponding input text. S300: A batch of entries is stored in the cache, representing vector-label pairs. ,in For the first k Embedding of cache entries, y k For the corresponding tags, K Given the cache size, the similarity between the input x and all cache entries is calculated using the dot product, i.e. Choose the most similar m There are 1 sample, denoted as _ . Statistics belong to the same category y Number of embeddings If the maximum value among all categories is max y ( d y Greater than If α ∈ (0, 1], it is considered a hit, and the label result is returned directly. y If max y ( d y ) ≤ If the cache misses, the input representation X obtained from the embedding layer is input into the BERT classification model, and the classification process of the BERT classification model begins.

2. The reasoning acceleration method for sentiment classification tasks according to claim 1, characterized in that, Store the cache in GPU memory.

3. The reasoning acceleration method for sentiment classification tasks according to claim 1, characterized in that, In step S100, the input text is first tokenized using a tokenizer. n Each token is used to generate a token embedding representation, position embedding, and segment embedding through an embedding layer. Each embedding has a dimension of [missing information]. h .

4. The reasoning acceleration method for sentiment classification tasks according to claim 1, characterized in that, In step S200, the method for calculating the salience score of each token embedding includes: , Where vi and xi represent the output log odds and embedding of the i-th token, respectively, and the operator ⊙ represents element-wise multiplication.

5. The inference acceleration method for emotion classification tasks according to claim 4, characterized in that, In step S200, the discarding q Embedsions with lower proportion significance scores (Saliency(xi)) include: sorting each token embedding according to its significance score (Saliency(xi)) from lowest to highest. q The proportion of token embeddings is discarded.

6. The reasoning acceleration method for sentiment classification tasks according to claim 1, characterized in that, The token filter includes an MLP network, which predicts the saliency score for each token embedding.

7. The reasoning acceleration method for emotion classification tasks according to claim 6, characterized in that, During training, the MLP network is updated by minimizing the mean squared error loss between the output of the MLP network and the significance value calculated through backpropagation.

8. The inference acceleration method for sentiment classification tasks according to claim 1, characterized in that, The input text includes multiple sentences.

9. A computer storage medium, characterized in that, The computer storage medium stores at least one executable instruction that causes the processor to perform the operation corresponding to the reasoning acceleration method for sentiment classification tasks as described in any one of claims 1 to 8.

10. A computer device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided. The processor, memory, and communication interface communicate with each other through the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the reasoning acceleration method for sentiment classification tasks as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Text sentiment classification method and device, electronic equipment and storage medium

    CN118113864A

  • Semi-supervised method and apparatus for public opinion text analysis

    WO2023092961A1