Multi-platform semantic feature fusion method based on large language model
By adopting an improved dynamic self-distillation algorithm based on BERT and a sparse principal component linear discriminant fusion algorithm in public opinion analysis, the problem of semantic feature fusion of multi-platform social media data is solved, and the accuracy and reliability of the analysis are improved.
Patent Information
- Application Number
- CN202510154905.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology is difficult to effectively integrate semantic features in multi-platform social media data, which has affected the accuracy and reliability of public opinion analysis, especially when facing data with multi-platform, multi-lingual and complex emotional expressions.
The improved dynamic self-distillation algorithm based on large language model (BERT) is adopted to optimize feature extraction and classification accuracy by dynamically adjusting the distillation strategy. Combining sparse principal component analysis (Sparse PCA) and linear discriminant analysis (LDA), the high-dimensional semantic features of multi-platform and multilingual species are mapped to a unified low-dimensional space, removing redundant features and enhancing category distinction.
The accuracy of large language models in complex public opinion analysis was improved, the semantic feature fusion in multi-platform and multi-lingual data was optimized, and the efficient adaptability and feature expression ability of the model in multi-platform data was improved.
Smart Images

Figure CN120086794A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data fusion and feature optimization, and particularly relates to a multi-platform semantic feature fusion method based on a large language model. Background Art
[0002] With the rapid development of social media platforms, public opinion analysis has become an important tool for social governance, public safety, and enterprise decision-making. Social media platforms such as Twitter, WeChat, and Weibo have significant differences in user groups, information dissemination methods, and content formats among different platforms, which poses challenges for public opinion analysis in terms of multi-platform information fusion. Currently, most research focuses on data analysis of a single social media platform, such as independent public opinion monitoring of Twitter, WeChat, or Weibo. However, these methods ignore the information complementarity across platforms. The differences in user behavior, information dissemination mechanisms, and content formats between different platforms lead to the fact that the results of public opinion analysis on a single platform often cannot comprehensively reflect the overall mood and views of society, and are prone to bias and one-sidedness.
[0003] Although existing research has attempted to conduct cross-platform public opinion analysis by aggregating data from different platforms, in practical applications, the fusion of cross-platform data still faces many difficulties. Existing methods are mainly limited to simple feature splicing or analysis based on a single language, lacking effective extraction and synthesis of different semantic features in multi-platform data. The heterogeneity, noise problems, and cross-platform semantic differences of data from different platforms greatly reduce the accuracy and reliability of public opinion analysis. In addition, although large language models (such as BERT) have made significant progress in the field of natural language processing, when facing public opinion data with multiple platforms, multiple languages, and complex emotional expressions, they still face the problem of being unable to accurately capture complex language phenomena such as metaphors, sarcasms, and puns, resulting in large errors in the classification and sentiment analysis of specific topics. Therefore, how to improve the accuracy of large language models in complex public opinion analysis, especially for the semantic feature fusion of multi-platform and multi-language data, has become an urgent technical problem to be solved. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a multi-platform semantic feature fusion method based on a large language model, including:
[0005] Obtain data from different social media platforms, preprocess the data from the different social media platforms to obtain a preprocessed data set;
[0006] Construct a BERT model, input the preprocessed data set into the BERT model for calculation to obtain a semantic representation data set;
[0007] Construct a teacher network based on the pre-trained BERT model, and process the BERT model based on the distillation method to obtain a student network;
[0008] Input the semantic representation dataset into the teacher network to obtain the output of the teacher network, and optimize the process of guiding the training of the student model by the output of the teacher network through a dynamic self-distillation mechanism to obtain a document semantic feature extraction model;
[0009] Process the data based on the document semantic feature extraction model to obtain a semantic feature matrix;
[0010] Construct a sparse principal component linear discriminant fusion algorithm, and input the semantic feature matrix into the sparse principal component linear discriminant fusion algorithm for calculation to obtain a final semantic feature fusion matrix.
[0011] Preferably, the process of preprocessing the data of different social media platforms to obtain a preprocessing dataset includes:
[0012] Perform unified character encoding processing on the data of different social media platforms, and perform standardization processing on the encoded data to obtain a standardized dataset;
[0013] Remove the noise and non-semantic information in the standardized dataset to obtain a cleaned dataset;
[0014] Perform multi-language word segmentation processing on the cleaned dataset to obtain a word-segmented dataset;
[0015] Perform lemmatization processing on the word-segmented dataset to obtain the preprocessing dataset.
[0016] Preferably, the process of obtaining the semantic representation dataset includes:
[0017] Convert the preprocessing dataset into an embedding representation;
[0018] Construct a BERT model, and calculate the context representation of each word in the embedding representation through the BERT model;
[0019] Obtain the semantic representation dataset based on the context representation of each word.
[0020] Preferably, the expression of the total loss function for optimizing the process of guiding the training of the student model by the output of the teacher network through a dynamic self-distillation mechanism is:
[0021] L total =DF t ·L distill +(1-DF t )·L task ;
[0022] Among them, L total is the total loss, DF t is the dynamically adjusted distillation factor, L task is the task loss, and L distill is the distillation loss.
[0023] Preferably, the expression of the dynamically adjusted distillation factor is:
[0024]
[0025] Among them, L train (t) represents the training loss at the t-th round of training, and L val (t) represents the validation loss at the t-th round of training. DF max and DF min represent the maximum and minimum values of the distillation factor.
[0026] Preferably, the process of obtaining the document semantic feature extraction model further includes: optimizing the dynamic self-distillation mechanism based on the KL divergence method to obtain an improved KL divergence self-distillation loss;
[0027] The expression of the improved KL divergence self-distillation loss is:
[0028] L(x i ) = λ · (KL improved (P(x i ) || Q(x i )));
[0029] Among them, x i represents the input text, P(x i ) is the prediction distribution of the student model, Q(x i ) is the prediction distribution of the teacher model, λ is the dynamic distillation factor, and KL improved (P(x i ) || Q(x i )) is the improved KL divergence algorithm.
[0030] Preferably, the expression of the improved KL divergence algorithm is:
[0031]
[0032] Among them, α and β are hyperparameters, and γ is a hyperparameter.
[0033] Preferably, the expression of the final semantic feature fusion matrix is:
[0034]
[0035] Among them, X’ is the standardized input data matrix, and Wfinal is the final semantic feature fusion matrix, W LDA is the linear discriminant analysis matrix, S B is the between-class scatter matrix, S W is the within-class scatter matrix, ν is the regularization parameter, is the data reconstruction error.
[0036] Compared with the prior art, the present invention has the following advantages and technical effects:
[0037] The present invention provides an improved dynamic self-distillation algorithm combined with a large language model (BERT). By dynamically adjusting the distillation strategy, an adaptive balance is achieved between the distillation loss and the task loss. This method can optimize feature extraction and classification accuracy for complex semantic features under different platforms and specific tasks, ensuring the efficient adaptability of the model in multi-platform and multi-language data.
[0038] The present invention proposes a semantic feature fusion algorithm combining sparse principal component analysis (Sparse PCA) and linear discriminant analysis (LDA), which maps high-dimensional semantic features of multiple platforms and multiple languages into a unified low-dimensional space. While reducing the dimension, redundant features are removed by combining sparsity constraints, and the model performance is optimized by enhancing class discrimination, enabling the features after multi-platform fusion to have stronger expressive and classification capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0040] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0042] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0043] Embodiment 1
[0044] As Figure 1 shown, in this embodiment, a multi-platform semantic feature fusion method based on a large language model is provided, including:
[0045] Obtain data from different social media platforms, preprocess the data of the different social media platforms to obtain a preprocessed data set;
[0046] Construct a BERT model, input the preprocessed data set into the BERT model for calculation to obtain a semantic representation data set;
[0047] Construct a teacher network based on the pre-trained BERT model, and process the BERT model based on the distillation method to obtain a student network;
[0048] Input the semantic representation data set into the teacher network to obtain the output of the teacher network, and optimize the process of guiding the training of the student model by the output of the teacher network through a dynamic self-distillation mechanism to obtain a document semantic feature extraction model;
[0049] Process data based on the document semantic feature extraction model to obtain a semantic feature matrix;
[0050] Construct a sparse principal component linear discriminant fusion algorithm, input the semantic feature matrix into the sparse principal component linear discriminant fusion algorithm for calculation to obtain a final semantic feature fusion matrix.
[0051] Specifically:
[0052] Step 1, Multi-platform and multi-language data preprocessing:
[0053] Multi-platform and multi-language data preprocessing includes two parts: one is unified encoding and format standardization of multi-platform data; the other is unified word segmentation and lemmatization of multiple languages.
[0054] Unified encoding and format standardization of multi-platform data:
[0055] The process of unified encoding and format standardization of multi-platform data is as follows:
[0056] First, perform unified character encoding (such as UTF-8) processing on the data of multiple platforms such as platform 1, platform 2,..., platform n. Different platforms may use different encodings, and unified encoding is a basic preprocessing step.
[0057] Furthermore, perform standardization processing on the multi-platform data format, use a unified data format to represent the data of all platforms, here use the CSV format, and at the same time unify the timestamp to a standard time format;
[0058] Unified word segmentation and lemmatization of multiple languages:
[0059] The process of unified word segmentation and lemmatization of multiple languages is as follows:
[0060] First, perform text preprocessing, the purpose of which is to remove noise and non-semantic information. Text preprocessing is the first step in natural language processing, aiming to remove noise and non-semantic information to make subsequent processing more efficient and accurate.
[0061] Remove useless characters: Remove punctuation marks, special characters, HTML tags, etc. from the text.
[0062] Convert to lowercase: Convert all text to lowercase to reduce the impact of case on the processing results.
[0063] Remove stop words: Remove frequently occurring words with little semantic contribution such as "the", "a", "is", etc. to reduce interference in semantic analysis.
[0064] Clean up emojis and emoticons: For social media data, emojis need to be removed or uniformly processed to avoid interference in sentiment analysis and topic classification.
[0065] Furthermore, perform multi-language tokenization on the data. Here, mainly perform tokenization skills processing on Chinese and English. The purpose of tokenization is to split the text into smaller language units (such as words, sub-words, etc.) for subsequent processing. Chinese tokenization: Chinese does not have spaces to separate words. Through the tokenization method of statistical models, divide the valid vocabulary in the text. Here, use the tokenization tool jieba to perform Chinese tokenization. English tokenization: English tokenization is relatively simple and is usually based on spaces and punctuation marks for splitting. Here, use the tool spaCy to implement this function.
[0066] Finally, perform lemmatization on the vocabulary. Lemmatization refers to restoring the vocabulary to its base form (such as the infinitive of a verb, the singular of a noun, etc.), thereby reducing the complexity brought by vocabulary changes and improving semantic understanding ability. Chinese lemmatization: The word forms of Chinese are standardized and normalized to a certain extent through techniques such as dictionary mapping and part-of-speech tagging, removing multiple variants of the vocabulary. Here, use jieba and HanLP to perform part-of-speech tagging and lemmatization on Chinese. English lemmatization, here use the Lemmatizer of spaCy to perform lemmatization on verbs and nouns, etc.
[0067] Thus, the preprocessing of multi-platform and multi-language data is completed.
[0068] Step 2, perform improved dynamic self-distillation semantic extraction processing based on BERT.
[0069] Extract document semantic features using the improved dynamic self-distillation algorithm based on BERT, which mainly includes two parts: one is the semantic feature extraction of the BERT model; the other is the improved dynamic self-distillation mechanism.
[0070] Semantic feature extraction of the BERT model:
[0071] The semantic feature extraction of BERT is achieved through the following process:
[0072] First, the input text is transformed into an embedding representation, which is encoded through three types of embeddings: one is token embeddings, representing the embedding vectors of each word or word segment, and segmented using the WordPiece technique; the second is segment embeddings, used to distinguish between input sentences A and B (for NLP tasks such as question answering, natural language inference, etc.); the third is position embeddings: used to represent the position of words in the sequence. Therefore, the input vector X of BERT = [x 1 , x 2 , …, x n is embedded as:
[0073] X = E token (X) + E segment (X) + E position (X) (1)
[0074] Among them, E token (X) represents token embeddings, E segment (X) represents segment embeddings, and E position represents position embeddings.
[0075] Furthermore, BERT uses the Transformer encoder structure to process the input text. The key part of the Transformer is the self-attention mechanism, which can calculate the context representation of each word based on the relevance of each position in the input sequence. Each Transformer layer in BERT consists of multiple self-attention sub-layers and a fully connected feed-forward network. In each layer, the calculation process of the self-attention mechanism is as follows:
[0076] Q = W Q ·X, K = W K ·X, V = W V ·X (2)
[0077] Among them, W Q , W K and W V are learned weight matrices, corresponding to the embeddings of query, key, and value respectively. The core of the self-attention mechanism is to calculate the similarity between the query and the key (i.e., the attention score):
[0078]
[0079] Among them,d k It is the dimension of the key, which is used to scale the similarity score.
[0080] Furthermore, after multi-layer Transformer encoding, BERT outputs the context representation of each input word. Suppose the input text is X = {x 1 , x 2 , …, x n}, and the context representation generated by BERT is H = [h 1 , h 2 , …, h n , where h i is the context representation of the word x i , that is, the word vector encoded by Transformer.
[0081] Finally, the semantic feature extraction of BERT is achieved through the hidden state output by the last layer of the model. In many NLP tasks, we usually use the vector corresponding to the [CLS] token (the representative at the sentence level) to represent the semantic features of the entire sentence. The input text is X = {x 1 , x 2 , …, x n}, the vector representation of the [CLS] token is h CLS , and it is used as the final semantic representation of this text:
[0082] X = h CLS = BERT(X) CLS (4)
[0083] where represents the semantic feature, n is the number of samples, d is the dimension of the semantic feature, which is the input of the dynamic distillation mechanism.
[0084] Through bidirectional context modeling and self-attention mechanism, BERT can effectively extract the semantic features of the input text. Its core lies in capturing the deep semantics of the text through the context representation of multiple layers of Transformer, thereby improving the performance in various NLP tasks.
[0085] Improved dynamic self-distillation mechanism optimizes semantic feature extraction:
[0086] To improve the performance of BERT in specific tasks, an improved dynamic self-distillation strategy is introduced, which specifically includes the following steps:
[0087] First, construct the target network and the student network. Target Network (Teacher Network): Train a target network on the pre-trained BERT model and obtain its output as the "teacher" signal. Let the output of the target network be T(X). Student Network (Student Network): Based on the BERT model, train a student network through the distillation method to make its output as close as possible to the representation of the target network. The output of the student network is S(X).
[0088] Furthermore, the core of improving the dynamic self-distillation strategy is to design a distillation dynamic adjustment mechanism. By observing the training loss of the model (such as cross-entropy loss) and the performance of the validation set, the h generated by BERT CLS is input as semantic features into the student model S(X) and the teacher model T(X). The output of the teacher model T(X) is used to guide the training of the student model. The loss consists of two parts: distillation loss and task loss. The distillation loss is calculated as follows:
[0089]
[0090] The task loss is:
[0091]
[0092] The total loss L total Combining the two, the weight is controlled by the dynamic adjustment factor DF t :
[0093] L total = DF t ·L distill +(1 - DF t )·L task (7)
[0094] where the dynamic adjustment distillation factor DF t has the following formula:
[0095]
[0096] where L_train(t) represents the training loss in the t-th round of training, L_val(t) represents the validation loss in the t-th round of training, DF_max and DF_min represent the maximum and minimum values of the distillation factor, and DF_t represents the distillation factor of the current round. This formula means that when the loss of the model on the training set is large and the loss of the validation set is small, the distillation factor increases and the distillation process becomes stronger; conversely, reducing the distillation factor can reduce the influence of distillation on the parameters.
[0097] In the treatment of the self-distillation loss function, in view of the problem that the traditional KL divergence formula is sensitive to outliers, this method proposes a new KL divergence formula, which combines the square term and the logarithmic term to improve the numerical stability and robustness. The self-distillation loss formula based on the improved KL divergence is as follows:
[0098] L(x i ) = λ·(KL improved (P(x i )||Q(x i ))) (9)
[0099] Among them, x i represents the input text, P(x i ) is the prediction distribution of the student model (current prediction step), Q(x i ) is the prediction distribution of the teacher model (past training step), λ is the dynamic distillation factor, used to control the weight of the distillation loss in the total loss, and KL improved (P(x i )||Q(x i )) is the improved KL divergence algorithm, and the specific formula is as follows:
[0100]
[0101] Among them, α and β are hyperparameters used to control the weight of the square term, and γ is a hyperparameter used to balance the weights of positive and negative samples. Adding -1 to the formula is mainly to improve the numerical stability and reduce the sensitivity of the KL divergence to outliers. Specifically, when the distributions of P(x i ) and Q(x i ) are very close, and P(x i ) / Q(x i ) is close to 1, (P(x i ) / Q(x i ) - 1) 2 will be very small, close to 0, which can effectively reduce the instability and floating-point errors in numerical calculations. When the distributions of P(x i ) and Q(x i ) are very different, P(x i ) / Q(x i ) will be very large, and (P(x i ) / Q(x i ) - 1) 2 limits the impact on the loss function within a reasonable range through -1, reducing the dominant role of outliers in the loss function.
[0102] The dynamic distillation factor λ in formula (9) is designed here as a dynamic weight related to the magnitude of the gradient of the loss function of the sample x i , and the calculation formula is as follows:
[0103]
[0104] Among them, α is a regulation factor used to control the intensity of the dynamic change of distillation, ▽ θ L total (x i ) is the gradient of the sample x i with respect to the total loss function.
[0105] So far, the result of extracting document semantic features by the improved dynamic self-distillation algorithm based on BERT is obtained n is the number of samples, and d is the dimension of semantic features.
[0106] Step 3, multi-platform semantic feature fusion processing based on the sparse principal component linear discriminant fusion algorithm:
[0107] The processing flow of the multi-platform semantic feature fusion algorithm based on Sparse PCA-LDA is as follows:
[0108] Suppose that each platform i has obtained a semantic feature matrix through the improved dynamic self-distillation semantic extraction algorithm based on BERT n i is the number of samples of platform i, and d i is the dimension of semantic features. Then, the semantic feature matrices of different platforms are concatenated into a matrix X, which is expressed as: where is the comprehensive dimension of features of all platforms. Then, the standard deviation is used to standardize the feature matrix X to obtain X', ensuring that each feature is in the same dimension. Then, the fusion matrix W final is obtained by jointly optimizing the Sparse PCA and LDA objective functions, and the joint optimization problem is expressed as:
[0109]
[0110] where W PCA is the principal component matrix obtained by Sparse PCA, and the form of the optimization problem is:
[0111]
[0112] where X' is the standardized input data matrix, with a size of n×d, is the data reconstruction error, W PCA is the sparse principal component matrix, with a size of d×m, where m is the number of selected principal components, ||W PCA || 1 = ∑ i,j |W PCA(i,j)is the L1 norm, which controls the sparsity of matrix W, penalizes the non-zero elements in matrix W, and ν is the regularization parameter used to adjust the balance between the reconstruction error and sparsity. Then, the coordinate descent method is used to solve this problem to obtain W PCA A sparse matrix used to represent the principal components of the data.
[0113] W in formula (5) LDA is the linear discriminant analysis matrix, such that in the projected space of X’, the between-class scatter matrix S B is maximized, and the within-class scatter matrix S W is minimized, that is, to maximize the ratio of the between-class scatter matrix S B to the within-class scatter matrix S W to achieve the best class separation. The optimization objective function is:
[0114]
[0115] where the between-class scatter matrix S B measures the difference between the mean of each class and the global mean, and the within-class scatter matrix S W measures the scatter degree of the samples within each class.
[0116] Finally, the final semantic feature fusion matrix W optimized during the feature fusion process final can be expressed as:
[0117]
[0118] Here, W is obtained by jointly optimizing Sparse PCA and LDA.
[0119] The optimized W final is directly used for feature fusion, converting the semantic features X of multiple platforms obtained in step two i into the fused feature X’ to provide input for subsequent classification or clustering tasks, while ensuring that the fused feature has better expressive ability and class discrimination ability.
[0120] So far, a multi-platform semantic feature fusion method based on a large language model has been completed.
[0121] The above is only a preferred specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A multi-platform semantic feature fusion method based on a large language model, characterized in that: include: Acquire data from different social media platforms, and preprocess the data from the different social media platforms to obtain preprocessed data sets; Constructing a BERT model, inputting the preprocessed data set into the BERT model for calculation, and obtaining a semantic representation data set; A teacher network is built based on the pre-trained BERT model, and the BERT model is processed based on the distillation method to obtain the student network; Inputting the semantic representation data set into the teacher network to obtain the teacher network output, optimizing the process of guiding the student model training by the teacher network output through a dynamic self-distillation mechanism, and obtaining a document semantic feature extraction model; Processing the data based on the document semantic feature extraction model to obtain a semantic feature matrix; A sparse principal component linear discriminant fusion algorithm is constructed, and the semantic feature matrix is input into the sparse principal component linear discriminant fusion algorithm for calculation to obtain a final semantic feature fusion matrix.
2. The method according to claim 1, characterized in that The process of preprocessing the data of the different social media platforms to obtain the preprocessed data set includes: Performing unified character encoding processing on the data of the different social media platforms, and standardizing the encoded data to obtain a standardized data set; Remove noise and non-semantic information from the standardized data set to obtain a cleaned data set; Performing multilingual word segmentation processing on the cleaned data set to obtain a word segmentation data set; The word segmentation dataset is subjected to lemma restoration processing to obtain the preprocessed dataset.
3. The method according to claim 1, characterized in that: The process of obtaining the semantic representation data set includes: Converting the preprocessed dataset into an embedding representation; Constructing a BERT model, and calculating the context representation of each word in the embedded representation through the BERT model; The semantic representation dataset is obtained based on the context representation of each word.
4. The method according to claim 1, characterized in that The expression of the total loss function for optimizing the process of guiding the student model training by the teacher network output through the dynamic self-distillation mechanism is: L total =DF t ·L distill +(1-DF t )·L task ; Among them, L total is the total loss, DF t To dynamically adjust the distillation factor, L task is the task loss, L distill is the distillation loss.
5. The method according to claim 4, characterized in that The expression for dynamically adjusting the distillation factor is: Among them, L train (t) represents the training loss during the tth round of training, L val (t) represents the validation loss during the tth round of training, DF max and DF min Indicates the maximum and minimum values of the distillation factor.
6. The method according to claim 1, characterized in that The process of obtaining the document semantic feature extraction model further includes: optimizing the dynamic self-distillation mechanism based on the KL divergence method to obtain an improved KL divergence self-distillation loss; The expression of the improved KL divergence self-distillation loss is: L(x i )=λ·(KL improved (P(x i )||Q(x i ))); Among them, x i represents the input text, P(x i ) is the predicted distribution of the student model, Q(x i ) is the prediction distribution of the teacher model, λ is the dynamic distillation factor, KL improved (P(x i )||Q(x i )) is an improved KL divergence algorithm.
7. The method according to claim 6, characterized in that The expression of the improved KL divergence algorithm is: Among them, α and β are hyperparameters, and γ is a hyperparameter.
8. The method according to claim 1, characterized in that: The expression of the final semantic feature fusion matrix is: Where X' is the normalized input data matrix, W final is the final semantic feature fusion matrix, W LDA is the linear discriminant analysis matrix, S B is the between-class scatter matrix, S W is the intra-class scatter matrix, ν is the regularization parameter, is the data reconstruction error.