A language sentiment analysis and labeling method for finance
By utilizing the BERT pre-trained model and partial text occlusion technology, a pre-defined public opinion analysis model was constructed, which solved the problems of large volume of public opinion text data and complex knowledge system, and achieved efficient and accurate sentiment analysis and annotation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHONGJINKE INFORMATION TECH CO LTD
- Filing Date
- 2023-03-22
- Publication Date
- 2026-04-24
AI Technical Summary
In existing technologies, public opinion text data is large in volume and has a complex knowledge system. Manual mining and analysis methods result in low analysis efficiency and insufficient accuracy.
Using the BERT pre-trained model as the basic framework, fine-tuning is performed using public opinion sample data to build a pre-set public opinion analysis model. During the model fine-tuning process, local text occlusion is performed to enhance the model's ability to learn contextual semantics and achieve automatic analysis of the sentiment polarity of public opinion texts.
It improves the efficiency and accuracy of analyzing massive amounts of public opinion text data, enhances the model's ability to learn contextual semantics, and enables more accurate sentiment classification and labeling.
Smart Images

Figure CN116578697B_ABST
Abstract
Claims
1. A language sentiment analysis and annotation method for finance, characterized in that, include: Obtain the public opinion text to be analyzed; The index corresponding to each character in the public opinion text is determined according to the preset text index table; Based on the index corresponding to each character, determine the input vector matrix corresponding to the public opinion text; The input vector matrix corresponding to the public opinion text is input into a preset public opinion analysis model for text sentiment classification to obtain the sentiment classification result corresponding to the public opinion text. During the fine-tuning training process of the preset public opinion analysis model, the public opinion training samples are partially occluded. Based on the sentiment classification results, the sentiment polarity of the public opinion text is labeled; The preset public opinion analysis model includes an enhanced semantic vector extraction model and a sentiment classification model. The enhanced semantic vector extraction model includes an attention layer. The step of inputting the input vector matrix corresponding to the public opinion text into the preset public opinion analysis model for text sentiment classification to obtain the sentiment classification result corresponding to the public opinion text includes: The input vector matrix is multiplied by the corresponding weight matrix of the attention layer to obtain the query vector matrix, key vector matrix, and value vector matrix corresponding to the input vector matrix; Calculate the attention matrix output by the attention layer based on the query vector matrix, the key vector matrix, and the value vector matrix; Based on the attention matrix, determine the enhanced semantic vector corresponding to each character in the public opinion text; The enhanced semantic vector corresponding to each character is input into the sentiment classification model to perform sentiment classification, thereby obtaining the sentiment classification result corresponding to the public opinion text. Specifically, before inputting the input vector matrix corresponding to the public opinion text into a preset public opinion analysis model for text sentiment classification to obtain the sentiment classification result corresponding to the public opinion text, the method further includes: Collect public opinion data samples; The public opinion data sample is preprocessed to obtain the preprocessed public opinion data sample; The public opinion training samples are determined based on the preprocessed public opinion data samples. Obtain an initial enhanced semantic vector extraction model and an initial sentiment classification model, wherein the initial enhanced semantic vector extraction model has been pre-trained; The input vector matrix corresponding to the public opinion training sample is input into the attention layer of the initial enhanced semantic vector extraction model for processing to obtain the initial attention matrix corresponding to the public opinion training sample; The initial attention matrix is adjusted to obtain an adjusted initial attention matrix; specifically, this includes: randomly determining occluded characters in the public opinion training samples; constructing an occlusion matrix based on the occluded and unoccluded characters, wherein the value at the position of the occluded character in the occlusion matrix is 1, and the value at the position of the unoccluded character is 0; and adjusting the initial attention matrix using the occlusion matrix to obtain the adjusted initial attention matrix. Based on the adjusted initial attention matrix, determine the initial enhanced semantic vector corresponding to each character in the public opinion training sample; The initial enhanced semantic vector corresponding to each character in the public opinion training sample is input into the initial sentiment classification model for sentiment classification, and the predicted sentiment classification result corresponding to the public opinion training sample is obtained. Based on the predicted sentiment classification result and the actual sentiment classification result corresponding to the public opinion training sample, the initial enhanced semantic vector extraction model and the initial sentiment classification model are jointly iteratively trained. The model iterative training process is repeated until the preset conditions are met, at which point the iterative training stops and the trained enhanced semantic vector extraction model and sentiment classification model are output. Based on the enhanced semantic vector extraction model and the sentiment classification model, the preset public opinion analysis model is determined.
2. The method according to claim 1, characterized in that, The step of determining the input vector matrix corresponding to the public opinion text based on the index corresponding to each character includes: Based on the index corresponding to each character, determine the character vector matrix corresponding to the public opinion text; Based on the position information of each character in the public opinion text, determine the position vector matrix corresponding to the public opinion text; Based on the sentence to which each character belongs in the public opinion text, determine the text vector matrix corresponding to the public opinion text; The input vector matrix corresponding to the public opinion text is obtained by adding the character vector matrix, the position vector matrix, and the text vector matrix.
3. A language sentiment analysis and annotation device for finance, used to implement the language sentiment analysis and annotation method for finance as described in any one of claims 1 to 2, characterized in that, include: The acquisition unit is used to acquire the public opinion text to be analyzed. The first determining unit is used to determine the index corresponding to each character in the public opinion text according to a preset text index table; The second determining unit is used to determine the input vector matrix corresponding to the public opinion text based on the index corresponding to each character; The analysis unit is used to input the input vector matrix corresponding to the public opinion text into a preset public opinion analysis model for text sentiment classification, and obtain the sentiment classification result corresponding to the public opinion text. In the fine-tuning training process of the preset public opinion analysis model, the public opinion training samples are partially occluded. The annotation unit is used to annotate the sentiment polarity of the public opinion text based on the sentiment classification results.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 2.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Target emovement analysis method, model training method, medium and equipment
CN111339255A