Information processing method, model training method, device, equipment, medium and product
By encoding multilingual news texts and generating fusion vectors through a multilingual model, efficient multi-dimensional public opinion analysis is achieved, solving the problems of low efficiency and insufficient accuracy in existing technologies and improving the timeliness and accuracy of public opinion analysis.
Patent Information
- Application Number
- CN202510784935.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-19
AI Technical Summary
Existing methods of obtaining and analyzing public opinion are inefficient and lack accuracy, especially in the analysis of multilingual news sources, which lack comprehensiveness and accuracy.
The first language model is used to encode the original multilingual news text to generate the original language vector, and the second language model is used to encode the translated target language text to generate the target language vector. By generating sentence-level and word-level fusion vectors, topic and sentiment classification is performed, and combined with entity extraction, multi-dimensional public opinion analysis results are generated.
The efficiency and accuracy of public opinion analysis have been improved. By combining multilingual model processing with timeliness and cross-language feature complementarity, the dependence on translation quality has been reduced, and the ability to capture public opinion correlations has been enhanced.
Smart Images

Figure CN120670671A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method for processing public opinion information, a public opinion information processing device, a model training method for public opinion information processing, a model training device for public opinion information processing, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Public opinion is a comprehensive reflection of the opinions, emotions, attitudes and behavioral tendencies expressed by the public in response to specific social events. By analyzing and understanding public opinion, we can accurately perceive environmental changes and opinions, and then respond to needs. However, the current methods of obtaining and analyzing public opinion still have the defects of inefficiency and lack of accuracy.
[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a public opinion information processing method, a processing device, a model training method for public opinion information processing, a training device, an electronic device, a storage medium and a computer program product, which can at least to some extent overcome the problems of low efficiency and insufficient accuracy in the acquisition and analysis methods of public opinion in related technologies.
[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0006] According to one aspect of the present disclosure, a method for processing public opinion information is provided, including: encoding a crawled original language text based on a first language model to obtain an original language vector, encoding a target language text based on a second language model to obtain a target language vector, wherein the target language text is generated by translating the original language text, and the first language model is a multilingual model; generating a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector; performing topic and sentiment classification based on the first fusion vector to obtain corresponding topic features and sentiment features; performing entity extraction on the second fusion vector to obtain entity features; and generating a public opinion analysis result based on the topic features, the sentiment features, and the entity features.
[0007] In one embodiment of the present disclosure, the target language vector includes a first target language vector and a second target language vector. The original language vector is obtained by encoding an original language text based on a first language model, and the target language vector is obtained by encoding the target language text based on a second language model. The method includes: encoding the original language text based on the first language model to obtain a first output sequence, and encoding the target language text based on the second language model to obtain a second output sequence; extracting hidden layer information corresponding to a global semantic tag in the first output sequence as the original language vector; extracting hidden layer information corresponding to the global semantic tag in the second output sequence as the first target language vector; and extracting hidden layer information corresponding to other tags in the second output sequence as the second target language vector. The second target language vector includes multiple first subvectors and second subvectors, where the first subvector is the hidden layer information corresponding to each subword in the target language text, and the second subvector is the hidden layer information corresponding to a delimiter tag in the second output sequence.
[0008] In one embodiment of the present disclosure, generating a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector includes: adding and fusing the original language vector with the first target language vector to obtain the first fusion vector; adding and fusing the original language vector with the multiple first sub-vectors and the second sub-vector respectively to obtain multiple third sub-vectors, and using the multiple third sub-vectors as the second fusion vector.
[0009] In one embodiment of the present disclosure, topic and emotion classification are performed respectively based on the first fusion vector to obtain corresponding topic features and emotion features, including: mapping the first fusion vector to the topic classification space based on a first linear layer to obtain a first logarithmic probability feature of the topic classification; outputting a first prediction probability for each topic category based on the first logarithmic probability feature by a first softmax function to determine the topic feature based on the first prediction probability; mapping the first fusion vector to the emotion classification space based on a second linear layer to obtain a second logarithmic probability feature of the emotion classification; outputting a second prediction probability for each emotion category based on the second logarithmic probability feature by a second softmax function to determine the emotion feature based on the second prediction probability.
[0010] In one embodiment of the present disclosure, entity extraction is performed on the second fusion vector to obtain entity features, including: mapping the multiple third sub-vectors to the label space respectively based on a third linear layer to obtain an initial label probability of each of the third sub-vectors; inputting the initial label probability into a conditional random field (CRF) layer as a state score, so that the CRF layer learns the transition scores between adjacent labels based on the state score, and performing Viterbi decoding based on the state score and the transition score to obtain a label sequence, so as to obtain the entity features based on the label sequence.
[0011] In one embodiment of the present disclosure, the sentiment classification includes positive, negative and neutral, and a public opinion analysis result is generated based on the topic feature, the sentiment feature and the entity feature, including: detecting that the topic feature belongs to the target topic and that the sentiment feature is negative, obtaining longitude and latitude information and / or matching administrative areas based on the geographic entity in the entity feature; performing a map marking operation of the public opinion location based on the longitude and latitude information, and / or dividing the public opinion influence range based on the matched administrative area as the public opinion analysis result.
[0012] In one embodiment of the present disclosure, a public opinion analysis result is generated based on the topic features, the sentiment features, and the entity features, including: detecting that the topic features belong to a target topic, extracting the subject entity, geographic entity, and event entity in the entity features; detecting that the subject entity includes institutional information, generating a business impact analysis result based on geographic association rules, subject association rules, and event and business association rules; and generating the public opinion analysis result based on the business impact analysis result and the sentiment features.
[0013] In one embodiment of the present disclosure, before generating the public opinion analysis result based on the topic features, the sentiment features and the entity features, it also includes: extracting and proposing multiple reference entities from the entity features; screening out a candidate set from a pre-stored database based on key entities among the multiple reference entities; performing entity matching based on the multiple reference entities and the candidate set to obtain a matching rate; obtaining a repeated tone determination result based on the relationship between the matching rate and the matching threshold, wherein, if the matching rate is detected to be greater than the matching threshold, it is determined to be repeated public opinion and the public opinion analysis result is not generated; if the matching rate is detected not to be greater than the matching threshold, the public opinion analysis result is generated based on the topic features, the sentiment features and the entity features.
[0014] According to another aspect of the present disclosure, a model training method for public opinion information processing is provided, comprising: preprocessing and performing three-task labeling operations on original historical language texts in multiple languages obtained by crawling to obtain a first set of training data; preprocessing and performing three-task labeling operations on target historical language texts obtained by translating the original historical language texts in the multiple languages to obtain a second set of training data, wherein the three-task labels include topic labels, sentiment labels, and entity labels; fine-tuning the output layer weights of a first pre-trained model and a second pre-trained model based on the first set of training data and the second set of training data to obtain a first language model and a second language model; and performing two-stage model training based on the first set of training data and the second set of training data to obtain a prediction model, wherein the prediction model is used to predict topics, sentiments, and entities for a fusion vector generated based on the original language text.
[0015] In one embodiment of the present disclosure, a two-stage model training is performed based on the first set of training data and the second set of training data to obtain a prediction model, including: in a first stage, based on the second set of training data, model training is performed on three original linear layers based on topic classification, sentiment analysis, and entity extraction, respectively, to obtain a first intermediate linear layer, a second intermediate linear layer, and a third intermediate linear layer, respectively, wherein the same weights are configured for the loss functions of the three original linear layers; in a second stage, the first set of training data and the second set of training data are fused to generate a first fused training vector and a second fused training vector; based on the first fused training vector and the second fused training vector, model training is performed on the first intermediate linear layer, the second intermediate linear layer, the third intermediate linear layer, and the original CRF layer, respectively, to obtain a first linear layer, a second linear layer, a third linear layer, and a CRF layer.
[0016] In one embodiment of the present disclosure, the original historical language texts in multiple languages obtained by crawling are preprocessed and subjected to three-task labeling operations to obtain a first set of training data, including: obtaining the original historical language texts in the multiple languages based on the configured news source; performing character-level cleaning and language normalization processing on the original historical language text to obtain preprocessed text; and performing text topic category labeling, sentiment polarity labeling, and entity type and boundary standards on the preprocessed text to obtain the first set of training data.
[0017] According to another aspect of the present disclosure, a public opinion information processing device is provided, including: an encoding module, configured to encode a crawled original language text based on a first language model to obtain an original language vector, and to encode a target language text based on a second language model to obtain a target language vector, wherein the target language text is generated by translating the original language text, and the first language model is a multilingual model; a fusion module, configured to generate a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector; a first prediction module, configured to perform topic and sentiment classification based on the first fusion vector, respectively, to obtain corresponding topic features and sentiment features; a second prediction module, configured to perform entity extraction on the second fusion vector, to obtain entity features; and an analysis module, configured to generate a public opinion analysis result based on the topic features, the sentiment features, and the entity features.
[0018] According to another aspect of the present disclosure, a model training device for public opinion information processing is provided, including: a first processing module for preprocessing and performing three-task labeling operations on original historical language texts in multiple languages obtained by crawling to obtain a first set of training data; a second processing module for preprocessing and performing three-task labeling operations on target historical language texts translated from the original historical language texts in the multiple languages to obtain a second set of training data, wherein the three-task labels include topic labels, sentiment labels, and entity labels; a fine-tuning module for fine-tuning the output layer weights of a first pre-trained model and a second pre-trained model based on the first set of training data and the second set of training data to obtain a first language model and a second language model; a training module for performing two-stage model training based on the first set of training data and the second set of training data to obtain a prediction model, wherein the prediction model is used to predict topics, sentiments, and entities for a fusion vector generated based on the original language text.
[0019] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; the processor is configured to execute the public opinion information processing method of the first aspect and the model training method for public opinion information processing of the second aspect by executing the executable instructions.
[0020] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the above-mentioned public opinion information processing method or the model training method for public opinion information processing.
[0021] According to another aspect of the present disclosure, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the above-mentioned public opinion information processing method or the model training method for public opinion information processing.
[0022] The public opinion information processing solution provided by the embodiments of the present disclosure encodes the original multilingual news text using a first language model to obtain an original language vector, encodes the translated target language text using a second language model to obtain a target language vector, and generates a first fusion vector focusing on sentence-level global semantics and a second fusion vector retaining word-level fine-grained information based on the two types of vectors, thereby respectively realizing topic and sentiment classification and entity extraction, and finally generating public opinion analysis results by integrating multi-dimensional features. Directly processing the original text through the multilingual model ensures the timeliness of processing, and combining bimodal feature fusion significantly improves processing efficiency. The generation of task-oriented fusion features is conducive to optimizing classification and entity extraction accuracy, and cross-language feature complementarity reduces dependence on translation quality. Multi-feature joint analysis enhances the ability to capture public opinion associations, thereby improving the efficiency and accuracy of public opinion analysis.
[0023] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0025] Figure 1 A schematic diagram showing a public opinion information processing system according to an embodiment of the present disclosure is shown;
[0026] Figure 2 A flowchart of a method for processing public opinion information in an embodiment of the present disclosure is shown;
[0027] Figure 3 A schematic diagram showing a language model in an embodiment of the present disclosure;
[0028] Figure 4 A flowchart of another method for processing public opinion information in an embodiment of the present disclosure is shown;
[0029] Figure 5 A schematic diagram showing another public opinion information processing system according to an embodiment of the present disclosure;
[0030] Figure 6 A flow chart of another method for processing public opinion information according to an embodiment of the present disclosure is shown;
[0031] Figure 7 A flow chart of a model training method for public opinion information processing according to an embodiment of the present disclosure is shown;
[0032] Figure 8 A flow chart of another model training method for public opinion information processing according to an embodiment of the present disclosure is shown;
[0033] Figure 9 A schematic diagram of a public opinion information processing device according to an embodiment of the present disclosure is shown;
[0034] Figure 10 A schematic diagram of a model training device for public opinion information processing according to an embodiment of the present disclosure is shown;
[0035] Figure 11 A structural block diagram of a computer device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0037] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0038] Existing public opinion acquisition technology mainly uses web crawler technology to crawl news data from relevant sites, and then filters the news data according to predefined rules to obtain relevant public opinion data. This method relies on a complex and accurate rule library. Rule formulation requires a lot of time and manpower. It does not have the ability to understand the semantics of the entire text and cannot accurately determine whether the text belongs to public opinion data. It is ineffective and has poor results.
[0039] In addition, traditional news and public opinion analysis methods usually rely on text processing technology in a single language. Single-language methods are limited by specific language analysis tools, and the breadth of public opinion analysis is insufficient, making it difficult to cover news sources in multiple languages, resulting in a lack of comprehensiveness in the analysis results.
[0040] In the present disclosure, the original multilingual news text is encoded using a first language model to obtain an original language vector, and the translated target language text is encoded using a second language model to obtain a target language vector. Based on the two types of vectors, a first fusion vector focusing on sentence-level global semantics and a second fusion vector retaining word-level fine-grained information are generated, thereby respectively realizing topic and sentiment classification and entity extraction, and finally integrating multi-dimensional features to generate public opinion analysis results. Directly processing the original text through the multilingual model ensures the timeliness of processing, and combining bimodal feature fusion significantly improves processing efficiency. The generation of task-oriented fusion features is conducive to optimizing classification and entity extraction accuracy, and cross-language feature complementarity reduces dependence on translation quality. Multi-feature joint analysis enhances the ability to capture public opinion correlations, thereby improving the efficiency and accuracy of public opinion analysis.
[0041] Figure 1 1 is a schematic diagram of the structure of a public opinion information processing system provided by an exemplary embodiment of the present application. The system includes: a plurality of user terminals 120 and a server terminal 140, wherein the user terminal 120 sends a user question and the server terminal 140 responds.
[0042] The user end 120 can be a mobile terminal such as a mobile phone, a game console, a tablet computer, an e-book reader, smart glasses, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a smart home device, an AR (Augmented Reality) device, a VR (Virtual Reality) device, etc., or the user end 120 can also be a personal computer (PC), such as a laptop computer and a desktop computer.
[0043] Among them, the user terminal 120 can be installed with an application for providing public opinion information processing.
[0044] The client 120 and the server 140 are connected via a communication network. Optionally, the communication network is a wired network or a wireless network.
[0045] Server 140 is a server, or a combination of multiple servers, a virtualization platform, or a cloud computing service center. Server 140 provides backend services for applications that process public opinion information. Optionally, server 140 performs primary computing tasks, while client 120 performs secondary computing tasks. Alternatively, server 140 performs secondary computing tasks, while client 120 performs primary computing tasks. Alternatively, client 120 and server 140 utilize a distributed computing architecture for collaborative computing.
[0046] In some optional embodiments, the server 140 is used to perform training program information of a model for processing public opinion information.
[0047] Optionally, the logistics client of the application installed on different user terminals 120 is the same, or the logistics client of the application installed on two user terminals 120 is the logistics client of the same type of application on different control system platforms. Based on the different terminal platforms, the specific form of the logistics client of the application can also vary. For example, the logistics client of the application can be a mobile phone logistics client, a PC logistics client, or a World Wide Web (Web) logistics client.
[0048] Those skilled in the art will appreciate that the number of the user terminals 120 may be greater or less. For example, there may be only one terminal, or there may be dozens, hundreds, or even more terminals. The embodiment of the present application does not limit the number and device type of the terminals.
[0049] Optionally, the system may further include a management device ( Figure 1 (not shown), the management device is connected to the server 140 via a communication network. Optionally, the communication network is a wired network or a wireless network.
[0050] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0051] To facilitate understanding, several terms involved in this application are first explained below.
[0052] XLM-RoBERTa (Cross-lingual Language Model – RoBERTa, cross-lingual pre-training model): is a multilingual version improved on the basis of the RoBERTa model, designed to solve cross-lingual tasks.
[0053] BERT (Bidirectional Encoder Representations from Transformers, bidirectional pre-trained language model): uses the Transformer encoder architecture, typically 12 layers (Base version) or 24 layers (Large version).
[0054] Below, each step of the public opinion information processing method in this example implementation will be described in more detail with reference to the accompanying drawings and examples.
[0055] Figure 2 A flow chart of a method for processing public opinion information in an embodiment of the present disclosure is shown.
[0056] like Figure 2 As shown, a method for processing public opinion information according to an embodiment of the present disclosure includes:
[0057] Step S202: Encode the crawled original language text based on the first language model to obtain an original language vector, and encode the target language text based on the second language model to obtain a target language vector. The target language text is generated by translating the original language text. The first language model is a multilingual model.
[0058] In some embodiments, the multilingual pre-trained model based on XLM-RoBERTa supports cross-lingual representation learning. By sharing the Transformer encoder, it can process multiple languages (such as English, Chinese, Russian, etc.) at the same time, and map texts in different languages to the same semantic space. The second language model is based on the single language (Chinese) pre-trained model of BERT, focusing on Chinese semantic understanding. Compared with the multilingual model, it is better at capturing Chinese grammar and contextual information.
[0059] like Figure 3 As shown, the first language model encodes the original language text, and the second language model encodes the target language text, and the outputs are both vector sequences, i.e., h[CLS], h1~h n , h[SEP], the original language vector and the target language vector are the vectors corresponding to the specified positions in the vector sequence, that is, a certain h, or several h.
[0060] Step S204 : generating a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector.
[0061] In some embodiments, both the first fusion vector and the second fusion vector can be obtained by fusing the original language vector and the target language vector through addition or attention mechanism. The difference between the first fusion vector and the second fusion vector is that the first fusion vector is obtained by fusing the original language vector with the first part of the vector sequence output by the second language model, and the second fusion vector is obtained by fusing the original language vector with the second part of the vector sequence output by the second language model.
[0062] In some embodiments, topic and sentiment analysis rely on the overall semantics of the text, and the first fusion vector is fused based on the relevant vectors in the two language vectors. Entity extraction requires precise positioning of the category of each token, and the second fusion vector is fused based on the relevant vectors in the two language vectors.
[0063] Step S206 , performing topic and sentiment classification based on the first fusion vector to obtain corresponding topic features and sentiment features.
[0064] In some embodiments, the first fused vector is a sentence-level global semantic representation that encodes the overall semantics of the text and is suitable for coarse-grained classification tasks.
[0065] Step S208: performing entity extraction on the second fusion vector to obtain entity features.
[0066] In some embodiments, the second fused vector is a word-level sequence semantic representation (including a hidden layer vector and a sentence boundary [SEP] vector for each subword), which retains the fine-grained structure of the text and is suitable for sequence labeling tasks.
[0067] Step S210: Generate public opinion analysis results based on topic features, sentiment features, and entity features.
[0068] In some embodiments, topic features, sentiment features (such as "negative"), and entity features (such as "foreign") can be input into a decision module to generate public opinion analysis results through rules or machine learning models.
[0069] In this embodiment, by using a first language model to encode the original multilingual news text to obtain an original language vector, using a second language model to encode the translated target language text to obtain a target language vector, and generating a first fusion vector that focuses on the global semantics at the sentence level and a second fusion vector that retains the fine-grained information at the word level based on the two types of vectors, and then respectively implementing topic and sentiment classification, entity extraction, and finally generating an opinion analysis result by comprehensively integrating multi-dimensional features. Directly processing the original text through the multilingual model ensures the timeliness of processing, and the combination of dual-modal feature fusion significantly improves the processing efficiency. The generation of task-oriented fusion features is beneficial to optimizing the accuracy of classification and entity extraction, and the cross-language feature complementarity reduces the dependence on translation quality. The joint analysis of multi-features enhances the ability to capture opinion associations, thereby improving the efficiency and accuracy of opinion analysis.
[0070] In an embodiment of the present disclosure, the target language vector includes a first target language vector and a second target language vector. Encoding the original language text based on the first language model to obtain an original language vector, and encoding the target language text based on the second language model to obtain a target language vector, including:
[0071] Encoding the original language text based on the first language model to obtain a first output sequence, and encoding the target language text based on the second language model to obtain a second output sequence.
[0072] As Figure 3 shown, in some embodiments, the output sequence includes:
[0073] (1) The hidden layer information of [CLS] (index 0).
[0074] Located at the first position (index 0) of the output sequence, it is the condensation of global semantics and is used for sentence-level tasks (such as classification, reasoning).
[0075] For example: The [CLS] vector of the Chinese sentence "This movie is very wonderful" encodes the overall semantics of "positive movie evaluation". <T
[0076] (2) The hidden layer information of Chinese subwords (indexes 1 to n).
[0077] Corresponding to the Chinese subwords after word segmentation in the input sequence (e.g., "natural language processing" → "zi", "ran", "yu", "yan", "chu", "li"), the hidden layer state of each subword contains its local semantics and context information, and is used for sequence annotation tasks and reading comprehension tasks. [[ID=!27]] <0000!78>(3) The hidden layer information of [SEP] (index n + 1).
[0079] Located at the end of the sequence, it is used to separate sentences (such as [CLS] sentence 1 [SEP] sentence 2 [SEP] when inputting sentence pairs), and its hidden layer state contains the semantics of sentence boundaries or relationships.
[0080] Extract the hidden layer information corresponding to the global semantic token of the first output sequence as the original language vector h[cls]_f.
[0081] Extract the hidden layer information corresponding to the global semantic token in the second output sequence as the first target language vector h[cls]_c.
[0082] In some embodiments, topic and sentiment analysis rely on the overall semantics of the text. The [CLS] vector, as a sentence-level representation, is suitable for such tasks. The original language vector (XLM-RoBERTa) provides cross-language generalization ability (such as understanding the semantic association between the English "trade" and the Chinese "trade"), while the first target language vector (BERT) enhances the Chinese expression accuracy (such as distinguishing the sentiment difference between "support" and "oppose").
[0083] In some embodiments, entity extraction requires precise positioning of the category of each Token. The first sub-vector (word-level hidden layer information) in the second fusion vector directly corresponds to each sub-word, providing the fine-grained features required for entity boundary recognition.
[0084] Extract the hidden layer information corresponding to other tokens in the second output sequence as the second target language vectors h1_c, h2_c,..., hn_c, h[sep]_c. The second target language vectors include multiple first sub-vectors and second sub-vectors. The first sub-vector is the hidden layer information corresponding to each sub-word in the target language text, and the second sub-vector is the hidden layer information corresponding to the separator token in the second output sequence.
[0085] In this embodiment, the first output sequence is obtained by encoding the original language text through the first language model, and the second output sequence is obtained by encoding the translated target language text through the second language model. The hidden layer information corresponding to the global semantic tokens is extracted from the two output sequences respectively as the original language vector and the first target language vector. At the same time, the hidden layer information corresponding to other tokens in the second output sequence is split into the first sub-vector containing the sub-word hidden layer information and the second sub-vector containing the separator token hidden layer information to form the second target language vector. By separating the sentence-level and word-level semantic representations, the refined adaptation of multi-task features is achieved, which not only retains the cross-language semantic association ability of the original language but also strengthens the fine-grained semantic expression of the Chinese translation text, enabling the model to capture more accurate global semantic information in the topic and sentiment classification tasks, precisely locate the entity boundaries and types in the entity extraction task, systematically improve the accuracy and robustness of multi-language public opinion analysis, and enhance the accuracy in cross-language semantic alignment and entity recognition.
[0086] In one embodiment of the present disclosure, generating a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector includes:
[0087] The original language vector is added to the first target language vector to obtain a first fused vector.
[0088] In some embodiments, the original language vector (XLM-RoBERTa's [CLS] vector) and the first target language vector (BERT's [CLS] vector) are both sentence-level representations. Although they come from different models, they have been mapped to similar semantic spaces through the pre-training mechanism. The original language vector retains the grammatical and cultural features of the source language (such as metaphorical expressions in English), while the first target language vector enhances the semantic accuracy in the Chinese context (such as the subtle differences in Chinese sentiment words). The addition operation realizes the direct fusion of the two features, forming a global semantic representation that takes into account both cross-language generalization and single-language accuracy.
[0089] The original language vector is added and fused with the multiple first sub-vectors and the second sub-vector respectively to obtain multiple third sub-vectors, and the multiple third sub-vectors are used as the second fused vector.
[0090] In some embodiments, the first sub-vector (BERT's word-level encoding of the translated text) provides fine-grained entity boundary information, while the cross-language features in the original language vector can align multilingual entity expressions, and achieve cross-language enhancement of entity features through addition operations. The second sub-vector (hidden layer information of the [SEP] tag) explicitly encodes sentence boundaries and paragraph structure. After being fused with the original language vector, it helps the model learn the dependency between entities and context. The generated multiple third sub-vectors (each corresponding to a word position) constitute a word-level sequence that retains cross-language information, directly adapting to the input requirements of the CRF layer and improving the accuracy of entity extraction.
[0091] In this embodiment, the first fusion vector effectively combines the cross-language understanding ability of the multilingual model with the precise semantic expression of the monolingual model by adding sentence-level features, enabling topic and sentiment classification to simultaneously process multiple language inputs and maintain stable performance in low-resource language scenarios, thereby preventing semantic loss caused by translation errors. The second fusion vector achieves precise alignment and positioning of multilingual entities by fusing word-level features with structural information. The collaborative design of the two fusion vectors enables the model to simultaneously optimize classification and sequence labeling tasks under a single architecture, thereby improving analysis efficiency.
[0092] like Figure 4 As shown, in one embodiment of the present disclosure, topic and sentiment classification are performed based on the first fusion vector to obtain corresponding topic features and sentiment features, including:
[0093] Step S402: Map the first fusion vector to the topic classification space based on the first linear layer to obtain the first logarithmic probability feature of the topic classification.
[0094] In some embodiments, the first linear layer (weight matrix W1) projects the high-dimensional semantic vector into the topic category space (e.g., 10 categories), i.e., logits1=W1·first fusion vector+b1, where b1 is a bias term.
[0095] In step S404, a first softmax function is used to output a first predicted probability for each topic category based on the first logarithmic probability feature, so as to determine the topic feature based on the first predicted probability.
[0096] In some embodiments, the first softmax function converts the raw scores (logits) output by the linear layer into a probability distribution.
[0097] Step S406: Map the first fusion vector to the sentiment classification space based on the second linear layer to obtain a second logarithmic probability feature of sentiment classification.
[0098] In some embodiments, although the topic and sentiment are both based on the first fusion vector, feature maps of different tasks are learned separately through independent linear layers (W2 and b2) to prevent interference between tasks. The linear transformation maps the semantic features to the sentiment polarity space (such as positive / negative / neutral), so that the model can capture the emotional tendency of the text.
[0099] Step S408: The second softmax function outputs a second predicted probability for each emotion category based on the second logarithmic probability feature, so as to determine the emotion feature based on the second predicted probability.
[0100] In some embodiments, the second softmax function works similarly to the first softmax function, converting the sentiment score into a probability distribution.
[0101] In this embodiment, through the task-decoupled linear mapping mechanism, the model can learn discriminative features for topic and sentiment classification tasks respectively. The independent linear layer maps the first fusion vector to the topic and sentiment classification space respectively, and combines the softmax function to output the probability distribution to determine the corresponding features, which can prevent interference between tasks and significantly improve the accuracy of classification. In addition, the simple and efficient structure of the linear layer and softmax ensures that the model can quickly process large-scale data and meet real-time analysis needs. The design based on multilingual fusion vectors enables the model to have cross-language semantic processing capabilities, flexibly supports topic and sentiment classification in multiple languages, enhances the generalization ability of public opinion in different languages, and improves the reliability and practicality of multi-dimensional public opinion analysis.
[0102] In one embodiment of the present disclosure, entity extraction is performed on the second fusion vector to obtain entity features, including:
[0103] Based on the third linear layer, multiple third sub-vectors are respectively mapped to the label space to obtain an initial label probability for each third sub-vector; the initial label probability is input into the conditional random field (CRF) layer as a state score, so that the CRF layer learns the transition scores between adjacent labels based on the state score, and performs Viterbi decoding based on the state score and the transition score to obtain a label sequence, so as to obtain entity features based on the label sequence.
[0104] In some embodiments, the third linear layer performs a linear transformation on each third sub-vector (word-level semantic vector) in the second fusion vector, projects it into the entity label space, generates an initial label probability (state score) corresponding to each word, and realizes a preliminary mapping from semantic features to the label space. After receiving the state score, the CRF layer learns the transition probability between entity labels and captures the long-distance dependency between labels. The Viterbi algorithm searches for the global optimal path in the label sequence space based on the state score and the transition score, and outputs the final entity label sequence.
[0105] In this embodiment, the word-level semantic vector in the second fusion vector is mapped to the entity label space through the third linear layer to generate the initial label probability, and then the CRF layer is used to learn the transfer scores between adjacent labels, and the entity label sequence is generated through Viterbi decoding to extract entity features. With the help of fine-grained utilization of word-level semantic information and sentence structure features, the entity boundaries and types are accurately located, and the label dependency relationship is modeled by the CRF layer to prevent the generation of invalid label sequences, thereby enhancing the coherence and rationality of entity recognition. At the same time, the fusion of the original language vector and the translated text vector is used to achieve the alignment of multilingual entity expressions, improve the uniformity of entity recognition in cross-language scenarios, and thus help improve the accuracy of entity recognition in public opinion analysis.
[0106] like Figure 5 As shown, a multilingual international public opinion analysis model according to the present disclosure includes a feature extraction module 502 , a feature fusion module 504 and a prediction module 506 .
[0107] The feature extraction module 506 performs feature extraction on the crawled original language news text and the news text translated into Chinese. The module input is the crawled original language news text and the Chinese news text obtained by translation using the translation API. The module uses two independent models to encode the two text sequences. BERT and XLM-RoBERTa are pre-trained natural language processing models that use the masked language modeling training method. The XLM-RoBERTa model is used to encode the original language news text, and the BERT model is used to encode the news text translated into Chinese. Both BERT and XLM-RoBERTa are pre-trained language models based on Transform. Among them, the cross-language pre-training model XLM-RoBERTa relies on the masked language model objective function and can process texts in 100 different languages.
[0108] The feature fusion module 504 is used to output a first fusion vector and a second fusion vector. For the original language text, the hidden layer information h[cls]_f of the first special tag [CLS] in the output sequence of the XLM-RoBERTa model is extracted as the original language vector. For Chinese text, the hidden layer information h[cls]_c of the first special tag [CLS] in the output sequence of the BERT model is extracted as the first target language vector. Other tags h1_c, h2_c, ..., hn_c, h[sep]_c in the output sequence of the BERT model are extracted as the second target language vector. The original language vector h[cls]_f is added to the first target language vector h[cls]_c to obtain The first fusion vector h[cls] will be used as the input feature vector of the subsequent news classification and sentiment analysis prediction layer, h[cls] = Add(h[cls]_f,h[cls]_c). The original language vector h[cls]_f is added to the second target language vector h1_c,h2_c,…,hn_c,h[sep]_c respectively to obtain the fused word-level information, that is, the second fusion vector (h1,h2,…,hn,h[sep]), which is used as the input feature vector of the subsequent entity extraction prediction layer, (h1,h2,…,hn,h[sep]) = Add(h[cls]_f,(h1_c,h2_c,…,hn_c,h[sep]_c)).
[0109] The prediction module 506 is used to input h[cls] into a news classifier composed of a first linear layer, and then perform probability prediction through a first softmax function, input h[cls] into a sentiment analyzer composed of a second linear layer, and then perform probability prediction through a second softmax function, input h1, h2,…, hn, h[sep] into an entity extractor composed of a third linear layer and a CRF (Conditional Random Field) layer, and then perform probability prediction through a third softmax function.
[0110] In one embodiment of the present disclosure, sentiment classification includes positive, negative, and neutral, and public opinion analysis results are generated based on topic features, sentiment features, and entity features, including:
[0111] It is detected that the topic feature belongs to the target topic and the sentiment feature is negative, and latitude and longitude information is obtained and / or administrative area matching is performed based on the geographic entity in the entity feature.
[0112] In some embodiments, it is detected that the topic feature belongs to the target topic, indicating that it belongs to the public opinion to be analyzed, and negative emotions are identified in combination with the emotional features, indicating that focused analysis is required.
[0113] In some embodiments, geographic entities (such as "Zhengzhou City" and "Yangtze River Basin") are extracted from entity features, and the geographic descriptions in the text are converted into computable spatial data (such as latitude and longitude coordinates, administrative division codes) through geocoding technology (such as address conversion to longitude and latitude) or administrative area matching rules (such as provincial / municipal boundary library).
[0114] Perform a map marking operation of the public opinion location based on the latitude and longitude information, and / or divide the public opinion influence range based on the matching administrative areas as the public opinion analysis result.
[0115] In some embodiments, the location of public opinion is marked on a map using latitude and longitude information, or the scope of influence is divided according to administrative regions (such as counting disaster-stricken areas by city-level administrative regions), and the geographical scope of public opinion is quantified through spatial analysis (such as buffer zone analysis).
[0116] In some embodiments, it is detected that the topic feature belongs to the target topic, indicating that it is international news data, sentiment analysis is performed on the international news data, news with negative sentiment tendencies are screened out, geographic information is obtained from the international public opinion data, longitude and latitude information is obtained, and dot operations are performed on the GIS map; geographic information is obtained from the international public opinion data, the administrative level is obtained, the influence range radius is set according to the administrative level and the area of the region, and the influence range area is circled.
[0117] In this embodiment, through the linkage analysis of topic-sentiment-geographic features, it is shown in turn that the text belongs to public opinion and has negative emotions that need to be analyzed. Further analysis is performed based on geographical features, and the geographical location and influence range of public opinion are intuitively displayed through map visualization, providing a geographical basis for decision-making, forming a complete analysis link from text semantics to spatial location, enriching the dimensions of public opinion analysis, and being particularly suitable for scenarios that require geographic spatial response, significantly improving the timeliness, spatial accuracy and visualization capabilities of public opinion analysis.
[0118] In one embodiment of the present disclosure, generating public opinion analysis results based on topic features, sentiment features, and entity features includes:
[0119] It is detected that the topic feature belongs to the target topic, and the subject entity, geographic entity and event entity in the entity feature are extracted.
[0120] In some embodiments, when the topic feature matches the target topic, it indicates that it belongs to the public opinion to be analyzed, and the subject entity, geographic entity, and event entity in the entity feature are extracted to construct a structured triple.
[0121] The subject entities detected include organization information, and business impact analysis results are generated based on geographic association rules, subject association rules, and event and business association rules.
[0122] In some embodiments, geographic association rules refer to matching pre-stored business layout data based on geographic entities to calculate the impact range of events. Subject association rules refer to identifying the related parties of subject entities through knowledge graphs and deducing the event diffusion path. Event-business association rules refer to mapping event entities to business indicators and evaluating the degree of impact in combination with historical data.
[0123] Generate public opinion analysis results based on business impact analysis results and sentiment characteristics.
[0124] In some embodiments, when using the multilingual international public opinion analysis model to perform international public opinion analysis, news crawling is performed based on pre-configured data sources to obtain original foreign news data; international news is classified and matched with pre-defined international public opinion definitions, and news data is screened; entity extraction is performed on international news, and the extracted entities are compared with the entities corresponding to the existing public opinion news texts in the database. When the entity matching rate reaches a pre-set threshold, the news is judged to be duplicate public opinion, and the public opinion data is deduplicated; entity extraction is performed on international news, and the obtained public opinion information is associated with the business of the international business agency in combination with the business-related resource information of the international business agency, and the scope of impact of the relevant international public opinion on the business of the international business agency is judged.
[0125] In this embodiment, by detecting the target topic, subject, geographic and event entities are extracted from entity features. For the subject entity containing institutional information, business impact analysis is performed by combining geographic association, subject association and event and business association rules. By combining structured entity extraction with multi-dimensional association rules, unstructured text is converted into quantifiable business impact indicators, realizing the transformation of public opinion from "event description" to "commercial value assessment". Geographic association rules accurately locate the spatial scope of public opinion impact, subject association rules explore the upstream and downstream transmission paths of events, and event and business association rules quantify the volatility risk of specific business indicators. Emotional characteristics are then combined to judge the market sentiment intensity of public opinion, which is conducive to improving the support ability of public opinion analysis for actual business decision-making.
[0126] In one embodiment of the present disclosure, before generating the public opinion analysis results based on the topic features, sentiment features, and entity features, the method further includes:
[0127] Extract and propose multiple reference entities from entity features; screen out a candidate set from a pre-stored database based on key entities among the multiple reference entities; perform entity matching with the candidate set based on the multiple reference entities to obtain a matching rate; obtain a repeated tone determination result based on the relationship between the matching rate and the matching threshold, wherein, if a matching rate greater than the matching threshold is detected, it is determined to be repeated public opinion and no public opinion analysis result is generated; if a matching rate not greater than the matching threshold is detected, a public opinion analysis result is generated based on the topic features, sentiment features and entity features.
[0128] In some embodiments, key entities are screened from entity features to construct a reference entity set. In a pre-stored database (such as the public opinion database of the past 24 hours), candidate documents containing any reference entity are quickly retrieved through an inverted index, and the degree of overlap between the reference entity and the candidate document is calculated. If it exceeds a threshold, it is determined to be a duplicate public opinion.
[0129] In this embodiment, before generating the public opinion analysis results, reference entities are first extracted from the entity features, candidate sets are screened based on key entities and entity matching is performed, and duplicate public opinions are determined by comparing the matching rate with the threshold to avoid generating analysis results for duplicate content. Through the entity-level precise matching mechanism, duplicate reports of the same event are automatically filtered out, reducing resource consumption of invalid analysis tasks and improving system processing efficiency. The candidate set screening of the pre-stored database is combined with the key entity weight calculation to ensure accurate identification of high-similarity public opinions. The matching rate determination rule avoids the subjectivity and lag of manual screening, allowing the model to focus on truly novel public opinion content, quickly locate effective information in massive data, reduce the cost of duplicate labor, and at the same time improve the purity of input data for subsequent modules such as topic classification and sentiment analysis, thereby enhancing the timeliness and accuracy of the entire public opinion analysis system.
[0130] In some embodiments, taking public opinion analysis based on news as an example, topic classification, that is, news classification task, pre-defines international public opinion in combination with specific business, that is, defines what kind of news is classified as international public opinion, such as defining the five categories of "war", "conflict", "crisis", "financial events" and "natural disasters" as the first-level classification that meets the definition of international public opinion.
[0131] Sentiment classification, that is, sentiment analysis tasks, is used to assist system users in identifying and locating potential risks posed by international events to international business organizations, that is, to discover public opinion data that is relevant to the business of international business organizations and may have a negative impact. In this disclosure, three types of public opinion sentiment tendencies are defined: "positive (active)", "negative (passive)", and "neutral".
[0132] For entity extraction tasks, it is necessary to determine the scope of business impact suspected to be caused by international public opinion and complete public opinion deduplication based on the public opinion geographical area information and equipment failure conditions. In this disclosure, entities such as "time", "place", "country", "event", "person", and "organization" are extracted to combine the business-related resource information of international business institutions to obtain public opinion information and associate it with the business of international business institutions, and determine the scope of the impact of relevant international public opinion on the business of international business institutions.
[0133] In some embodiments, a public opinion deduplication operation can be performed based on the above-mentioned entities, and entities can be extracted from the newly crawled public opinion news text. The extracted entities are then compared with the entities corresponding to the existing public opinion news texts in the database. When the matching rate of the entities reaches a preset threshold, the news is judged to be duplicate public opinion, and the weight coefficient is increased by one.
[0134] like Figure 6 As shown, according to another embodiment of the present disclosure, a method for processing public opinion information includes:
[0135] Step S602: crawl news based on pre-configured data sources to obtain original international news data.
[0136] Step S604: classify international news, match it with predefined international public opinion definitions, and filter the news data.
[0137] In step S606, entities are extracted from international news, and the extracted entities are compared with the entities corresponding to the existing public opinion news texts in the database. When the matching rate of the entities reaches a preset threshold, the news is judged to be duplicate public opinion, and the public opinion data is deduplicated.
[0138] Step S608: extract entities from international news, and associate the acquired public opinion information with the business of the international business organization in combination with the business-related resource information of the international business organization to determine the impact of the relevant international public opinion on the business of the international business organization.
[0139] Step S610: Perform sentiment analysis on the international news data to filter out news with negative sentiment tendencies.
[0140] Step S612: Obtain geographic information of international public opinion data, obtain latitude and longitude information, and perform dot operations on the GIS map.
[0141] Step S614, obtain geographic information of international public opinion data, obtain the administrative division level, set the influence range radius according to the administrative division level, that is, the corresponding regional area, and circle the influence range area.
[0142] In this embodiment, international news can be understood as foreign language news. By crawling news portal sites, obtaining original foreign news data, and classifying international news, the news data type can be screened. Furthermore, entity extraction is performed on international news, and duplication processing is performed on public opinion data. Entity extraction is performed on international news to determine the scope of business impact of relevant public opinion on relevant business organizations. Sentiment analysis is performed on international news data to screen out news with negative sentiment tendencies. Geographic information is obtained from international public opinion data to obtain latitude and longitude information, and dot operations are performed on the GIS map.
[0143] Furthermore, geographic information of international public opinion data is obtained, the administrative level is obtained, the influence radius is set according to the administrative level and the area of the region, the influence area is circled, and semantic analysis such as news classification, entity extraction, and sentiment analysis is performed on news data based on a multilingual model. This solves the limitations of public opinion judgment based on rules, such as the lack of effectiveness, individual uncertainty, non-exhaustive rules, and lack of full-text comprehension.
[0144] like Figure 7 As shown, a model training method for public opinion information processing according to an embodiment of the present disclosure includes:
[0145] Step S702 : Preprocessing and three-task labeling operations are performed on the crawled original historical language texts in multiple languages to obtain a first set of training data.
[0146] Step S704 , preprocessing and three-task labeling operations are performed on the target historical language text obtained by translating the original historical language texts in multiple languages to obtain a second set of training data, where the three-task labels include topic labels, sentiment labels, and entity labels.
[0147] Step S706: Fine-tune the output layer weights of the first pre-trained model and the second pre-trained model based on the first set of training data and the second set of training data to obtain a first language model and a second language model.
[0148] In some embodiments, the first pre-trained model uses a multilingual model (such as XLM-RoBERTa), whose shared vocabulary and cross-language encoder can handle inputs in multiple languages, and the second pre-trained model selects a single language model (such as Chinese BERT) to optimize semantic representation for the target language.
[0149] In some embodiments, the structures of the Transform-based pre-trained language model, i.e., the first pre-trained model and the second pre-trained model, are as follows: Figure 3 As shown in the figure, in the modeling process, for the input language text, data processing is first required to generate regular text sequence data, including sentence segmentation, truncation of too long text, padding of too short text, and other operations. Then, the Embedding layer (including Token Embedding, Position Embedding, Segment Embedding) is used to convert the text into a vector. Finally, a multi-layer Transform layer stacking structure is used to encode the text sequence to obtain the semantic vector representation of the text. In this disclosure, considering the complexity of the task and the limitations of computing resources, the base version pre-trained models BERT_base and XLM-RoBERTa_base are selected. The base version pre-trained model contains a 12-layer Transformer encoder. The size of each hidden layer is 778 hidden units, and the multi-head attention mechanism contains 12 attention heads.
[0150] Step S708 , performing two-stage model training based on the first set of training data and the second set of training data to obtain a prediction model, which is used to predict topics, emotions, and entities of the fusion vector generated based on the original language text.
[0151] In some embodiments, the model used for public opinion information processing can be called a multilingual public opinion analysis model. The training is divided into two stages. The first stage uses Chinese news corpus and annotation results to train and construct a multilingual public opinion analysis model. The second stage uses the original multilingual news text corpus, Chinese news corpus and annotation results to train and construct a multilingual public opinion analysis model. After the model training is completed, the optimized parameters and weights need to be saved to a persistent storage medium for future use. This process usually involves saving the model structure, weights and optimizer status to a file system or cloud storage.
[0152] In this embodiment, by preprocessing the original and translated historical texts in multiple languages and completing the three-task labeling of topics, emotions, and entities to generate two sets of training data, the model's ability to uniformly process texts in different languages is enhanced, enabling the model to effectively capture cross-language semantic commonalities. The three-task labeling mechanism promotes the collaborative learning of topic, emotion, and entity features. Combined with the two-stage model training, it effectively enhances the cross-language entity alignment capability and multi-task feature complementarity, enabling the model to perform better in tasks such as topic classification, sentiment analysis, and entity extraction.
[0153] In one embodiment of the present disclosure, a two-stage model training is performed based on the first set of training data and the second set of training data to obtain a prediction model, including:
[0154] In the first stage, based on the second set of training data, the three original linear layers are trained based on topic classification, sentiment analysis, and entity extraction, respectively, to obtain the first intermediate linear layer, the second intermediate linear layer, and the third intermediate linear layer, respectively. The same weights are configured for the loss functions of the three original linear layers.
[0155] In some embodiments, the loss function for the three original linear layers is L = αL1 + βL2 + γL3, where L1 corresponds to the loss function for topic classification, L2 represents the loss function for sentiment analysis, and L3 represents the loss function for entity extraction. α, β, and γ are hyperparameters that control the weights of the tasks. Here, all three are set to 1, indicating that topic classification, sentiment analysis, and intent recognition are equally important.
[0156] In some embodiments, a second set of training data (translated target language text, such as Chinese) is used for training because it usually has more sufficient annotation resources and more accurate monolingual semantic expression (such as the delicacy of Chinese sentiment words). An equal-weight loss function is used for the original linear layers of the three tasks of topic classification, sentiment analysis, and entity extraction to ensure that the model pays equal attention to the three types of tasks in the initial stage, avoiding the dominance of the training process due to the larger loss value of a certain type of task. The first, second, and third intermediate linear layers obtained through monolingual data training lay the foundation for the model's basic discrimination ability in the target language.
[0157] In the second stage, the first set of training data and the second set of training data are fused to generate a first fused training vector and a second fused training vector.
[0158] Based on the first fused training vector and the second fused training vector, the first intermediate linear layer, the second intermediate linear layer, the third intermediate linear layer and the original CRF layer are trained respectively to obtain the first linear layer, the second linear layer, the third linear layer and the CRF layer.
[0159] In some embodiments, the sentence-level semantic vectors ([CLS] vectors) of the original language text (such as English) and the target language text are fused to capture cross-language global semantic alignment, the word-level semantic vectors (including subwords and separator markers) of the original language text and the target language text are fused, the cross-language entity boundary alignment is strengthened, the intermediate linear layer is trained twice using the fused vectors, the model is guided to learn cross-language semantic differences (such as the emotional mapping of English metaphors and Chinese straightforward expressions), the multilingual generalization ability is improved, the original CRF layer training is introduced, and the dependency of the entity label sequence is optimized in combination with cross-language word-level features to prevent cross-language entity fragmentation caused by single-language training.
[0160] In one embodiment of the present disclosure, the original historical language texts in multiple languages obtained by crawling are preprocessed and subjected to three-task labeling operations to obtain a first set of training data, including:
[0161] Based on the configured news sources, original historical language texts in multiple languages are obtained; character-level cleaning and language normalization processing are performed on the original historical language texts to obtain preprocessed texts; text topic category labeling, sentiment polarity labeling, entity type and boundary standards are performed on the preprocessed texts to obtain the first set of training data.
[0162] In some embodiments, the original multilingual international news text is first obtained from a configured news source and translated into the corresponding Chinese news text using a translation API. The news text is preprocessed. In the present disclosure, preprocessing includes removing special characters, deleting non-text elements (hyperlinks, HTML tags, etc.), removing stop words, processing accents and diacritics, unifying capitalization, and removing duplicate characters. After preprocessing, the news data is manually categorized, sentiment analyzed, and entity extracted and annotated.
[0163] like Figure 8 As shown, a model training method for public opinion information processing according to another embodiment of the present disclosure includes:
[0164] Step S802: The multilingual international news text is translated into Chinese news text.
[0165] In step S804, the two texts enter the text preprocessing phase together.
[0166] Step S806: After pre-processing, the text is annotated with news classification, sentiment analysis, and entity extraction.
[0167] Multilingual model training consists of two stages:
[0168] Step S808: In the first stage, a multilingual model is trained using Chinese news corpus.
[0169] Step S810, the second stage uses multilingual news corpus to train a multilingual model to enhance the model's cross-language generalization capability.
[0170] Step S812: Finally, the trained model is saved through model persistence.
[0171] In this embodiment, through the strategy of single-language initialization and multi-language optimization, Chinese annotation resources are used to improve the model's basic task capabilities in the target language, and then multi-language data is used to expand cross-language adaptability. A model training system that supports multi-language public opinion analysis is constructed to achieve collaborative learning and semantic alignment of topic, emotion, and entity tasks in cross-language scenarios.
[0172] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0173] Refer to the following Figure 9 To describe the public opinion information processing device 900 according to an embodiment of the present invention. Figure 9 The public opinion information processing device 900 shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0174] The public opinion information processing device 900 is implemented in the form of a hardware module. The components of the public opinion information processing device 900 may include, but are not limited to: an encoding module 902 for encoding the crawled original language text based on a first language model to obtain an original language vector, and encoding the target language text based on a second language model to obtain a target language vector, wherein the target language text is generated by translating the original language text, and the first language model is a multilingual model; a fusion module 904 for generating a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector; a first prediction module 906 for performing topic and sentiment classification based on the first fusion vector, respectively, to obtain corresponding topic features and sentiment features; a second prediction module 908 for performing entity extraction on the second fusion vector to obtain entity features; and an analysis module 910 for generating public opinion analysis results based on topic features, sentiment features, and entity features.
[0175] Refer to the following Figure 10 To describe the model training device 1000 for public opinion information processing according to an embodiment of the present invention. Figure 10 The model training device 1000 for public opinion information processing shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0176] The model training device 1000 for processing public opinion information is implemented in the form of a hardware module. The components of the model training device 1000 for processing public opinion information may include, but are not limited to: a first processing module 1002 for preprocessing and performing three-task labeling operations on the original historical language texts in multiple languages obtained by crawling to obtain a first set of training data; a second processing module 1004 for preprocessing and performing three-task labeling operations on the target historical language texts translated from the original historical language texts in multiple languages to obtain a second set of training data, where the three-task labels include topic labels, sentiment labels, and entity labels; a fine-tuning module 1006 for fine-tuning the output layer weights of the first pre-trained model and the second pre-trained model based on the first set of training data and the second set of training data to obtain a first language model and a second language model; a training module 1008 for performing two-stage model training based on the first set of training data and the second set of training data to obtain a prediction model, where the prediction model is used to predict the topic, sentiment, and entity of the fusion vector generated based on the original language text.
[0177] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Therefore, various aspects of the present invention may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."
[0178] Refer to the following Figure 11 An electronic device 1100 according to this embodiment of the present invention will be described. Figure 11 The electronic device 1100 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0179] like Figure 11 As shown, electronic device 1100 is implemented as a general-purpose computing device. Components of electronic device 1100 may include, but are not limited to, the aforementioned at least one processing unit 1110, the aforementioned at least one storage unit 1120, and a bus 1130 connecting various system components (including storage unit 1120 and processing unit 1110).
[0180] The storage unit stores program codes, which can be executed by the processing unit 1110, so that the processing unit 1110 performs the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification. For example, the processing unit 1110 can perform the following steps: Figure 2 The described scheme.
[0181] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 11201 and / or a cache memory unit 11202 , and may further include a read-only memory unit (ROM) 11203 .
[0182] The storage unit 1120 may also include a program / utility 11204 having a set (at least one) of program modules 11205, such program modules 11205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0183] The bus 1130 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0184] Electronic device 1100 can also communicate with one or more external devices 1170 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1100, and / or any device that enables electronic device 1100 to communicate with one or more other computing devices (e.g., a router, modem, etc.). Such communication can occur via input / output (I / O) interface 1150. Furthermore, electronic device 1100 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via network adapter 1160. As shown, network adapter 1160 communicates with other modules of electronic device 1100 via bus 1130. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with electronic device 1100, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0185] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0186] In exemplary embodiments of the present disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the above-described methods of this specification is stored. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is executed on an electronic device, the program code is used to cause the electronic device to perform the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section of this specification.
[0187] According to an embodiment of the present invention, a program product for implementing the above-mentioned method can be a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on an electronic device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0188] The program product may be implemented in any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0189] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0190] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0191] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0192] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0193] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0194] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0195] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. A method for processing public opinion information, characterized in that: include: Encoding the crawled original language text based on a first language model to obtain an original language vector, and encoding the target language text based on a second language model to obtain a target language vector, wherein the target language text is generated by translating the original language text, and the first language model is a multilingual model; Generate a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector; Performing topic and sentiment classification based on the first fusion vector to obtain corresponding topic features and sentiment features; performing entity extraction on the second fusion vector to obtain entity features; Generate public opinion analysis results based on the topic features, the sentiment features, and the entity features.
2. The method for processing public opinion information according to claim 1, characterized in that: The target language vector includes a first target language vector and a second target language vector. The original language vector is obtained by encoding the original language text based on the first language model, and the target language vector is obtained by encoding the target language text based on the second language model, including: Encoding the original language text based on the first language model to obtain a first output sequence, and encoding the target language text based on the second language model to obtain a second output sequence; Extracting hidden layer information corresponding to the global semantic tag of the first output sequence as the original language vector; extracting hidden layer information corresponding to the global semantic tag in the second output sequence as the first target language vector; Hidden layer information corresponding to other tokens in the second output sequence is extracted as the second target language vector, where the second target language vector includes multiple first subvectors and second subvectors, where the first subvector is the hidden layer information corresponding to each subword in the target language text, and the second subvector is the hidden layer information corresponding to the separator token in the second output sequence.
3. The method for processing public opinion information according to claim 2, characterized in that: Generating a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector includes: Adding and fusing the original language vector and the first target language vector to obtain the first fused vector; The original language vector is added and fused with the multiple first sub-vectors and the second sub-vector respectively to obtain multiple third sub-vectors, and the multiple third sub-vectors are used as the second fused vector.
4. The method for processing public opinion information according to claim 1, characterized in that: Based on the first fusion vector, topic and sentiment classification are performed respectively to obtain corresponding topic features and sentiment features, including: Mapping the first fusion vector to the topic classification space based on the first linear layer to obtain a first logarithmic probability feature of the topic classification; Outputting a first predicted probability for each topic category by a first softmax function based on the first log-probability feature, so as to determine the topic feature based on the first predicted probability; Mapping the first fusion vector to the sentiment classification space based on the second linear layer to obtain a second logarithmic probability feature of sentiment classification; A second softmax function is used to output a second predicted probability for each emotion category based on the second log-probability feature, so as to determine the emotion feature based on the second predicted probability.
5. The method for processing public opinion information according to claim 3, characterized in that: Performing entity extraction on the second fusion vector to obtain entity features includes: Mapping the plurality of third sub-vectors to the label space respectively based on a third linear layer to obtain an initial label probability for each of the third sub-vectors; The initial label probability is input into a conditional random field (CRF) layer as a state score, so that the CRF layer learns the transition scores between adjacent labels based on the state score, and performs Viterbi decoding based on the state score and the transition score to obtain a label sequence, so as to obtain the entity feature based on the label sequence.
6. The method for processing public opinion information according to claim 1, characterized in that: The sentiment classification includes positive, negative and neutral, and generating a public opinion analysis result based on the topic feature, the sentiment feature and the entity feature includes: Detecting that the topic feature belongs to the target topic and the sentiment feature is negative, acquiring longitude and latitude information and / or matching administrative regions based on the geographic entity in the entity feature; A map marking operation of the public opinion location is performed based on the latitude and longitude information, and / or the public opinion influence range is divided based on the matched administrative areas to serve as the public opinion analysis result.
7. The method for processing public opinion information according to claim 1, characterized in that: Generating a public opinion analysis result based on the topic feature, the sentiment feature, and the entity feature includes: Detecting that the topic feature belongs to a target topic, extracting a subject entity, a geographic entity, and an event entity from the entity feature; It is detected that the subject entity includes organization information, and a business impact analysis result is generated based on a geographic association rule, a subject association rule, and an event and business association rule; The public opinion analysis result is generated based on the business impact analysis result and the sentiment characteristics.
8. The method for processing public opinion information according to claim 1, characterized in that: Before generating a public opinion analysis result based on the topic feature, the sentiment feature, and the entity feature, the method further includes: Extracting and proposing a plurality of reference entities from the entity features; Filtering a candidate set from a pre-stored database based on key entities among the multiple reference entities; Perform entity matching on the candidate set based on the multiple reference entities to obtain a matching rate; Based on the relationship between the matching rate and the matching threshold, a repeated tone determination result is obtained, wherein, if the matching rate is detected to be greater than the matching threshold, it is determined to be repeated public opinion and the public opinion analysis result is not generated. If the matching rate is detected not to be greater than the matching threshold, the public opinion analysis result is generated based on the topic features, the emotional features and the entity features.
9. A model training method for public opinion information processing, characterized in that: include: The original historical language texts in multiple languages obtained by crawling are preprocessed and labeled with three tasks to obtain the first set of training data; performing preprocessing and three-task labeling operations on target historical language texts obtained by translating the original historical language texts in the plurality of languages to obtain a second set of training data, wherein the three-task labeling includes a topic label, a sentiment label, and an entity label; Fine-tune the output layer weights of the first pre-trained model and the second pre-trained model based on the first set of training data and the second set of training data to obtain a first language model and a second language model; A two-stage model training is performed based on the first set of training data and the second set of training data to obtain a prediction model, which is used to predict topics, emotions, and entities for a fusion vector generated based on the original language text.
10. The model training method for public opinion information processing according to claim 9, characterized in that: Performing two-stage model training based on the first set of training data and the second set of training data to obtain a prediction model includes: In the first stage, based on the second set of training data, the three original linear layers are trained based on topic classification, sentiment analysis, and entity extraction, respectively, to obtain a first intermediate linear layer, a second intermediate linear layer, and a third intermediate linear layer, wherein the loss functions of the three original linear layers are configured with the same weights; In the second stage, the first set of training data and the second set of training data are fused to generate a first fused training vector and a second fused training vector; Based on the first fusion training vector and the second fusion training vector, model training is performed on the first intermediate linear layer, the second intermediate linear layer, the third intermediate linear layer and the original CRF layer to obtain a first linear layer, a second linear layer, a third linear layer and a CRF layer.
11. The model training method for public opinion information processing according to claim 9, characterized in that: The original historical language texts in multiple languages obtained by crawling are preprocessed and labeled with three tasks to obtain the first set of training data, including: Obtaining original historical language texts in the plurality of languages based on configured news sources; Performing character-level cleaning and language normalization processing on the original historical language text to obtain a preprocessed text; The preprocessed text is annotated with text topic categories, sentiment polarity, entity types, and boundary standards to obtain the first set of training data.
12. A public opinion information processing device, characterized in that: include: An encoding module, configured to encode the crawled original language text based on a first language model to obtain an original language vector, and to encode the target language text based on a second language model to obtain a target language vector, wherein the target language text is generated by translating the original language text, and the first language model is a multilingual model; A fusion module, configured to generate a first sentence-level fusion vector and a second word-level fusion vector based on the original language vector and the target language vector; A first prediction module is used to perform topic and sentiment classification based on the first fusion vector to obtain corresponding topic features and sentiment features; A second prediction module is used to perform entity extraction on the second fusion vector to obtain entity features; An analysis module is used to generate public opinion analysis results based on the topic features, the sentiment features and the entity features.
13. A model training device for public opinion information processing, characterized in that: include: The first processing module is used to perform preprocessing and three-task labeling operations on the original historical language texts in multiple languages obtained by crawling to obtain a first set of training data; a second processing module for performing preprocessing and three-task labeling operations on the target historical language texts obtained by translating the original historical language texts in the plurality of languages to obtain a second set of training data, wherein the three-task labels include topic labels, sentiment labels, and entity labels; a fine-tuning module, configured to fine-tune the output layer weights of the first pre-trained model and the second pre-trained model based on the first set of training data and the second set of training data to obtain a first language model and a second language model; A training module is used to perform two-stage model training based on the first set of training data and the second set of training data to obtain a prediction model, wherein the prediction model is used to predict topics, emotions and entities of a fusion vector generated based on the original language text.
14. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the public opinion information processing method described in any one of claims 1 to 8 or the model training method for public opinion information processing described in 9 to 11 by executing the executable instructions.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the public opinion information processing method described in any one of claims 1 to 8 or the model training method for public opinion information processing described in 9 to 11.
16. A computer program product having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the public opinion information processing method described in any one of claims 1 to 8 or the model training method for public opinion information processing described in 9 to 11.
Citation Information
Cited By
Cascade condition-based multi-task learning bad language information detection method and system
CN121834621A
Cascade condition-based multi-task learning poor language information detection method and system
CN121834621B