Public opinion monitoring method, system and equipment based on language large model and storage medium
Through automated collection and language big model classification methods, the time-consuming and inaccurate problems of traditional public opinion monitoring systems are solved, real-time and accurate public opinion monitoring is achieved, and the safety and stability of the enterprise is improved.
Patent Information
- Application Number
- CN202510306089.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-08-05
AI Technical Summary
Traditional public opinion monitoring systems rely on manual reading and manual data collection, which is time-consuming and labor-intensive, not real-time enough, and have subjective deviations and inaccuracies, so they cannot fully monitor online public opinion.
Automatic collection tools are used to simulate user operations to screen public opinion data, and use trained language models for classification to achieve automated and real-time public opinion monitoring.
It improves the efficiency and accuracy of public opinion monitoring, can detect negative public opinion in a timely manner, improves the automation and intelligence level of the system, and saves manpower and material resources.
Smart Images

Figure CN120429489A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence and data processing technology, and in particular to a method, system, device and storage medium for public opinion monitoring based on a large language model. Background Art
[0002] During the implementation of this application, the inventors discovered that the prior art has at least the following problems:
[0003] Traditional public opinion monitoring systems typically rely on manual reading, screening, and analysis of large amounts of online public opinion information. This approach is time-consuming, labor-intensive, and prone to omissions. Data collection is often manual, making real-time, comprehensive monitoring impossible. The determination of public opinion categories also relies primarily on manual judgment, which is subject to subjective bias and a certain degree of inaccuracy.
[0004] In addition to traditional public opinion monitoring systems that rely on manual judgment, there are other traditional public opinion monitoring system solutions, such as:
[0005] Rule-based matching systems: These systems filter and categorize public opinion information by setting a series of rules and keywords. This approach is based on simple keyword matching. However, this approach relies on static rules and cannot flexibly respond to the diversity and changes in public opinion expression.
[0006] Sentiment lexicon-based systems: These systems use sentiment lexicons (positive, negative, and neutral sentiment words) to analyze text and determine public opinion. However, sentiment lexicons may not fully capture the context of public opinion, leading to biased analysis results. Summary of the Invention
[0007] The main purpose of this application is to provide a method, system, device and storage medium for public opinion monitoring based on a large language model, aiming to solve the technical problem in the existing technology that manual public opinion monitoring cannot be carried out comprehensively, in real time and accurately.
[0008] In a first aspect, the present application provides a method for monitoring public opinion based on a large language model. The method comprises:
[0009] Use automated collection tools to access the target webpage and simulate user operations to filter out raw public opinion data using keywords on the target webpage;
[0010] Use the trained language model to classify the original public opinion data and obtain the classification results.
[0011] This application also provides a public opinion monitoring system based on a large language model, which includes:
[0012] The data collection module is used to access the target web page using automated collection tools, simulate user operations, and filter out raw public opinion data through keywords on the target web page;
[0013] The classification module is used to classify the original public opinion data using the trained language model to obtain classification results.
[0014] The third aspect of the present application provides a computer device, comprising: a memory and at least one processor, wherein instructions are stored in the memory; at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned public opinion monitoring method based on the language large model.
[0015] The fourth aspect of the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the above-mentioned public opinion monitoring method based on the language large model.
[0016] One of the above technical solutions has the following beneficial effects:
[0017] This application realizes automatic real-time monitoring of public opinion through the automated collection and processing of online public opinion, improves the efficiency and accuracy of public opinion monitoring, and can timely detect negative public opinion; this application introduces advanced large language models for automated public opinion sentiment analysis, realizes automatic identification and classification of negative public opinion, and greatly improves the accuracy and efficiency of public opinion sentiment analysis; this application effectively solves the limitations of traditional public opinion monitoring systems and can accurately classify public opinion data of various types and forms of expression; this application can train the public opinion classification capabilities of large language models without data labeling and feature extraction, saving manpower and material resources; this application improves the automation, accuracy, real-time and intelligence level of the public opinion monitoring system, and further ensures the security and stability of corporate business. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of the first embodiment of the method for public opinion monitoring based on a large language model in the embodiment of the present application;
[0019] Figure 2 This is a flow chart of a second embodiment of the method for monitoring public opinion based on a large language model in the embodiment of the present application;
[0020] Figure 3 This is a functional module diagram of an embodiment of a public opinion monitoring system based on a large language model in an embodiment of the present application;
[0021] Figure 4 This is a schematic diagram of an embodiment of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0022] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0023] Public opinion monitoring plays a crucial role in businesses, such as corporate operations. Its value lies in helping companies understand public opinion about their brands, products, or services, promptly identifying and addressing negative public opinion, protecting their brand reputation, and enhancing their public image. Public opinion monitoring can help companies understand competitor dynamics and market changes, providing crucial information for developing marketing strategies and addressing market challenges. It can also help companies forewarn of potential crises, enabling rapid response and developing effective crisis public relations strategies to mitigate the negative impact on their business. By monitoring and analyzing public opinion information from various channels, companies can gain more accurate insights into the market, consumer demand, and social trends, providing strong support for business decision-making. The importance of public opinion monitoring lies in providing information support, risk warnings, and decision-making insights, helping companies remain agile, flexible, and competitive in a highly competitive market.
[0024] Traditional public opinion monitoring systems have limitations. For example, they typically rely on manual reading, filtering, and analysis of large amounts of online public opinion information, a time-consuming and labor-intensive process that can easily miss information. Data collection is often manual, making real-time and comprehensive monitoring impossible. Public opinion categorization also relies heavily on manual judgment, which can be subject to subjective bias and a certain degree of inaccuracy.
[0025] Based on this, this application provides a public opinion monitoring solution based on a large language model.
[0026] refer to Figure 1 In one embodiment, the present application provides a method for monitoring public opinion based on a large language model, the method comprising:
[0027] S100: Use automated collection tools to access the target web page, simulate user operations, and filter out original public opinion data through keywords on the target web page.
[0028] Specifically, automated collection tools can be web crawlers. For example, Playwright can be used to automatically collect online public opinion data. Playwright is an automated testing and web crawling tool built on the Node.js framework. It automates browser operations and supports various network requests and interactions. Playwright supports multiple browsers, including Chromium, Firefox, and WebKit, meeting cross-browser testing requirements. It also provides a rich API that can simulate various user actions such as clicks, input, and scrolling, as well as set cookies and modify browser features.
[0029] The automated collection tool can automatically log in to the target website and access the target webpage therein, and automatically simulate real-life user operations on the target webpage, where the simulated user operations include at least one of pull-down operations, scrolling operations, expansion operations, click operations, page turning operations, jump operations, and search operations. For example, accessing user comments on products, user evaluations and messages on companies, etc., and then using keywords to find the target content containing keywords in these contents as raw public opinion data. More specifically, for example, to find whether there are comments discussing (or evaluating) Company A or Company A's products on the target webpage, the comments containing Company A or Company A's products will be filtered out as raw public opinion data.
[0030] Automated collection tools can also flip through pages on the target web page to enable multi-page browsing and collection of target public opinion data.
[0031] The automated collection tool will also detect whether there is a next page button on the current web page to indicate that the page can be turned to the next page. If there is no next page button, or the next page button is in a prohibited operation state, it is determined that the content of the target web page and its upper and lower web pages have been crawled, and this round of crawling ends.
[0032] Alternatively, based on a series of rules: for example, whether a certain keyword has collected a certain amount of original public opinion data, or whether all relevant public opinions on a certain platform have been collected (which can be determined by whether there is a page-turning button), it can be determined whether to end the crawling.
[0033] If there is a next page button, the automated collection tool will simulate page turning, for example, clicking the next page button will go to the next page for web crawling.
[0034] In addition, on the current web page, the automated collection tool also has operations such as simulated pull-down or simulated sliding scroll bars. For example, if the current web page has a lot of content and requires pulling down or sliding the scroll bar to display more web page content, the automated collection tool can perform operations such as simulated pull-down or simulated sliding scroll bars to display more data.
[0035] By simulating operations such as pulling down or sliding the scroll bar, you can also locate the page turning button on the web page to turn the page.
[0036] S200: Use the trained language model to classify the original public opinion data to obtain the classification results.
[0037] Specifically, the collected raw public opinion data is classified to determine whether it is negative public opinion. In this process, this embodiment uses a language model. The language model can simulate the human language cognition and generation process to a certain extent, and can understand and generate human language. The language model is mainly used for natural language processing tasks such as text generation, translation, summarization, question and answer, etc. Therefore, it can accurately identify and classify negative public opinion.
[0038] Some traditional public opinion monitoring systems based on machine learning algorithms utilize traditional machine learning algorithms (such as Naive Bayes and Support Vector Machines) to classify and analyze public opinion text. These systems require large amounts of labeled data and manual feature extraction for supervised training, and are difficult to adapt to new forms of public opinion expression.
[0039] Large language models have strong contextual understanding capabilities, enabling them to process long texts and understand the contextual relationships within them, enabling more accurate classification. Compared to simpler text classification models, large language models are able to understand and reason about abstract concepts, enabling them to handle more challenging text classification tasks. Furthermore, large language models do not require extensive labeled data or feature extraction; they can achieve public opinion classification through unsupervised or weakly supervised training.
[0040] Based on the powerful natural language processing capabilities of the language big model, this embodiment uses the language big model to classify public opinion data.
[0041] More specifically, the language model may use, for example, Wenxin Yiyan, etc., and this application does not impose any restrictions on this.
[0042] The classification result of each piece of original public opinion data is one of negative public opinion, positive public opinion and neutral public opinion.
[0043] In contrast, the automated public opinion monitoring system based on the large language model can analyze online public opinion more accurately and efficiently, and improve the intelligence and practicality of the public opinion monitoring system.
[0044] In one embodiment, the public opinion monitoring method based on the language large model may further include: associating and storing the original public opinion data with the corresponding classification results.
[0045] Specifically, associating and storing the original public opinion data with the corresponding classification results is equivalent to classifying and labeling the original public opinion data to facilitate distinguishing various types of public opinion data.
[0046] This embodiment realizes automatic real-time monitoring of public opinion through the automated collection and processing of online public opinion, improves the efficiency and accuracy of public opinion monitoring, and can timely detect negative public opinion; this embodiment introduces advanced language large models for automated public opinion sentiment analysis, realizes automatic identification and classification of negative public opinion, and greatly improves the accuracy and efficiency of public opinion sentiment analysis; this embodiment effectively solves the limitations of traditional public opinion monitoring systems, and can accurately classify public opinion data of various types and forms of expression; this embodiment can train the public opinion classification capability of the large language model without data labeling and feature extraction, saving manpower and material resources; this embodiment improves the automation, accuracy, real-time and intelligence level of the public opinion monitoring system, and further ensures the security and stability of corporate business.
[0047] In one embodiment, in step S200, the original public opinion data is classified using the trained language model to obtain classification results, including:
[0048] Insert the original public opinion data into the target queue;
[0049] Monitor the amount of raw public opinion data inserted into the target queue. If the amount of raw public opinion data inserted into the target queue reaches a preset threshold, read the same amount of raw public opinion data as the preset threshold from the target queue, and generate a model access request based on the read raw public opinion data.
[0050] Send a model access request to the language model by calling the batch processing interface of the language model;
[0051] The language big model is used to batch classify the original public opinion data in the model access request to obtain the classification result corresponding to each original public opinion data, where the classification result is one of positive public opinion, negative public opinion and neutral public opinion.
[0052] Specifically, the language model includes a batch execution interface for executing tasks in batches, namely a batch processing interface (eg, a batch interface) and a single execution interface for executing tasks one by one.
[0053] By calling a single execution interface of the language model, each piece of raw public opinion data can be classified individually. That is, each time the automated collection tool captures a piece of raw public opinion data, the classification module generates a prompt based on the raw public opinion data. By calling a single execution interface of the language model, the prompt is input into the language model, which then performs classification based on the prompt and produces a classification result for the raw public opinion data.
[0054] If the language model's batch execution interface, also known as the batch processing interface (batch interface), is called, multiple pieces of raw public opinion data can be batch-classified. To achieve batch classification, this embodiment first inserts the collected raw public opinion data into a target queue for consumption. Simultaneously, the collected raw public opinion data can also be written to a data storage unit to ensure immediate data storage. The data storage unit can, for example, be Elasticsearch.
[0055] As web pages are crawled, the raw public opinion data in the target queue will increase accordingly. The amount of raw public opinion data in the target queue will be monitored. If the amount of raw public opinion data in the target queue reaches a preset threshold, such as 10, 20, 50, or other numbers not limited to these, the raw public opinion data will be read from the target queue and a model access request will be generated based on the raw public opinion data.
[0056] In addition, the original public opinion data in the target queue will be destroyed or deleted after reading, and new original public opinion data will be accumulated in the target queue.
[0057] This embodiment utilizes the batch task execution function of the language model. Therefore, this embodiment invokes the language model's batch processing interface to send a model access request to the language model. This model access request includes prompts or questions corresponding to a preset threshold of raw public opinion data. Therefore, the language model can perform batch classification on this preset threshold of raw public opinion data, obtaining classification results for each piece of raw public opinion data.
[0058] In one specific embodiment, the requests library can be used to automatically send requests. The requests library is a convenient HTTP request library in Python that can simulate browser requests and easily send HTTP requests. The requests library allows for sending various HTTP requests and includes various request templates. Therefore, model access requests can be automatically generated and sent to the large language model.
[0059] The classification results returned by the language model include the classification results of each original public opinion data in the model access request. The classification results returned by the language model are used to label and classify the corresponding original public opinion data in the data storage unit, or each classification result is stored corresponding to the original public opinion data in the data storage unit to achieve the associated storage of the original public opinion data and the corresponding classification results.
[0060] Using the batch processing API generally does not reduce the accuracy of model predictions. This is because the prediction algorithm and learned parameters are the same regardless of whether single or batch processing is used. Batch processing does not change the model's internal architecture or affect its prediction methods.
[0061] This embodiment uses the batch processing interface of the queue and language model to process multiple text messages at a time and return the prediction results for each one. Batch processing means that multiple tasks can be performed simultaneously instead of processing only one piece of data at a time, which can significantly improve the processing efficiency of large amounts of data.
[0062] In one embodiment, the public opinion monitoring method based on the language large model further includes:
[0063] Storing the original public opinion data in Elasticsearch;
[0064] Call the Bulk interface of Elasticsearch and use the classification results returned by the large language model to batch label and classify the corresponding original public opinion data in Elasticsearch.
[0065] Specifically, in the prior art, data storage and retrieval use traditional databases or file systems, which have slow query speeds and are not conducive to real-time monitoring and analysis.
[0066] Elasticsearch is a distributed, scalable, real-time search and data analysis engine that can store and process big data. To facilitate storage and data processing, this embodiment stores the classified public opinion data (i.e., the original public opinion data and the corresponding classification results) in Elasticsearch.
[0067] The collected raw public opinion data is immediately written into Elasticsearch to ensure real-time storage of the data.
[0068] After all the prediction results (i.e., classification results) are returned, the Elasticsearch Bulk interface (BulkAPI) is called to perform batch updates. That is, the classification results are written to Elasticsearch and stored in association with the corresponding original public opinion data, so as to label and classify the original public opinion data in Elasticsearch. This not only ensures the integrity of data storage but also improves the efficiency of data processing. Elasticsearch's Bulk API is mainly used to perform multiple indexing / deleting / writing / reading operations in a single request, which helps reduce the number of network round trips and allows Elasticsearch to use parallel processing internally, thereby improving data processing performance.
[0069] This embodiment uses Elasticsearch to store data and calls Elasticsearch's Bulk interface to batch update data, which can efficiently store and retrieve data, improve query speed, and facilitate real-time monitoring, analysis, and visual analysis of public opinion data.
[0070] In one embodiment, before classifying the original public opinion data using the trained language model in step S200, the public opinion monitoring method based on the language model further includes:
[0071] Constructing a current prompt based on input training corpus, and inputting the current prompt into the language model, wherein the training corpus includes the current sample public opinion data to be classified, or the training corpus includes the current sample public opinion data to be classified and the corresponding public opinion category description;
[0072] receiving and displaying a response result returned by the language macromodel based on the current prompt;
[0073] Feedback information given by the user in response to the response result is input into the language model, so that the language model optimizes classification parameters according to the feedback information.
[0074] Specifically, the language model can process various types of natural language, including negative sentences, interrogative sentences, etc., and can classify various forms of public opinion relatively accurately. If the language model is inaccurate in some cases, manual intervention can be easily performed. By utilizing the powerful human-computer interaction capabilities of the language model, it can interact with users at a deeper level when processing the public opinion classification process and obtain more attitude and intention information.
[0075] During the dialogue interaction process, the computer device will display a human-computer dialogue interface, receive user input, and construct prompts or prompt questions based on the question template, the sample public opinion data input by the user, and the specified public opinion category description (standard definition). For example, the public opinion category description is spliced before the sample public opinion data to be judged to obtain a prompt, and the prompt is sent to the language model. Alternatively, the computer device will construct a prompt based on the feedback information of the previous prompt input by the user and the current sample public opinion data to be classified and the specified public opinion category description. The computer device will receive the response result of the language model through the human-computer dialogue interface. Each time a new input from the user is received, the above operation is repeated to realize the human-computer interaction between the user and the language model, thereby achieving the purpose of training the language model.
[0076] The prompt can be a prompt. A prompt is a piece of text input that guides the model to generate the corresponding text output. A prompt can be considered an initial condition or a question given to the model. The model will generate coherent and relevant subsequent text based on this prompt. A prompt question is equivalent to a prompt input to the language model.
[0077] The following example illustrates the detailed steps for training the language model for public opinion classification through human-computer interaction with the language model:
[0078] A piece of public opinion information is collected as sample public opinion data, and a question (prompt) is constructed based on the sample public opinion data. For example, the content is: "I give you a public opinion message. Can you determine whether it is negative information based on this information?" At this step, the language model will respond and explain the various clues behind the information: "Based on the information given, this public opinion message can be regarded as negative information. My reasons for this judgment are as follows: 1. XXXX, 2. XXXXX, 3. XXXXXX".
[0079] Based on the clear definition of negative public opinion, the next prompt is constructed, prepending this definition to the sample to be classified, and then querying the language model. This step may require multiple rounds of dialogue with the language model to help it understand what constitutes negative information. Users can also employ some techniques, such as assigning the language model a role name and task, such as public opinion analyst. This requires it to respond according to a specific template and return data in JSON format.
[0080] If the language model correctly identifies negative public opinion, the user can give a positive response to the language model, such as praising it: "You did a great job! You successfully identified this negative public opinion message", and then provide the next new data.
[0081] If the language model mistakenly judges a piece of information as positive or neutral, users can help it clarify the misinterpretation that led to it. Together, they can summarize the reasons for the error and incorporate this experience as a warning or reminder into the next prompt, implementing a secondary prompt process to correct the language model. For example, "Your judgment is incorrect. As we discussed earlier, negative public opinion information often contains the following factors: 1, ****, 2****, 3****. Let's try another one." This allows the language model to re-evaluate in the new context.
[0082] After pre-training the language model on a few samples, such as approximately 100, users can use it to classify and label all collected public opinion information. When constructing a question (prompt), users add the definition of negative public opinion before the sample to be judged. The model then performs automated weakly supervised judgment on the data and extracts all negative public opinion.
[0083] The following is an example of a specific conversation:
[0084] Question: You are Jiazheng, an outstanding public opinion analysis engineer. Your task is to determine whether public opinion information is negative. For example, you can identify offensive, malicious, or negative opinions and comments online. Now, according to the above definition, I will give you some public opinion information, and you need to determine whether it is negative. Do you understand the task?
[0085] Answer: I know very well. I am Jiazheng. I will complete your task of classifying negative public opinion.
[0086] Question: Very good, let's get started. For the following data (starting with **** and ending with ****):
[0087] ****
[0088] {"user":"John Doe","comment":"I think this product is of very poor quality and I will never buy it again.","post_date":"2021-09-01"}
[0089] ****
[0090] Is this public opinion information negative?
[0091] The result should be returned in strict JSON format, containing two key values: "Analysis" and "Answer." "Analysis" is a detailed description of the reason for the judgment. Note that the value in "Analysis" should not be enclosed in double quotes.
[0092] The answer part only needs to be a concise "yes" or "no".
[0093]
[0094] Through the human-computer dialogue interaction described above, the language model can have strong model training capabilities based on large-scale data training, and can continuously optimize and improve model performance through continuous learning in actual public opinion classification work.
[0095] In one specific embodiment, before training the language model through human-computer interaction, some sample public opinion data can be collected and classified and labeled. Training samples are generated based on the classified and labeled sample public opinion data. The training samples are input into the language model for preliminary training. This improves the expressiveness of the language model during human-computer interaction training, shortens the training time, and improves training efficiency.
[0096] This embodiment achieves accurate training of the public opinion classification capability of the language large model through human-computer interaction.
[0097] In one embodiment, the public opinion monitoring method based on the language large model may further include:
[0098] Visualize the classified raw public opinion data;
[0099] and / or,
[0100] If the negative raw public opinion data reaches the set threshold, an alarm will be triggered.
[0101] Specifically, for example, you can build a data visualization panel on Grafana to visualize the stored and categorized raw public opinion data, making it easier for users to observe and analyze the data. At the same time, you can configure Grafana to alert users of fluctuations in the number of public opinion posts. When the number of public opinion posts reaches a set threshold, an alert will be triggered, allowing you to identify potential risks in a timely manner.
[0102] Grafana is an open-source metric analysis and visualization tool widely used in large-scale data environments, including cloud computing, the Internet of Things, and artificial intelligence. It provides powerful visualization, flexible custom dashboards, and robust alerting mechanisms.
[0103] Visualization: Grafana supports a variety of graphical panels, such as graphs, tables, heat maps, maps, and pie charts, providing a wealth of visualization options for users to choose from. Grafana allows you to create clear and intuitive data display panels, allowing you to understand and interpret data in a more intuitive way.
[0104] Custom Panels: Users can customize multiple panels to form a large dashboard. Each panel can define its data source, type, and display method. A large dashboard can display a collection of multiple panels. This allows users to see the relationship between multiple indicators in the same dashboard.
[0105] Alerting: Grafana has a powerful alerting mechanism. By setting thresholds, alerts are automatically sent when data exceeds or falls below these thresholds. Alerts can be sent via various methods, including email, Slack messages, and webhooks.
[0106] Overall, Grafana is a powerful and flexible tool that can help users effectively monitor and understand complex data.
[0107] In this public opinion monitoring system solution, the data visualization part is designed using Grafana, which mainly includes two types of panels: bar charts and detailed data tables.
[0108] Histogram Panel: This panel provides a visual representation of the changing trends in public opinion data by displaying time series data. Each column represents the total amount of data over a period of time, and its height represents the volume of public opinion during that period. Colors can be set to change gradually based on the volume. This panel allows users to quickly grasp the general trends in public opinion volume and identify unusual fluctuations in a timely manner.
[0109] Detailed Data Table: The detailed data table panel displays detailed information about each piece of public opinion. In this table, each row corresponds to a separate piece of public opinion, and each column corresponds to an attribute of the public opinion, such as source, release time, and content. Users can use this table to view the specific information of each piece of public opinion for detailed analysis.
[0110] By combining these two types of panels, users can not only grasp the public opinion situation from a macro perspective, but also gain an in-depth understanding of the detailed information of each piece of public opinion, greatly improving the efficiency and accuracy of public opinion management.
[0111] This solution uses Grafana's built-in threshold alarm function to generate negative public opinion alerts, which is a very intelligent and real-time way to monitor public opinion. When the system detects that public opinion data exceeds the set threshold, it will automatically trigger an alert.
[0112] Furthermore, the solution ensures accurate and real-time responses to alerts. Rather than simply pushing all public opinion information, it selects and pushes only highly reliable negative public opinion. This prevents information overload from impacting security personnel's judgment.
[0113] Finally, by pushing this important public opinion information to email or to security personnel in the form of text messages or messages, security personnel can obtain the dynamics of public opinion changes in the first time and conduct manual review so as to make corresponding strategic decisions as soon as possible.
[0114] In general, this solution is very efficient in the design of negative public opinion warnings and subsequent processing, enabling enterprises to detect, handle and resolve various possible public opinion risks early.
[0115] This embodiment can quickly retrieve the classified raw public opinion data and build a data visualization panel, making the information presentation more intuitive and easy to analyze. It also configures the relevant public opinion quantity fluctuation alarm, which can timely warn of real-time data changes and respond to risks in a targeted manner. Configuring fluctuation alarms can respond to public opinion changes in a timely manner and reduce risks. It can realize the automated collection, processing and monitoring of online public opinion, improve the efficiency and accuracy of public opinion monitoring, and can timely discover and deal with public opinion risks.
[0116] In one embodiment, accessing the target webpage using an automated collection tool in step S100 includes:
[0117] Use automated collection tools to obtain the corresponding cookie values generated when successfully logging into each website;
[0118] Associate the cookie value with the unique website identifier of the corresponding website and store it;
[0119] Determine the target website where the target webpage to be visited is located;
[0120] Match the target cookie value based on the unique identifier of the target website;
[0121] Setting the target cookie value in the initial access request of the target web page to generate a target access request;
[0122] By sending a target access request to the server of the target website, you can log in to the target website and enter the target web page.
[0123] Specifically, the automated collection tool can simulate a user logging into each target website using a browser. Logging into different websites generates different cookies. The user can perform a real-person login during the initial login. The automated collection tool can then crawl the cookie value after the user logs into the target website for the first time or logs back in after logging out. The cookie value is then associated and stored with the corresponding target website's unique identifier. This allows the automated collection tool to directly log into the target website based on the cookie value the next time the user accesses the target webpage, eliminating the need for the user to log in repeatedly, thus achieving highly automated data collection and crawling.
[0124] Among them, the unique identifier of the website can be the website's domain name, website's full name, etc. This application does not impose any restrictions on this.
[0125] For example, by calling the setCookie API of the Playwright tool, you can set the user's login information. After logging into the website, you can get the cookie value and save it.
[0126] If an automated scraping tool is currently crawling data for a target webpage, it can determine the target website of the target webpage based on the website source, website domain, or initial access request. The target cookie value can be determined based on the correspondence between the website's unique identifier and the cookie value.
[0127] The automated collection tool will input the target source on the application side and control the application side, such as changing the domain name or access link in the browser. The application side will generate an initial access request. The automated collection tool will set the target cookie value in the initial access request of the target web page, such as setting it in the header of the initial access request, to generate the target access request.
[0128] The application sends the target access request to the server of the target website, logs in to the target website and enters the target web page in the logged-in state, thereby crawling the target web page in the logged-in state.
[0129] After the target access request is sent, the application end, such as a browser, will load the data source, render and display the content of the target web page under the target domain name or target website based on the data obtained from the server.
[0130] This embodiment can realize automatic web page crawling by using cookie modification requests through automated collection tools.
[0131] In one embodiment, accessing the target webpage using an automated collection tool in step S100 further includes:
[0132] Use automated collection tools to monitor the login status of the target website after sending the target access request;
[0133] If the target website is not logged in or the login is invalid, instruct the user to perform the target website login operation, or simulate the user login operation according to the pre-stored target website account and password to log in to the target website and enter the target webpage;
[0134] Obtain the new cookie value generated after successfully logging into the target website, and use the new cookie value to update the local cookie file, wherein the local cookie file is used to store the association relationship between the cookie value and the website unique identifier of the corresponding website.
[0135] Specifically, cookies generally have an expiration date, such as one week or two weeks, which may vary from website to website. Furthermore, the expiration date of a cookie is usually monitored by the website being visited.
[0136] Therefore, if an automated scraping tool uses an expired cookie value to log in to a target website, the target website will deny the login or require a new login, indicating an unsuccessful login. To this end, the automated scraping tool monitors the target website's login status after sending the target access request, specifically, whether the login is successful. If a successful login is detected, the cookie is valid, and the automated scraping tool can crawl the webpage normally while logged in.
[0137] If the target website fails to log in successfully, for example, the target website prompts you to log in again or displays the login interface instead of the target page, then the cookie has expired. Based on this, the automated collection tool will instruct the user to perform a real-person login operation to assist the automated collection tool in logging in to the target website.
[0138] Users can log in to the target website by scanning, entering their account and password, logging in via mobile phone, etc. This application does not impose any restrictions on this.
[0139] Among them, the automated collection tool can remind the user to log in to the target website by sending text messages, emails, system messages, etc., but not limited to these methods.
[0140] In addition, after logging back into the target website, the automated collection tool will capture the new cookie value generated by the browser or other application after successfully logging into the target website. It will then update the old cookie value corresponding to the target website in the local cookie file based on the new cookie value, thus updating the cookie file.
[0141] In addition, automated collection tools can also simulate user login operations based on pre-stored target website accounts and passwords and attempt to log in to the target website.
[0142] This embodiment automatically monitors the login status of the website, and can notify the user to assist in logging into the website when not logged in, and re-acquire a new cookie value to facilitate automatic login to the website next time, avoiding the situation where the user is stuck in the non-logged in state for a long time and cannot collect web page data.
[0143] In one embodiment, accessing the target webpage using an automated collection tool in step S100 further includes:
[0144] Use automated collection tools to modify the property value of navigator.webdriver on the application side;
[0145] and / or,
[0146] Use automated collection tools to simulate users' non-automated operations and web browsing behaviors on the target web page.
[0147] Specifically, the automated collection tool can log in to the target website and access the target web page through an application terminal such as a browser or application.
[0148] navigator.webdriver helps detect whether the current browser is controlled by an automated testing tool or a crawler tool. For example, if it returns true, it means the current environment is an automated testing environment or a crawler environment. If it returns false or undefined, it means the current environment is a normal user browsing environment.
[0149] Based on this, this embodiment can use the automated collection tool to tamper with the attribute value of navigator.webdriver on the application side to an empty value, false, or undefined, etc. before sending an access request, which is determined according to the actual situation, so that the application side mistakenly believes that the crawler operation of the automated collection tool is a normal human operation, and successfully crawls the data of the target web page.
[0150] In addition, this embodiment can also use automated collection tools to simulate real users to perform some operations other than web crawling in the target web page, that is, simulate the user's non-automated operations and web browsing behaviors, such as one or more operations such as finding, inputting, searching, scrolling web pages, commenting or leaving messages, so that the application end such as the browser mistakenly believes that it is a normal human operation and successfully crawls the data of the target web page.
[0151] In this embodiment, data of a target web page can be successfully crawled by tampering with the navigator.webdriver property value and / or simulating non-automated browsing behavior.
[0152] This application introduces an advanced language model for automated public opinion analysis. This trained model automatically identifies and categorizes negative public opinion, significantly improving the accuracy and efficiency of analysis. Collected public opinion information is stored in Elasticsearch, enabling rapid retrieval and the construction of data visualization panels, making information presentation more intuitive and easier to analyze. Furthermore, alarms for fluctuations in the volume of relevant public opinion are configured to provide timely warnings of real-time data changes and enable targeted risk responses.
[0153] Figure 2 This is a flowchart of a method for monitoring public opinion based on a large language model in a specific embodiment. Figure 2, data collection begins, and then the collected raw data is preprocessed to obtain raw public opinion data. Data preprocessing includes, for example, removing abnormal characters and converting the raw data into JSON format, which is not limited by this application. Then, the preprocessed raw public opinion data is classified using a large language model, the classified raw public opinion data is stored, and the classified raw public opinion data is displayed through a visualization panel. When the negative raw public opinion data reaches the alarm threshold, the negative public opinion alarm is generated.
[0154] refer to Figure 3 In one embodiment, the present application further provides a public opinion monitoring system based on a large language model, the public opinion monitoring system based on a large language model comprising:
[0155] The data collection module 100 is used to access the target webpage using an automated collection tool, and simulate user operations to filter out raw public opinion data through keywords on the target webpage;
[0156] The classification module 200 is used to classify the original public opinion data using the trained language model to obtain classification results.
[0157] In one embodiment, the public opinion monitoring system based on the language large model further includes:
[0158] The storage module is used to associate and store the original public opinion data with the corresponding classification results.
[0159] In one embodiment, the classification module 200 includes:
[0160] Insertion and storage module, used to insert the original public opinion data into the target queue;
[0161] The model access request generation module is used to monitor the amount of original public opinion data inserted into the target queue. If the amount of original public opinion data inserted into the target queue reaches a preset threshold, the module reads out the same amount of original public opinion data as the preset threshold from the target queue, and generates a model access request based on the read original public opinion data.
[0162] A request sending module is used to send a model access request to the language model by calling the batch processing interface of the language model;
[0163] The batch classification module is used to use the language large model to batch classify the original public opinion data in the model access request to obtain the classification result corresponding to each original public opinion data, where the classification result is one of positive public opinion, negative public opinion and neutral public opinion.
[0164] In one embodiment, the public opinion monitoring system based on the language big model also includes: a storage module for storing the original public opinion data in Elasticsearch; calling the Bulk interface of Elasticsearch and using the classification results returned by the language big model to batch label and classify the corresponding original public opinion data in Elasticsearch.
[0165] In one embodiment, the public opinion monitoring system based on the language large model further includes: a training module;
[0166] The training module is specifically used to: construct a current prompt based on input training corpus, and input the current prompt into the language model, wherein the training corpus includes the current sample public opinion data to be classified, or the training corpus includes the current sample public opinion data to be classified and the corresponding public opinion category description;
[0167] receiving and displaying a response result returned by the language macromodel based on the current prompt;
[0168] Feedback information given by the user in response to the response result is input into the language model, so that the language model optimizes classification parameters according to the feedback information.
[0169] In one embodiment, the data acquisition module 100 includes:
[0170] The cookie capture module is used to use automated collection tools to obtain the corresponding cookie values generated when successfully logging into each website;
[0171] A first association storage module is used to associate and store the cookie value with the unique website identifier of the corresponding website;
[0172] A target website determination module is used to determine the target website where the target webpage to be accessed is located;
[0173] The cookie matching module is used to match the target cookie value based on the unique identifier of the target website;
[0174] A request rewriting module is used to set the target cookie value in the initial access request of the target web page to generate a target access request;
[0175] The request sending module is used to log in to the target website and enter the target web page by sending a target access request to the server of the target website.
[0176] In one embodiment, the data acquisition module 100 further includes:
[0177] A login status monitoring module is used to monitor the login status of the target website after sending a target access request using an automated collection tool;
[0178] A re-login module is used to instruct the user to perform a login operation on the target website if the target website is not logged in or the login is invalid, or to simulate a user login operation based on the pre-stored account and password of the target website to log in to the target website and enter the target webpage;
[0179] The cookie update module is also used to obtain the new cookie value generated after successfully logging into the target website, and use the new cookie value to update the local cookie file, wherein the local cookie file is used to store the association relationship between the cookie value and the website unique identifier of the corresponding website.
[0180] In one embodiment, the data acquisition module 100 further includes:
[0181] The property rewriting module is used to modify the property value of navigator.webdriver on the application side using automated collection tools;
[0182] and / or,
[0183] The data acquisition module 100 further includes:
[0184] The simulation operation module is used to use automated collection tools to simulate the user's non-automated operations and web browsing behaviors on the target web page.
[0185] The implementation principle of the public opinion monitoring system based on the language large model of this application is specifically described in the above description of the public opinion monitoring method based on the language large model, which will not be repeated here.
[0186] This application collects online public opinion information through automated scripts, reducing the workload and time cost of manual collection.
[0187] This application automatically performs positive and negative analysis of public opinion through a large language model, improving the efficiency and accuracy of information processing.
[0188] This application uses Elasticsearch for data storage, which can provide fast and accurate search services and enable fast data retrieval of classified public opinion information.
[0189] This application can help companies quickly understand the changing trends of public opinion through real-time data visualization panels, which is more intuitive and clear.
[0190] This application is configured with a fluctuation alarm. When there are abnormal fluctuations in the amount of public opinion, it can respond to the changes in public opinion in a timely manner and automatically issue an alarm so that the company can know and take corresponding measures at the first time to reduce risks.
[0191] This application can help companies obtain relevant information more quickly, thereby formulating strategies and making decisions more quickly, thereby improving the company's response speed and decision-making efficiency.
[0192] This application can timely discover and correct negative public opinions that may cause crises through real-time and accurate public opinion monitoring, minimize business risks caused by public opinion, and ensure the stable and safe operation of corporate business.
[0193] Figure 4 7 is a structural diagram of a computer device provided in an embodiment of the present application. The computer device 700 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 710 (for example, one or more processors) and a memory 720, and one or more storage media 730 (for example, one or more mass storage devices) storing application programs 733 or data 732. Among them, the memory 720 and the storage medium 730 can be temporary storage or permanent storage. The program stored in the storage medium 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 700. Furthermore, the processor 710 can be configured to communicate with the storage medium 730 to execute a series of instruction operations in the storage medium 730 on the computer device 700.
[0194] The computer device 700 may further include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input and output interfaces 760, and / or one or more operating systems 731, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 4 The illustrated computer device structure does not limit the computer device and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0195] The present application also provides a computer device, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the public opinion monitoring method based on the language large model in the above-mentioned embodiments.
[0196] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the public opinion monitoring method based on a large language model.
[0197] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0198] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0199] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for monitoring public opinion based on a large language model, characterized in that: The public opinion monitoring method based on the language large model includes: Use automated collection tools to access the target webpage, simulate user operations on the target webpage and filter out raw public opinion data using keywords; The trained language model is used to classify the original public opinion data to obtain a classification result.
2. The method for monitoring public opinion based on a large language model according to claim 1 is characterized in that: The method of using the trained language model to classify the original public opinion data to obtain a classification result includes: Inserting the original public opinion data into the target queue; Monitor the amount of original public opinion data inserted into the target queue. If the amount of original public opinion data inserted into the target queue reaches a preset threshold, read out original public opinion data equal to the preset threshold from the target queue, and generate a model access request based on the read original public opinion data. Sending the model access request to the language large model by calling the batch processing interface of the language large model; The language big model is used to batch classify the original public opinion data in the model access request to obtain a classification result corresponding to each original public opinion data, wherein the classification result is one of positive public opinion, negative public opinion and neutral public opinion.
3. The method for monitoring public opinion based on a large language model according to claim 2 is characterized in that: The public opinion monitoring method based on the language large model also includes: Storing the original public opinion data in Elasticsearch; Call the Bulk interface of the Elasticsearch and use the classification results returned by the language large model to batch label and classify the corresponding original public opinion data in Elasticsearch.
4. The method for monitoring public opinion based on a large language model according to claim 1, characterized in that: Before using the trained language model to classify the original public opinion data, the public opinion monitoring method based on the language model further includes: Constructing a current prompt based on input training corpus, and inputting the current prompt into the language model, wherein the training corpus includes the current sample public opinion data to be classified, or the training corpus includes the current sample public opinion data to be classified and the corresponding public opinion category description; receiving and displaying a response result returned by the language macromodel based on the current prompt; Feedback information given by the user in response to the response result is input into the language model, so that the language model optimizes classification parameters according to the feedback information.
5. The method for monitoring public opinion based on a large language model according to any one of claims 1 to 4, characterized in that: The use of an automated acquisition tool to access a target webpage includes: Use automated collection tools to obtain the corresponding cookie values generated when successfully logging into each website; Associating and storing the cookie value with the unique website identifier of the corresponding website; Determine the target website where the target webpage to be visited is located; Match the target cookie value based on the unique identifier of the target website; Setting the target cookie value in the initial access request of the target webpage to generate a target access request; The target website is logged in and the target web page is accessed by sending the target access request to the server of the target website.
6. The method for monitoring public opinion based on a large language model according to claim 5 is characterized in that: The use of an automated collection tool to access the target webpage also includes: Using an automated collection tool to monitor the login status of the target website after sending the target access request; If the target website is not logged in or the login is invalid, instruct the user to perform the target website login operation, or simulate the user login operation according to the pre-stored account and password of the target website to log in to the target website and enter the target webpage; A new cookie value generated after successfully logging into the target website is obtained, and a local cookie file is updated using the new cookie value, wherein the local cookie file is used to store an association relationship between the cookie value and the website unique identifier of the corresponding website.
7. The method for monitoring public opinion based on a large language model according to claim 5, characterized in that: The use of an automated collection tool to access the target webpage also includes: Use automated collection tools to modify the property value of navigator.webdriver on the application side; and / or, Utilize automated collection tools to simulate the user's non-automated operations and web browsing behaviors on the target web page.
8. A public opinion monitoring system based on a large language model, characterized by: The public opinion monitoring system based on the language large model includes: The data collection module is used to access the target webpage using an automated collection tool, and simulate user operations on the target webpage to filter out raw public opinion data using keywords; The classification module is used to classify the original public opinion data using the trained language model to obtain classification results.
9. A computer device, characterized in that: The computer device includes: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the computer device executes the public opinion monitoring method based on a large language model as described in any one of claims 1-7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the public opinion monitoring method based on the language large model as described in any one of claims 1-7 is implemented.