Public opinion monitoring method and device based on large model, storage medium and product

By acquiring customer conversation texts and incoming call recordings, and using deep learning models to extract acoustic features and text data, the problem of lag and unreliability in public opinion monitoring in existing technologies has been solved, enabling early risk detection and accurate public opinion assessment.

CN121836698APending Publication Date: 2026-04-10INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2025-12-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies rely on publicly available online platforms for public opinion monitoring, resulting in high levels of noise, distortion, and delays in information delivery. Furthermore, they ignore the true emotional state of customers, leading to unreliable public opinion assessments.

Method used

By acquiring customer conversation texts and incoming call recordings, deep learning models are used to extract acoustic features and text data. Combined with business classification tags and sentiment states, public opinion early warning information is generated.

Benefits of technology

It enables early detection of public opinion risks, accurate capture of customers' true emotions, avoidance of information distortion, and improvement of the accuracy of public opinion judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121836698A_ABST
    Figure CN121836698A_ABST
Patent Text Reader

Abstract

The invention provides a public opinion monitoring method and device based on a large model, a storage medium and a product, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring voice data of a dialogue text and an incoming call record, performing text conversion on the voice data, and extracting acoustic features in the voice data; the dialogue text and the text obtained through conversion serve as text data, the text data and the acoustic features are input into a classification model based on deep learning, and a service classification label and an emotional state output by the classification model are obtained; the business classification model is used for carrying out business type classification identification according to input data and a preset hierarchical relationship of business classification labels; and public opinion early warning information is generated based on the business classification label and the emotional state. According to the method provided by the invention, the problems of data distortion and hysteresis in public opinion monitoring are avoided by using the voice data of the dialogue text and the incoming call record, and the customer emotion is accurately judged by using the acoustic characteristics in the language data, so that the accuracy of public opinion monitoring is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method, device, storage medium and product for public opinion monitoring based on a large model. Background Technology

[0003] Existing technologies typically employ methods that monitor publicly available online information. Specifically, existing public opinion monitoring solutions mainly use web crawlers and other technologies to collect and aggregate publicly available text information on social media platforms and forums. Subsequently, natural language processing (NLP) techniques are used to analyze the collected text, such as keyword matching, and simple statistical rules (such as a sudden increase in the frequency of specific keywords) are applied to determine whether a public opinion alert has been triggered.

[0004] However, the aforementioned existing technical solutions have the following drawbacks. First, their data sources rely on publicly available information from online platforms. This information is characterized by its diverse origins, high noise levels, and long dissemination chains, and often represents the "results" after a public opinion crisis has erupted. This results in an inherent lag in monitoring, making it impossible to provide early warnings. Second, since the emotional state of customers is a major factor leading to significant public opinion events, existing technologies primarily utilize textual information for public opinion assessment, ignoring the true emotional state of customers, thus making public opinion assessment unreliable. Summary of the Invention

[0005] This application provides a method, device, storage medium, and product for public opinion monitoring based on a large model, in order to solve the following technical problems in the prior art: the prior art obtains information from messy and public network platforms, the information is noisy and may be distorted after dissemination, and there is a lag in using public information for public opinion monitoring; the prior art mainly uses text information for public opinion monitoring, ignoring the real emotional state of customers, resulting in unreliable public opinion judgment.

[0006] Firstly, this application provides a public opinion monitoring method based on a large model, the method comprising:

[0007] Acquire dialogue text and incoming call recording voice data, perform text conversion on the voice data and extract acoustic features from the voice data;

[0008] The dialogue text and the converted text are used as text data. The text data and the acoustic features are input into a deep learning-based classification model to obtain the business classification labels and sentiment states output by the classification model. The business classification model is used to classify and identify business types according to the input data and the hierarchical relationship of the preset business classification labels.

[0009] Based on the aforementioned business category tags and sentiment status, public opinion early warning information is generated.

[0010] Secondly, this application provides a public opinion monitoring device based on a large model, the device comprising:

[0011] The feature extraction module is used to acquire dialogue text and incoming call recordings of voice data, perform text conversion on the voice data, and extract acoustic features from the voice data.

[0012] The classification module is used to take the dialogue text and the converted text as text data, input the text data and the acoustic features into a deep learning-based classification model, and obtain the business classification labels and sentiment states output by the classification model; the business classification model is used to classify and identify business types according to the input data and the hierarchical relationship of the preset business classification labels.

[0013] The information generation module is used to generate public opinion early warning information based on the business classification tags and sentiment status.

[0014] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0015] The memory stores computer-executed instructions;

[0016] The processor executes computer execution instructions stored in the memory to implement the method as described in the first aspect.

[0017] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in the first aspect.

[0018] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0019] The public opinion monitoring method, equipment, storage medium, and products based on large models provided in this application have the following technical effects:

[0020] 1. This method directly analyzes customer service call content by acquiring dialogue text and voice data from incoming call recordings, avoiding the distortion, amplification, and noise interference that may occur during the dissemination of information on social networks. It captures the most original and authentic demands and emotional expressions of customers, avoiding data distortion. Since customers' preferred communication method when encountering problems is remote calling, which occurs much earlier than when they express their opinions on social platforms, potential risks of public opinion can be identified earlier based on dialogue text and voice data from incoming call recordings.

[0021] 2. This method does not rely on single text analysis, but combines acoustic features in speech, which can accurately capture the strong negative emotions that customers reveal through their tone of voice in addition to the text. Since these strong negative emotions are more likely to trigger high-urgency public opinion events, this method can make more accurate judgments on public opinion by analyzing customer emotions. Attached Figure Description

[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0023] Figure 1 A flowchart illustrating the public opinion monitoring method based on a large model provided in this application embodiment. Figure 1 ;

[0024] Figure 2 A flowchart illustrating the public opinion monitoring method based on a large model provided in this application embodiment. Figure 2 ;

[0025] Figure 3 A schematic diagram of the structure of a public opinion monitoring device based on a large model provided in this application embodiment;

[0026] Figure 4 This is a schematic diagram of the electronic device structure provided in an embodiment of this application.

[0027] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.

[0030] It should be noted that the public opinion monitoring method, equipment, storage medium and products based on large models provided in this application can be used in the field of artificial intelligence, or in any field other than artificial intelligence. The application fields of the public opinion monitoring method, equipment, storage medium and products based on large models in this application are not limited.

[0031] In response to the problems of existing technologies that obtain information from chaotic and publicly available online platforms, resulting in high levels of noise and potential distortion after dissemination, and the lag in using publicly available information for public opinion monitoring, the inventors of this application have considered that they can access customers' remote inbound call channels to obtain the text of conversations between customers and the company's customer service, as well as the voice data from the call recordings. By using this first-hand, real-time, and highly authentic data for public opinion analysis, the problem of data authenticity can be solved. At the same time, since this data can directly reflect the issues that customers are most concerned about and most urgent, using this data for public opinion analysis in advance, rather than relying on publicly available information from social media platforms, can solve the problem of the lag in public opinion monitoring.

[0032] In view of the problem that existing technologies mainly use text information for public opinion monitoring, ignoring the true emotional state of customers, resulting in unreliable public opinion judgment, the inventors of this application have considered that acoustic features can be extracted from customer language data. Since these acoustic features can reflect the true emotional state of customers, customer emotions can be analyzed using acoustic features, thereby improving the accuracy of public opinion judgment.

[0033] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0034] Example 1

[0035] Figure 1 A flowchart illustrating the public opinion monitoring method based on a large model provided in this application embodiment. Figure 1 This method can be applied to servers, such as... Figure 1 As shown, the method includes:

[0036] S101. Acquire the dialogue text and the voice data of the incoming call recording, perform text conversion on the voice data and extract the acoustic features from the voice data.

[0037] In this step, the server can initiate a data monitoring service, establishing a connection with the business system (such as a call center system or online customer service platform) through a predefined application programming interface (API). The server continuously listens for and receives data streams from the business system through this interface. Specifically, for incoming call recordings, the server can receive encoded audio streams or audio files (e.g., voice data packets with a sampling rate of 8kHz, mono, and G.711 encoding), while for dialogue text, the server can receive structured or semi-structured log data (e.g., JSON format data packets containing fields such as session ID, timestamp, user identifier, and text content).

[0038] Furthermore, the server can decode the received voice data packets and convert them into a standard Pulse Code Modulation (PCM) format for subsequent processing. The server can then perform preprocessing operations on the decoded data, including but not limited to: background noise suppression using spectral subtraction-based algorithms, and endpoint detection of the audio using silence detection algorithms (such as those based on short-time energy and zero-crossing rate) to remove invalid silence segments.

[0039] In the text-to-speech stage, the server can invoke a speech recognition engine to convert the preprocessed clean audio data into text. It's important to note that this speech recognition engine can be an end-to-end deep learning model, such as one employing a Conformer or RNN-Transducer architecture. Specifically, the server inputs the preprocessed clean audio data into the speech recognition engine. The engine first extracts the Mel-frequency cepstral coefficients or filter bank features from the audio, then performs joint decoding using an acoustic and language model, ultimately outputting the text sequence with the highest probability. The server records the text output by the speech recognition engine and associates it with the session ID of the original audio.

[0040] While the aforementioned speech data is being converted to text, the server can initiate a parallel feature extraction thread to extract acoustic features from the same preprocessed clean audio data. For example, the amplitude of the sound signal (generally, high-intensity emotions such as anger and excitement are accompanied by higher amplitudes, while emotions such as sadness and fatigue correspond to lower amplitudes) and the contour of sound energy (i.e., the curve of sound energy change over time. For example, anger may be characterized by sudden bursts and violent fluctuations in energy, while calm emotions correspond to a stable energy contour).

[0041] Optionally, acoustic features are extracted from the speech data, specifically including:

[0042] Extract acoustic features from speech data that include at least one of the following: speech rate, pitch, and level of emotional arousal.

[0043] Specifically, the above acoustic features can be extracted using the following feature extraction operations:

[0044] Speech rate: The server first segments the text output by the speech recognition engine, then counts the number of valid words, and divides it by the effective duration of the audio (after removing silent segments) to obtain the speech rate value in "words / second".

[0045] Pitch (fundamental frequency): The server extracts the fundamental frequency (F0) from the audio signal using autocorrelation or cepstral method, and calculates the mean, maximum, minimum and standard deviation of the fundamental frequency within a segment to characterize the pitch and fluctuation.

[0046] Emotional arousal level: The server calculates the short-time energy (loudness) of the audio and analyzes its dynamic range. Simultaneously, it can extract advanced features such as spectral centroid and roll-off point. Emotional arousal is typically characterized by increased energy and a raised spectral centroid.

[0047] The server can combine the extracted acoustic features into corresponding acoustic feature vectors. Finally, the server generates a corresponding sound feature vector for each segment of speech data and stores the text output by the language recognition model and the original dialogue text into a temporary database for subsequent processing.

[0048] S102. The dialogue text and the converted text are used as text data. The text data and acoustic features are input into a deep learning-based classification model to obtain the business classification labels and sentiment states output by the classification model. The business classification model is used to classify and identify business types according to the input data and the hierarchical relationship of the preset business classification labels.

[0049] In this step, the server reads all text data for the current batch from a temporary database, including dialogue text directly obtained from the customer service system and text converted through speech recognition. Further, before inputting the text data and acoustic features into the deep learning-based classification model, the server can convert the text data into standardized text data in a uniform format according to a preset standardization format. For example, the server can uniformly convert the text data to UTF-8 encoding, remove special characters and extra spaces, and perform spell correction and standardization processing. Since text data from different channels (such as speech-to-text and online customer service text) may have differences in format, encoding, punctuation usage, etc., standardization processing can eliminate this "noise," providing clean and consistent input data for the deep learning model. This reduces the computational burden or judgment errors caused by data format chaos, improving the accuracy of business classification and sentiment analysis.

[0050] Furthermore, the server can associate each cleaned text data with its corresponding acoustic feature vector. Each associated data sample contains two parts: a text sequence and an acoustic feature vector. The server then batches these associated data samples into a pre-trained deep learning classification model for inference. It should be noted that this deep learning classification model can, for example, employ a multi-task learning framework. For business classification, the deep learning classification model can output a probability distribution of dimension N, where N is the total number of preset business classification labels, and the label with the highest probability is the predicted business classification result. Hierarchical classification can be performed during the business classification stage. Taking banking business as an example, the major category (e.g., "personal loans") can be determined first, followed by the minor category (e.g., "interest rate consultation").

[0051] For sentiment classification, a deep learning classification model can output a probability distribution of dimension M, where M is the number of sentiment state categories (e.g., anger, anxiety, calm, satisfaction). The category with the highest probability is the predicted sentiment state.

[0052] Finally, the server can record the business classification label and sentiment status label output by the model for each data sample, along with the corresponding confidence level.

[0053] S103. Generate public opinion early warning information based on business classification tags and sentiment status.

[0054] In this step, the server can use a sliding time window (e.g., a window size of 30 minutes and a sliding step size of 1 minute) to perform real-time aggregation calculations on the analysis results continuously output by the previous steps. Specifically, for each business category tag (especially sub-category tags), the server can, for example, count the total number of sessions that appear within the time window, the number of sessions marked as negative emotions (such as "anger" or "anxiety") and their proportion, and the month-on-month change rate of the number of sessions (compared to the previous time window).

[0055] The server can be configured with an alert rule engine, which can compare the above statistical results with preset alert rules. For example, the alert rules can be:

[0056] Conditions: Within a 30-minute window, the total number of conversations for a certain business subcategory (such as "transfer failure") exceeds 50, and the proportion of conversations identified as having negative sentiment exceeds 40%, and the total number of conversations increases by more than 100% month-on-month.

[0057] Once the conditions of the warning rules are met, the warning rule engine triggers a warning. The server then generates a structured public opinion warning message, which may include, for example, the warning ID, trigger time, associated business category tags (from broad category to subcategory), core negative sentiment type, related session statistics, and trigger rule details. Subsequently, the server can publish the warning event to a specified topic by calling the internal message queue API. Other business systems subscribed to this topic (such as monitoring dashboard systems and work order systems) can then consume and display the warning information in real time.

[0058] The public opinion monitoring method provided in this embodiment has the following technical effects:

[0059] 1. This method directly analyzes customer service call content by acquiring dialogue text and voice data from incoming call recordings, avoiding the distortion, amplification, and noise interference that may occur during the dissemination of information on social networks. It captures the most original and authentic demands and emotional expressions of customers, avoiding data distortion. Since customers' preferred communication method when encountering problems is remote calling, which occurs much earlier than when they express their opinions on social platforms, potential risks of public opinion can be identified earlier based on dialogue text and voice data from incoming call recordings.

[0060] 2. This method does not rely on single text analysis, but combines acoustic features in speech, which can accurately capture the strong negative emotions that customers reveal through their tone of voice in addition to the text. Since these strong negative emotions are more likely to trigger high-urgency public opinion events, this method can make more accurate judgments on public opinion by analyzing customer emotions.

[0061] Figure 2A flowchart illustrating the public opinion monitoring method based on a large model provided in this application embodiment. Figure 2 ,like Figure 2 As shown,

[0062] S201. Obtain the dialogue text and the voice data of the incoming call recording, perform text conversion on the voice data and extract the acoustic features from the voice data.

[0063] The specific implementation process of this step can be referred to in Embodiment 1, which will not be elaborated here.

[0064] S202. The dialogue text and the converted text are used as text data. The text data and acoustic features are input into a deep learning-based classification model to obtain the business classification label and sentiment state output by the classification model. The business classification model extracts the semantic features of the text data through the feature extraction layer, processes the semantic features and acoustic features through the fully connected layer, and outputs the business classification label and sentiment state.

[0065] In this step, the server first feeds the text data into the model's feature extraction layer, which converts the text into a high-dimensional semantic feature vector. Then, the server concatenates the semantic feature vector with the acoustic feature vector to form a multimodal fusion feature. This fusion feature is then fed into the fully connected layer of the deep learning model for classification decisions. The fully connected layer outputs two results in parallel: a business classification label (identified according to a pre-defined hierarchical system of 36 major categories and 159 minor categories) and an emotional state (such as anger, anxiety, calmness, etc., and their confidence levels).

[0066] It's important to note that the feature extraction layer can be composed of a pre-trained large language model. Its task is to convert each word or token in the text into a high-dimensional numerical vector (called a word vector or embedding vector), thus transforming unstructured text data into a mathematical representation that the model can process. Subsequently, through the self-attention mechanism in the Transformer architecture, the feature extraction layer analyzes the relationship between each word and all other words in the text, thereby understanding the contextual meaning of the words. For example, in the sentence "My mobile banking transfer failed," the model can capture the specific meaning of the word "transfer" in the context of "mobile banking." After processing the text data, the feature extraction layer outputs a semantic feature vector rich in contextual information. Through the feature extraction layer, a deep semantic understanding of customer needs can be achieved, enabling accurate business classification.

[0067] The fully connected layer is responsible for comprehensively judging the fused features and outputting the final classification result. Specifically, the server concatenates the semantic feature vector output by the feature extraction layer with the acoustic feature vector directly extracted from the speech data to form a rich multimodal fusion feature vector, thereby associating language content with speech emotion signals. This fusion process allows the model to learn the association between text semantics and sound features simultaneously in a unified feature space. For example, the model can learn that when the text contains "transfer failed" accompanied by "rapid speech," the confidence level of its negative emotion is much higher than when either feature appears alone. This early fusion strategy allows the model to explore deeper cross-modal associations, improving the accuracy of classification and sentiment judgment.

[0068] Next, the fused feature vector is fed into a fully connected layer, which consists of multiple layers of neurons. Each neuron is connected to all outputs of the previous layer. The fully connected layer performs complex nonlinear transformations on the input features through a series of weighted sums and nonlinear activation functions (such as ReLU and Softmax). During the transformation process, the model uses knowledge learned during training (such as the high correlation between "rapid speech" and "anger" and "financial service" complaints) to make decisions and ultimately outputs two results in parallel:

[0069] Business category tags: Based on a preset hierarchical classification system (such as 36 major categories and 159 minor categories), output the business category to which this dialogue is most likely to belong.

[0070] Emotional state: Output the emotional tendency (such as anger, anxiety, calm) and the corresponding confidence level of the dialogue.

[0071] In this embodiment, the fully connected layer, by fusing textual semantics and acoustic features, can perform cross-modal comprehensive analysis, making the judgments of the deep learning model closer to the comprehensive cognitive level of humans, and improving the accuracy and reliability of business classification and emotional state recognition.

[0072] S203. Determine the priority of public opinion events based on the hierarchical relationship of business category tags; the hierarchical relationship is that each parent category tag includes at least one child category tag;

[0073] In this step, the server first parses the business classification labels output by the deep learning model to determine the priority of public opinion events. For example, if the server identifies the business subcategory as "deposit security rumors", since its parent category "personal account" is a high-risk business, the server can assign a high priority to the event according to the built-in rule base (e.g., the priority of "deposit security" events is "extremely high"). In contrast, the priority of "page crash" events is defined as "medium" or "low".

[0074] S204. Quantify and score the emotional state to obtain the intensity assessment results corresponding to the emotional state, and determine the urgency of the warning based on the intensity assessment results;

[0075] In this step, the server can use a weighted calculation method to quantify and score the output of emotional state. For example, the server can weight the confidence level of the "anger" emotion (e.g., 0.95) with the score of emotional agitation in the acoustic features (e.g., rapid speech rate score) to obtain a comprehensive emotional intensity value between 0 and 1, which serves as the intensity assessment result corresponding to the anger emotion. After obtaining the comprehensive emotional intensity value, the server can determine the urgency of the warning based on preset thresholds. For example, the preset thresholds can be 0.8 or 0.6, with a comprehensive emotional intensity value > 0.8 indicating "urgent" and a comprehensive emotional intensity value > 0.6 indicating "important".

[0076] S205. Generate public opinion warning information based on priority and urgency.

[0077] In this step, the server can combine event priority (from S203) and warning urgency (from S204) to generate the final public opinion warning information. For example, if an event is determined to have extremely high priority and high urgency, the server will generate an emergency warning message. The warning message content includes an event description, business category, sentiment analysis summary, and timestamp. The server can generate different levels of warning messages based on different priorities and warning urgency levels. Warning messages can be identified by different colors, for example, red for the highest level and yellow for the second highest level.

[0078] The S203-S205 method, by linking and integrating information from two different dimensions—the importance (priority) of the business and the external emotional expression (urgency) of users—can more accurately assess the true risk level of potential public opinion.

[0079] S206. Push public opinion early warning information to the designated early warning notification module;

[0080] In this step, after generating the warning information, the server can push the warning information data packet to the warning notification service (i.e., the warning notification module) responsible for message distribution through an internal message queue or by directly calling the API.

[0081] S207. Send public opinion warning information via email or SMS through the warning notification module.

[0082] Specifically, after receiving the information, the early warning notification service can immediately send the early warning information to the relevant personnel's mobile devices or email addresses through an integrated SMS gateway or email server, based on a pre-set subscription list (such as the person in charge of the business line or the operations and maintenance team).

[0083] The methods S206-S207 can automatically and instantly push early warning information to the specific responsible persons, ensuring that public opinion incidents can be handled as soon as possible.

[0084] Figure 3 This is a schematic diagram of the structure of a public opinion monitoring device based on a large model provided in an embodiment of this application, as shown below. Figure 3 As shown, the device 30 includes:

[0085] The feature extraction module 301 is used to acquire dialogue text and incoming call recording voice data, perform text conversion on the voice data, and extract acoustic features from the voice data.

[0086] The classification module 302 is used to take the dialogue text and the converted text as text data, input the text data and acoustic features into a deep learning-based classification model, and obtain the business classification labels and sentiment states output by the classification model; the business classification model is used to classify and identify business types according to the input data and the hierarchical relationship of the preset business classification labels.

[0087] The information generation module 303 is used to generate public opinion early warning information based on business classification tags and sentiment status.

[0088] Figure 4 This is a schematic diagram of the electronic device structure provided in the embodiments of this application, such as... Figure 4 As shown, the device 40 includes at least one processor 401 and a memory 402. Optionally, the device 40 also includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.

[0089] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.

[0090] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0091] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0092] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0093] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0094] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0095] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0096] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0097] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.

[0098] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0099] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0100] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0101] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A public opinion monitoring method based on a large model, characterized in that, The method includes: Acquire dialogue text and incoming call recording voice data, perform text conversion on the voice data and extract acoustic features from the voice data; The dialogue text and the converted text are used as text data. The text data and the acoustic features are input into a deep learning-based classification model to obtain the business classification labels and sentiment states output by the classification model. The business classification model is used to classify and identify business types according to the input data and the hierarchical relationship of the preset business classification labels. Based on the aforementioned business category tags and sentiment status, public opinion early warning information is generated.

2. The method according to claim 1, characterized in that, The generation of public opinion early warning information based on the business classification tags and sentiment status specifically includes: The priority of public opinion events is determined based on the hierarchical relationship of the business category tags; the hierarchical relationship is that each parent category tag includes at least one child category tag. The emotional state is quantitatively scored to obtain the intensity assessment result corresponding to the emotional state, and the urgency of the warning is determined based on the intensity assessment result; Based on the priority and the urgency of the warning, public opinion warning information is generated.

3. The method according to claim 1, characterized in that, The extraction of acoustic features from the speech data specifically includes: The acoustic features extracted from the speech data include at least one of the following: speech rate, pitch, and level of emotional arousal.

4. The method according to claim 3, characterized in that, The step of inputting the text data and the acoustic features into a deep learning-based classification model specifically includes: The extracted acoustic features are fused with text information to obtain multimodal data; The multimodal data is input into the deep learning-based classification model.

5. The method according to claim 1, characterized in that, Before inputting the text data and the acoustic features into the deep learning-based classification model, the method further includes: The text data is converted into standardized text data in a unified format according to a preset standardized format.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: The public opinion warning information is pushed to the designated warning notification module; The public opinion warning information is sent via email or SMS through the warning notification module.

7. The method according to any one of claims 1-5, characterized in that, The classification model includes a feature extraction layer and a fully connected layer. The step of classifying and identifying business types based on the input data and according to the preset hierarchical relationship of business classification labels specifically includes: The semantic features of the text data are extracted through the feature extraction layer; The semantic and acoustic features are processed by the fully connected layer to output business classification labels and sentiment states.

8. A public opinion monitoring device based on a large model, characterized in that, The device includes: The feature extraction module is used to acquire dialogue text and incoming call recordings of voice data, perform text conversion on the voice data, and extract acoustic features from the voice data. The classification module is used to take the dialogue text and the converted text as text data, input the text data and the acoustic features into a deep learning-based classification model, and obtain the business classification labels and sentiment states output by the classification model; the business classification model is used to classify and identify business types according to the input data and the hierarchical relationship of the preset business classification labels. The information generation module is used to generate public opinion early warning information based on the business classification tags and sentiment status.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.