Multi-modal public opinion data intelligent classification and grading method, device and equipment
By converting multimodal public opinion data into a unified text mode and using large language models for intelligent classification and grading, the problem of insufficient cross-modal integration and dynamic adaptability in traditional public opinion analysis is solved, and efficient and flexible public opinion analysis is achieved.
Patent Information
- Application Number
- CN202510532061.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional public opinion analysis methods rely on single-modal data, lack cross-modal information fusion capabilities, lack intelligence, weak dynamic adaptability, and difficult to respond to new public opinion events quickly. The classification system adjustment requires retraining the model, which lacks flexibility.
By uniformly representing multimodal public opinion data as text modes, using large language models for intelligent classification and grading, building public opinion classification and grading instructions, dynamically adjusting the public opinion classification and grading system, and avoiding retraining the model.
It has realized efficient and intelligent classification and grading of multimodal public opinion data, maintained good generalization performance and intelligence level, and responded flexibly to changes in public opinion, improving the efficiency and accuracy of public opinion analysis.
Smart Images

Figure CN120448543A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of Internet information processing technology, and specifically relates to a method, device and equipment for intelligent classification and grading of multimodal public opinion data. Background Art
[0002] The impact of online public opinion on government governance, corporate brand image, and social stability is becoming increasingly significant. With the rapid development of the internet, data from multiple sources, including social media, news portals, and short video platforms, has exploded, and traditional public opinion analysis methods are increasingly unable to keep pace with this development.
[0003] Traditional public opinion analysis relies primarily on single-modal data (such as text analysis) and focuses primarily on text analysis. It lacks the ability to integrate cross-modal information, resulting in low efficiency in processing unstructured data. Secondly, due to insufficient intelligence and reliance on manual rules and feature engineering, traditional machine learning models have poor generalization capabilities when faced with complex public opinion scenarios and struggle to effectively understand semantic associations. In addition, current public opinion classification models are all fixed categories, and users cannot dynamically adjust them, resulting in weak dynamic adaptability of the system and an inability to quickly respond to new public opinion events. Adjustments to the classification system and grading standards often require retraining the model, which lacks flexibility. Summary of the Invention
[0004] In response to the above problems, the embodiments of the present disclosure propose a multimodal public opinion data intelligent classification and grading solution based on a large language model.
[0005] A first aspect of the embodiments of the present disclosure provides a method for intelligently classifying and grading multimodal public opinion data, including:
[0006] Acquire multimodal public opinion data and represent the multimodal public opinion data in a unified manner, wherein the multimodal public opinion data is one or more of video, image, audio, and text data;
[0007] Constructing a public opinion classification and grading instruction based on the uniformly represented multimodal public opinion data and a preset public opinion classification and grading system;
[0008] The public opinion classification and grading instructions are input into a large language model, and the classification and grading results of the multimodal public opinion data are output.
[0009] In some embodiments of the present disclosure, obtaining multimodal public opinion data includes:
[0010] The multimodal public opinion data is collected from the Internet.
[0011] In some embodiments of the present disclosure, the unified representation of the multimodal public opinion data includes:
[0012] The multimodal public opinion data is uniformly represented as text data.
[0013] In some embodiments of the present disclosure, the step of uniformly representing the multimodal public opinion data as text data includes:
[0014] Converting the video data into image data and audio data, and using an image description model to convert the image data into a text describing the image content, thereby generating first text data;
[0015] Extracting text data from the audio data using a speech recognition model to generate second text data;
[0016] The first text data, the second text data, and the text data are aggregated into unified text data.
[0017] In some embodiments of the present disclosure, converting the video data into image data and audio data includes:
[0018] Performing frame cutting on the video data to obtain frame image data;
[0019] Extract the audio from the video data to obtain audio data.
[0020] In some embodiments of the present disclosure, constructing a public opinion classification and grading instruction based on the uniformly represented multimodal public opinion data and a preset public opinion classification and grading system includes:
[0021] Based on the unified text data and the public opinion classification and grading system, a public opinion classification and grading prompt word is constructed, and the public opinion classification and grading prompt word is used to instruct the large language model to classify the unified text data based on the public opinion classification and grading system.
[0022] In some embodiments of the present disclosure, the method further includes:
[0023] Adjust the public opinion classification and grading system.
[0024] In some embodiments of the present disclosure, adjusting the public opinion classification and grading system includes:
[0025] The public opinion classification and grading system is adjusted according to the classification and grading results of the multimodal public opinion data.
[0026] A second aspect of the embodiments of the present disclosure provides a device for intelligently classifying and grading multimodal public opinion data, including:
[0027] An acquisition module, configured to acquire multimodal public opinion data and uniformly represent the multimodal public opinion data, wherein the multimodal public opinion data is one or more of video, image, audio, and text data;
[0028] A construction module, configured to construct a public opinion classification and grading instruction based on the uniformly represented multimodal public opinion data and a preset public opinion classification and grading system;
[0029] The inference module is used to input the public opinion classification and grading instructions into the large language model and output the classification and grading results of the multimodal public opinion data.
[0030] A third aspect of the present disclosure provides a multimodal public opinion data intelligent classification and grading device, including a memory and a processor.
[0031] The memory is used to store computer programs;
[0032] The processor is configured to implement the method according to the first aspect of the present disclosure when executing the computer program.
[0033] In summary, the methods, devices, and equipment for intelligent classification and grading of multimodal public opinion data provided by the embodiments of the present disclosure convert all multimodal public opinion data into a unified text modality, making it possible to use large language model technology for intelligent classification and grading. Thanks to the powerful understanding ability of the large language model, it can maintain good generalization performance and intelligence level without the need for complex public opinion judgment rules and logic; at the same time, based on the large language model, this method can build the user's preset public opinion classification and grading indicator system into the instructions, so as to perform flexible public opinion judgment. When the indicator system changes, there is no need to retrain the model, which ensures the flexibility of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The features and advantages of the present disclosure will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present disclosure in any way. In the accompanying drawings:
[0035] Figure 1 This is the overall framework diagram of the multimodal public opinion data intelligent classification and grading algorithm based on the large language model shown in this disclosure;
[0036] Figure 2 This is a flowchart of a method for intelligent classification and grading of multimodal public opinion data according to some embodiments of the present disclosure;
[0037] Figure 3 This is an example of multimodal public opinion data collected from the Internet, including video, audio, and text, and its processing flow;
[0038] Figure 4 is a schematic diagram of a multimodal public opinion data intelligent classification and grading device according to some embodiments of the present disclosure;
[0039] Figure 5This is a schematic diagram of a multimodal public opinion data intelligent classification and grading device according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0040] In the detailed description that follows, many specific details of the present disclosure are set forth by way of example in order to provide a thorough understanding of the relevant disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure can be implemented without these details. It should be understood that the use of the terms "system," "device," "unit," and / or "module" in the present disclosure is a method for distinguishing between different parts, elements, parts, or assemblies at different levels in a sequential arrangement. However, these terms may be replaced by other expressions if they can achieve the same purpose.
[0041] It should be understood that when a device, unit, or module is referred to as being "on," "connected to," or "coupled to" another device, unit, or module, it may be directly on, connected to, coupled to, or in communication with the other device, unit, or module, or there may be intervening devices, units, or modules, unless the context clearly indicates an exception. For example, the term "and / or" as used in this disclosure includes any and all combinations of one or more of the associated listed items.
[0042] The terms used in this disclosure are only for describing specific embodiments and are not intended to limit the scope of this disclosure. As shown in the specification and claims of this disclosure, unless the context clearly indicates an exception, the words "a", "an", "a kind" and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of clearly identified features, wholes, steps, operations, elements and / or components, and such expressions do not constitute an exclusive list, and other features, wholes, steps, operations, elements and / or components may also be included.
[0043] These and other features and characteristics of the present disclosure, as well as the methods of operation, the functions of the related elements of the structure, the combination of parts, and the economy of manufacture may be better understood with reference to the following description and accompanying drawings, which form a part of this specification. However, it is to be expressly understood that the drawings are for illustration and description purposes only and are not intended to limit the scope of protection of the present disclosure. It is to be understood that the drawings are not drawn to scale.
[0044] Various structural diagrams are used in this disclosure to illustrate various variations of the embodiments of the present disclosure. It should be understood that the preceding or following structures are not intended to limit the present disclosure. The scope of protection of the present disclosure is subject to the claims.
[0045] With the rapid development of the internet, multi-source data from social media, news portals, and short video platforms has exploded, and the impact of online public opinion on government governance, corporate brand image, and social stability has become increasingly significant. Consequently, government agencies, enterprises, and institutions are increasingly demanding public opinion monitoring, hoping to achieve real-time monitoring, analysis, and early warning through efficient, accurate, and intelligent technologies. However, traditional public opinion analysis primarily relies on single-modal data (such as text analysis) and fails to fully utilize multimodal information (such as images, audio, and video). This results in low information utilization, incomplete analytical data, and difficulty in accurately understanding and identifying public opinion trends, resulting in limitations.
[0046] Existing public opinion analysis technologies still have numerous shortcomings in terms of multimodal data processing, intelligence, dynamic adaptability, and interpretability. First, multimodal data processing capabilities are limited. Current methods primarily focus on text analysis and lack the ability to integrate cross-modal information, resulting in low efficiency in processing unstructured data. Second, due to insufficient intelligence and reliance on manual rules and feature engineering, traditional machine learning models have poor generalization capabilities when faced with complex public opinion scenarios and struggle to effectively understand semantic associations. Furthermore, current public opinion classification models are all fixed categories and cannot be dynamically adjusted by users, resulting in weak dynamic adaptability and an inability to quickly respond to new public opinion events. Adjustments to the classification system and grading standards often require retraining the model, lacking flexibility. Finally, current mainstream public opinion analysis models suffer from poor interpretability, with a "black box" decision-making process and a lack of a transparent credibility assessment system. This makes it difficult to trace the results and hinders accurate analysis and decision support.
[0047] In order to solve the above problems, this paper proposes a method for intelligent classification and grading of multimodal public opinion data based on a large language model. By converting multimodal public opinion data into a unified text mode, the large language model technology is used to intelligently classify and grade the multimodal public opinion data. The overall framework of the multimodal public opinion data intelligent classification and grading algorithm shown in this paper is as follows: Figure 1 In some embodiments, the flowchart of the multimodal public opinion data intelligent classification and grading method is as follows: Figure 2 As shown, the specific steps include:
[0048] S210, obtaining multimodal public opinion data, and uniformly representing the multimodal public opinion data, wherein the multimodal public opinion data is one or more of video, image, audio, and text data.
[0049] The multimodal public opinion data in this disclosure is one or more of video, image, audio, and text data collected from the Internet. This disclosure uniformly represents the multimodal public opinion data in text form. Specifically:
[0050] First, the video data is preprocessed, the video data is frame-cut, the frame image data is obtained, and the audio is extracted to obtain the audio data;
[0051] Then, based on the image description model, the image data (including the frame image data) is converted into a first text describing the image content;
[0052] And extract the second text in the audio based on the speech recognition model.
[0053] Finally, the first character, the second character and the character data are integrated into unified text data.
[0054] Figure 3 This is an example of multimodal public opinion data collected from the Internet, including video, audio, and text, and its processing flow.
[0055] Figure 3 The multimodal public opinion data collected in the survey include videos and video texts.
[0056] The video's text is plain text and reads: A well-known electronics company has recently received numerous complaints from consumers alleging repeated quality issues with its products, poor after-sales service, and shirking of responsibility, resulting in serious damage to consumer rights. Consumers have repeatedly contacted the company's customer service, but their issues remain unresolved, with some even being disconnected. Consumers have expressed dissatisfaction and called on the company to prioritize product quality and after-sales service to ensure consumer rights are effectively protected.
[0057] The video copy also includes multiple hot tags: #consumer rights #product quality #after-sales service #business integrity.
[0058] The image description model is used to generate a video description based on the frame image data obtained after video data segmentation. The video shows a storefront of a well-known electronics company. In the center of the screen, a few lines of text read: "Products have repeatedly experienced quality issues, and after-sales service has shirked responsibility and failed to resolve them, harming consumer rights."
[0059] The text extraction of the audio data extracted from the video data did not obtain any valid information.
[0060] Finally, the video copy and video description containing effective information are integrated into a textual summary:
[0061] Information summary:
[0062] Video description: The video shows a store of a well-known electronics company. In the middle of the screen are written a few lines of words: "The product has repeatedly had quality problems, the after-sales service is evasive and the problem is not resolved, and the rights of consumers are damaged."
[0063] Video caption: A well-known electronics company has recently received numerous complaints from consumers alleging repeated quality issues with its products, poor after-sales service, and shirking of responsibility, resulting in serious damage to consumer rights. Consumers have repeatedly contacted the company's customer service, but their issues remain unresolved, and in some cases, the company has even hung up on them. Consumers are expressing their dissatisfaction and calling on the company to prioritize product quality and after-sales service and effectively protect consumer rights. #ConsumerRights #ProductQuality #AfterSalesService #BusinessIntegrity
[0064] S220, constructing a public opinion classification and grading instruction based on the uniformly represented multimodal public opinion data and a preset public opinion classification and grading system.
[0065] This paper first constructs a public opinion classification and grading system based on prior knowledge. Figure 3 The classification system includes:
[0066] Public opinion categories include product quality, service complaints, commercial fraud, and labor disputes.
[0067] The public opinion level is 1-5, with level 1 being the mildest and level 5 being the most severe.
[0068] Then, a public opinion classification and grading instruction is constructed based on the information summary in text form and the public opinion classification and grading system. In some embodiments of the present disclosure, the public opinion classification and grading instruction is a prompt word (prompt) of the large language model, and the indicator word is used to instruct the large language model to classify and grade the information summary in text form based on the public opinion classification and grading system.
[0069] Figure 3 The prompt words constructed in are:
[0070] #Task
[0071] You are a public opinion analysis assistant and need to analyze the public opinion category and level corresponding to the given materials.
[0072] #Public Opinion Classification and Grading System
[0073] Public opinion categories include product quality, service complaints, commercial fraud, and labor disputes.
[0074] The public opinion level is 1-5, with level 1 being the mildest and level 5 being the most severe.
[0075] #Material Content
[0076] Video description: The video shows a store of a well-known electronics company. In the middle of the screen are written a few lines of words: "The product has repeatedly had quality problems, the after-sales service is evasive and the problem is not resolved, and the rights of consumers are damaged."
[0077] Video caption: A well-known electronics company has recently received numerous complaints from consumers alleging repeated quality issues with its products, poor after-sales service, and shirking of responsibility, resulting in serious damage to consumer rights. Consumers have repeatedly contacted the company's customer service, but their issues remain unresolved, and in some cases, the company has even hung up on them. Consumers are expressing their dissatisfaction and calling on the company to prioritize product quality and after-sales service and effectively protect consumer rights. #ConsumerRights #ProductQuality #AfterSalesService #BusinessIntegrity
[0078] In this disclosure, the public opinion classification and grading system can be adjusted according to needs, such as adding public opinion categories or subdividing public opinion levels. By flexibly adjusting the public opinion classification and grading system, it is possible to quickly respond to new public opinion events and conduct refined management based on the subdivided public opinion levels.
[0079] In other embodiments of the present disclosure, the public opinion classification and grading system can be dynamically adjusted according to the public opinion classification and grading results output by the large language model, thereby dynamically optimizing the public opinion classification and grading system.
[0080] S230: Input the public opinion classification and grading instruction into a large language model, and output the classification and grading result of the multimodal public opinion data.
[0081] The present disclosure inputs the constructed public opinion classification and grading prompt words into a large language model for inference, and outputs the classification and grading results of the multimodal public opinion data.
[0082] Figure 3 The inference results output in are:
[0083] Public opinion category: Product quality
[0084] Public opinion level: Level 3
[0085] This disclosure benefits from the powerful understanding ability of the large language model to intelligently classify and grade multimodal public opinion data, while maintaining good generalization performance and intelligence level without the need for complex public opinion judgment rules and logic.
[0086] Figure 4 Schematic diagram of a multimodal public opinion data intelligent classification and grading device according to some embodiments of the present disclosure. Figure 4 As shown, the multimodal public opinion data intelligent classification and grading device 400 includes an acquisition module 410, a construction module 420, and an inference module 430. Among them:
[0087] An acquisition module 410 is configured to acquire multimodal public opinion data and to uniformly represent the multimodal public opinion data, wherein the multimodal public opinion data is one or more of video, image, audio, and text data;
[0088] A construction module 420 is used to construct a public opinion classification and grading instruction based on the uniformly represented multimodal public opinion data and a preset public opinion classification and grading system;
[0089] The reasoning module 430 is used to input the public opinion classification and grading instructions into the large language model and output the classification and grading results of the multimodal public opinion data.
[0090] Figure 5 Schematic diagram of a multimodal public opinion data intelligent classification and grading device according to some embodiments of the present disclosure. Figure 5 As shown, the multimodal public opinion data intelligent classification and grading device 500 includes a memory 520 and a processor 510, wherein the memory 520 is used to store a computer program; the processor 510 is used to implement when executing the computer program. Figure 2 The intelligent classification and grading method for multimodal public opinion data described in S210-S230.
[0091] In summary, the methods, devices, and equipment for intelligent classification and grading of multimodal public opinion data provided by the embodiments of the present disclosure convert all multimodal public opinion data into a unified text modality, making it possible to use large language model technology for intelligent classification and grading. Thanks to the powerful understanding ability of the large language model, it can maintain good generalization performance and intelligence level without the need for complex public opinion judgment rules and logic; at the same time, based on the large language model, this method can build the user's preset public opinion classification and grading indicator system into the instructions, so as to perform flexible public opinion judgment. When the indicator system changes, there is no need to retrain the model, which ensures the flexibility of the method.
[0092] Although the subject matter described herein is provided in the general context of being executed in conjunction with the execution of an operating system and application programs on a computer system, those skilled in the art will recognize that other implementations may also be performed in conjunction with other types of program modules. Generally speaking, program modules include routines, programs, components, data structures, and other types of structures that perform specific tasks or implement specific abstract data types. Those skilled in the art will appreciate that the subject matter described herein may be practiced using other computer system configurations, including handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like, and may also be used in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.
[0093] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0094] It should be understood that the above-described specific embodiments of the present disclosure are merely illustrative of or explanation of the principles of the present disclosure and do not constitute limitations on the present disclosure. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present disclosure shall be included within the scope of protection of the present disclosure. In addition, the claims appended to the present disclosure are intended to cover all variations and modifications that fall within the scope and metes and bounds of the appended claims, or equivalents of such scope and metes and bounds.
Claims
1. A method for intelligent classification and grading of multimodal public opinion data, characterized in that: include: Acquire multimodal public opinion data and represent the multimodal public opinion data in a unified manner, wherein the multimodal public opinion data is one or more of video, image, audio, and text data; Constructing a public opinion classification and grading instruction based on the uniformly represented multimodal public opinion data and a preset public opinion classification and grading system; The public opinion classification and grading instructions are input into a large language model, and the classification and grading results of the multimodal public opinion data are output.
2. The method according to claim 1, characterized in that The obtaining of multimodal public opinion data includes: The multimodal public opinion data is collected from the Internet.
3. The method according to claim 1, characterized in that The unified representation of the multimodal public opinion data includes: The multimodal public opinion data is uniformly represented as text data.
4. The method according to claim 3, characterized in that The step of uniformly expressing the multimodal public opinion data as text data includes: Converting the video data into image data and audio data, and using an image description model to convert the image data into a text describing the image content, thereby generating first text data; Extracting text data from the audio data using a speech recognition model to generate second text data; The first text data, the second text data, and the text data are aggregated into unified text data.
5. The method according to claim 4, characterized in that: The converting of the video data into image data and audio data comprises: Performing frame cutting on the video data to obtain frame image data; Extract the audio from the video data to obtain audio data.
6. The method according to claim 4, characterized in that: The public opinion classification and grading instructions based on the multimodal public opinion data represented in a unified manner and a preset public opinion classification and grading system include: Based on the unified text data and the public opinion classification and grading system, a public opinion classification and grading prompt word is constructed, and the public opinion classification and grading prompt word is used to instruct the large language model to classify the unified text data based on the public opinion classification and grading system.
7. The method according to claim 1, characterized in that: The method further comprises: Adjust the public opinion classification and grading system.
8. The method according to claim 7, characterized in that: The adjustment of the public opinion classification and grading system includes: The public opinion classification and grading system is adjusted according to the classification and grading results of the multimodal public opinion data.
9. A multimodal public opinion data intelligent classification and grading device, characterized in that: include: An acquisition module, configured to acquire multimodal public opinion data and uniformly represent the multimodal public opinion data, wherein the multimodal public opinion data is one or more of video, image, audio, and text data; A construction module, configured to construct a public opinion classification and grading instruction based on the uniformly represented multimodal public opinion data and a preset public opinion classification and grading system; The inference module is used to input the public opinion classification and grading instructions into the large language model and output the classification and grading results of the multimodal public opinion data.
10. A multimodal public opinion data intelligent classification and grading device, characterized in that: including memory and processor, The memory is used to store computer programs; The processor is configured to implement the method according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Intelligent news public opinion early warning system
CN115098756A
Public opinion event monitoring method based on multi-modal fusion large model
CN117350301A
Online public opinion risk event monitoring, studying and judging method
CN117910575A
Public opinion processing method and device, equipment, medium and program product
CN117933204A
Chemical safety inspection method and system based on large language model
CN118333409A