Social media text denoising method based on large model, electronic equipment and computer readable storage medium
Through a large-model-based social media text denoising method, utilizing atomic category labels, multi-label classification, and LoRA fine-tuning, the adaptability and accuracy issues in different business scenarios in social media data processing are solved, achieving efficient and low-cost data classification and denoising.
Patent Information
- Application Number
- CN202510871941.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-30
AI Technical Summary
Existing social media data processing technologies are unable to effectively distinguish and process data according to the needs of different business scenarios, resulting in the loss of valid data and the proliferation of invalid data. In addition, existing text classification technologies ignore multi-dimensional features when processing social media data, resulting in inaccurate classification results.
A large-model-based social media text denoising method is adopted to achieve refined classification and denoising of social media text through the combination of atomic category labels, multi-label classification, LoRA fine-tuning and vector database technologies.
It improves the flexibility and adaptability of data processing, enhances the accuracy of classification, reduces costs, and maintains the efficiency of data processing through text cleaning and vector retrieval technology.
Smart Images

Figure CN120723914A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of social media text data processing, and specifically relates to a social media text denoising method based on a large model, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of social media, user-generated text content is increasing. This text content covers various industries and contains multiple attributes. This data has different values for different business analyses. For example, second-hand transfers and marketing recommendations may be valid data in marketing analysis, but may be considered invalid data in customer experience analysis.
[0003] Existing social media data processing technologies often adopt a one-size-fits-all approach and are unable to effectively distinguish and process data based on the needs of different business scenarios. This processing method leads to the loss of valid data and the proliferation of invalid data, reducing data utilization efficiency and the accuracy of analysis results. In addition, when processing social media data, existing text classification technologies often ignore the multi-dimensional characteristics of text content. For example, the popular soft-text advertising, in essence, contains two attributes: user experience or product reviews plus advertising. The existence of this situation can easily lead to inaccurate classification results and fail to meet refined business needs. Summary of the Invention
[0004] In order to address the deficiencies in the prior art, the present invention provides a social media text denoising method based on a large model, an electronic device, and a computer-readable storage medium that are flexible, adaptable, accurate, and low-cost.
[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions: The social media text denoising method based on a large model includes the following steps: S1. Collect social media text data as initial social media data, and then formulate atomic category labels for the initial social media data; S2: Use several large models to perform multi-label classification and annotation on the initial social media data in S1, and then use a voting mechanism to obtain the annotation results; S3. After manual proofreading, the proofread data is used to perform LoRA fine-tuning on the large model, enabling it to accurately identify and classify social media text content. The labeled data results are stored in the vector database to obtain the domain large model and the initial vector database. S4. The social media texts that need to be denoised are passed through a cleaning pipeline to filter out data that does not meet the requirements; S5: Perform similarity retrieval on the social media text and the data in the initial vector database. If the retrieval is successful, the label category is directly output. If the retrieval fails, the process proceeds to S6. S6. Perform multi-label classification on social media texts that failed to be retrieved using the domain model, and then output the classification results.
[0006] Preferably, the step S6 further includes storing the classification result in an initial vector database.
[0007] Preferably, the atomized category label in S1 is a category label of the smallest unit that cannot be further subdivided in the social media data.
[0008] Preferably, in S4, the social media text to be denoised is passed through a cleaning pipeline to remove one or more of emoticons, HTML, and irregular character content.
[0009] The present invention also discloses an electronic device, comprising: one or more processors; as well as A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are enabled to perform the above-mentioned large model-based social media text denoising method.
[0010] The present invention also discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the above-mentioned large model-based social media text denoising method.
[0011] By adopting the above technical solution, the present invention has the following beneficial effects: (1) The present invention adopts atomic category labels and a multi-label classification scheme, which can flexibly respond to data denoising requirements in different business scenarios and improve the flexibility and adaptability of data processing, especially social media text data; (2) The present invention uses a large model to replace the traditional text classification model and adopts the LoRA fine-tuning scheme, which enables the present invention to more accurately identify and classify social media text content and reduce the problems of category overlap and confusion; (3) The present invention significantly reduces costs while maintaining data processing efficiency through a comprehensive text cleaning and strategy filtering method, as well as a similarity determination technology based on vector retrieval; In summary, the present invention has the advantages of high flexibility and adaptability, high accuracy and low cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a partial flow chart of the large model-based social media text denoising method of the present invention; Figure 2 This is another partial flow chart of the large model-based social media text denoising method of the present invention. DETAILED DESCRIPTION
[0013] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0014] The components of the embodiments of the present invention generally described and shown in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the invention.
[0015] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.
[0016] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0017] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0018] Example 1 In this embodiment, the present invention proposes a social media text denoising method based on a large model, which aims to solve the problem that existing social media data processing technologies are unable to effectively distinguish and process data in different business scenarios. Specifically, the technical problem to be solved by the present invention is how to use a large model to achieve refined classification of social media text content while keeping the complexity of engineering operation and maintenance controllable, so as to meet the data denoising requirements in different business scenarios.
[0019] like Figure 1 and Figure 2 As shown, in one embodiment of the present invention, the social media text denoising method based on a large model of the present invention includes the following steps. Specifically, the social media text denoising method based on a large model of the present invention includes two parts, namely the system end part and the service end part, wherein S1-S3 and Figure 1 For the system side, S4-S6 and Figure 2 For the server part: S1. Collect social media text data as initial social media data, and then formulate atomic category labels for the initial social media data. Atomized category labels are category labels for the smallest unit in social media data that cannot be further subdivided. The smallest unit specifically refers to the smallest category commonly found in social media text, such as product reviews, strategy records, user-generated content (UGC), demonstrations and tutorials, and emotional expressions. In S1, during the initialization phase of the model, a batch of social media text data is collected and corresponding atomic category labels are formulated. More specifically, during the initialization phase of the technical system or model, before the system is started or the model is trained, a certain amount of text content is first collected from social media platforms (such as Weibo, Douyin, etc.) as the initial data reserve. This operation usually serves the purposes of data preprocessing, model training, and functional testing. The initialization phase refers to the preparation phase before the formal operation of the system or model (such as development testing, pre-launch deployment, and initial startup). At this time, basic data needs to be prepared in advance to ensure the normal operation of subsequent functions. Social media text data comes from user-generated text content such as Weibo blog posts, Douyin comments, forum posts, user private messages, social dynamics, etc. Its specific forms include plain text, text with emoticons, topic tags (#), @ mentions, links, and other unstructured data; Collecting social media text data during initialization is essentially a "warm-up" for the system or model (mainly the large model in this invention) - by storing representative initial data, it can quickly respond to user needs (such as recommendations, analysis, and interaction) after startup. This operation is common in natural language processing, public opinion analysis, and social products. Its core lies in balancing data scale, compliance, and matching with application scenarios, laying the foundation for subsequent functional iterations.
[0020] S2. Use several large models to perform multi-label classification and annotation on the initial social media data in S1. According to the characteristics of social media text content, a multi-label classification scheme is adopted to allow a social media text to simultaneously meet the characteristics of multiple types of text, effectively solving the category overlap confusion problem caused by single-label classification. Then, a voting mechanism is used to obtain the annotation results. The specific large models are all existing open source large models. In the specific multi-label classification and annotation scenario of initial social media data in the present invention, using several large models to perform multi-label classification and annotation on the initial social media data in S1 is a method of improving annotation accuracy through model integration. Its core logic is to use multiple large models to independently annotate the same batch of data, and then use a voting mechanism to integrate the judgments of each model to finally obtain more reliable annotation results. Specifically, multi-label classification refers to tasks where each data sample can belong to multiple categories (labels) at the same time. For example, a social media text may contain multiple labels such as "technology", "entertainment", and "hot events". Integration of multiple models refers to calling multiple different large language models (such as GPT-4, Claude, LLaMA, etc.) or different versions of the same model to label the same batch of data. The voting mechanism refers to integrating the labeling results of multiple models and determining the final label in a "minority obeys majority" manner. It is specifically divided into "hard voting" and "soft voting". Hard voting refers to directly counting the number of occurrences of the labels labeled by each model. If more than half of them appear, they will be adopted. Soft voting refers to weighted calculation based on the confidence level of the model (such as probability value). Models with higher confidence have higher voting weights. A more specific operation method is to first select a large model and specific data, and in the present invention, refer to Figure 1, you can select three large models and the initial social media data in S1, and then input the initial social media data in S1 (such as user comments and posts) into each large model in batches, requiring the large model to output the label set corresponding to each text. In the specific labeling stage, each large model makes independent predictions, which can improve the labeling accuracy and avoid the defects of a single model. A single large model may make misjudgments (such as missing key labels) due to training data deviation or algorithm limitations. The integration of multiple models can reduce the error rate through "complementarity" and reduce the cost of manual labeling. For massive data (such as millions of texts), manual labeling is time-consuming and labor-intensive, while the integration of large models can automatically generate labeling results, requiring only a small amount of manual verification (such as spot checking 10% of the data to correct deviations). It can also adapt to complex scenarios, such as in multi-label scenarios. Social media texts often involve multi-dimensional labels (such as "emotion + field + event type"), and a single model may miss labels. A voting mechanism can improve coverage. For example, in cross-domain scenarios, it can simultaneously process technology, entertainment, and livelihood texts. Different models perform better in their respective fields, and the overall effect is more stable after integration. The essence of multi-model voting annotation is to improve the reliability of data annotation through "collective wisdom". It is particularly suitable for scenarios such as social media with complex content and diverse labels. Its core value lies in replacing some manual labor with algorithm integration, while reducing the uncertainty of a single model through model complementarity. In practical applications, it is necessary to balance model diversity, voting strategy, and business costs. The ultimate goal is to generate high-quality, reliable labeled data to lay the foundation for subsequent model training or system analysis. S3. After manual proofreading, the proofread data is used to perform LoRA fine-tuning on the large model, so that the large model can accurately identify and classify social media text content, and the labeled data results are stored in the vector database to obtain the domain large model and the initial vector database. Specifically, based on the general categories of business scenarios, the open source large model is fine-tuned with LoRA (Low-Rank Adaptation) to adapt to specific business needs and improve the adaptability and accuracy of the open source large model. Based on the set minimum unit, the large model is used for pre-labeling, and then manually proofread to obtain a standard data set (training / testing). By comparing the results of the LoRA models trained under different LoRA ranks and alphas, the model with the best performance on the test set is finally selected, which is the domain large model; In the scenario of social media text processing, the process combining manual proofreading, LoRA fine-tuning, and vector database storage is a complete technical solution for model optimization and efficient retrieval. Data annotation and proofreading first generate initial annotations through multi-model voting and then correct the deviations through manual verification to form a high-quality annotated dataset. Model optimization uses the proofread data to perform LoRA fine-tuning on the open-source large model to improve its recognition and classification capabilities for social media text. Data storage and application convert the annotated data into vectors and store them in a vector database to support subsequent text retrieval based on similarity. Among them, manual proofreading is a key link to improve data quality. The purpose of proofreading is to correct possible errors in multi-model voting (such as label omission, redundancy, classification errors) to ensure the accuracy and consistency of the annotated data. The key points of proofreading include label integrity, semantic accuracy, and label system unity. Label integrity is to check whether key dimensions are omitted (such as sentiment tendency, domain classification, event type). Semantic accuracy is to correct the misjudgment of the model on Internet buzzwords and ambiguous sentences (such as whether "amazing" is correctly recognized as a positive evaluation). Label system unity is to ensure that the label expressions of the same semantics are consistent (such as "mobile phone" is unified as "electronic product" instead of "digital product"); In this invention, LoRA fine-tuning makes the model more suitable for the social media scenario. In the initial stage, fine-tuning data preparation is carried out first, including data format preparation and data division. Data format preparation refers to organizing the proofread text and labels into "text-label pairs". Data division is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1 to ensure the generalization ability of the model. The key point of LoRA fine-tuning target layer selection is to insert LoRA parameters into the attention layer and MLP layer of the model to avoid modifying the underlying semantic representation module. The performance of the model on the test set is evaluated through indicators such as accuracy (Accuracy) and F1 score to verify the fine-tuning effect; This invention constructs an initial vector database to improve the retrieval ability. The process of text vectorization uses the fine-tuned large model or a dedicated vectorization model (such as SBERT) to convert the text into high-dimensional vectors (such as 768 dimensions). The vector dimension needs to match the index structure of the vector database. Through the technical combination of "manual proofreading to improve data quality, LoRA fine-tuning to endow the model with scenario capabilities, and then the vector database to achieve efficient retrieval", this invention constructs a complete link from data processing to application implementation. Its core value lies in: using the generalization ability of the large model to reduce the annotation cost and realizing the secondary utilization of data through vector retrieval, finally forming a virtuous cycle of "model optimization - data accumulation - application efficiency improvement", which is especially suitable for scenarios such as social media where data is updated quickly and has high semantic diversity. Aiming at the problem of high cost of the large model solution, this invention subsequently continues to apply text cleaning and strategy filtering methods, as well as similarity determination techniques based on vector retrieval to reduce costs while maintaining the efficiency of data processing; The above is the processing flow on the system side. The following will expand the processing flow on the server side. S4. Run the social media text that needs to be denoised through a cleaning pipeline. This involves removing one or more of the following: emojis, HTML, irregular characters, and filtering out data that doesn't meet the requirements. A "cleaning pipeline" is a process that pre-processes raw text using a series of rules and algorithms to remove noise and filter out valid data. Social media text is user-generated content from social media platforms (such as Weibo posts, TikTok comments, and forum posts). It is characterized by diverse formats, unstructured content, and a high level of noise. A cleaning pipeline is a standardized data pre-processing process that gradually purifies the text through multiple stages, laying the foundation for subsequent analysis (such as classification, retrieval, and modeling). Specific operations in the cleaning pipeline include removing non-text elements, normalizing character processing, and filtering and screening data. For example, for emojis, Unicode emojis are directly deleted or converted into text descriptions (such as "smile" and "heart") to prevent non-semantic symbols from interfering with model understanding. For HTML tags, web page tags can be directly filtered. For special symbols, @mentions (@users) are removed to reduce irrelevant information. Normalizing character processing refers to correcting irregular characters, deduplicating, and replacing special characters, while also normalizing the language. Data filtering and screening involves length filtering, content compliance filtering, and semantic validity filtering. The essence of the media text cleaning pipeline is to "remove noise and retain semantics." Through standardization, unstructured user-generated content is converted into clean data suitable for machine analysis. This process is similar to the purification of raw materials in industrial production and is the foundation of subsequent AI applications (such as large model fine-tuning, vector retrieval, and public opinion analysis). In practice, the cleaning intensity must be balanced according to business objectives, removing interference information while retaining key semantic features, ultimately improving the efficiency and accuracy of the entire data processing chain. S5: Perform similarity retrieval on the social media text and the data in the initial vector database. If the retrieval is successful, the label category is directly output. If the retrieval fails, the process proceeds to S6. S6. Perform multi-label classification on the social media texts that failed to be retrieved using the domain big model, and then output the classification results. At the same time, the classification results are stored in the initial vector database. Specifically, since the social media texts that failed to be retrieved are not in the vector database, they need to be stored in the vector database while performing multi-label classification on the social media texts that failed to be retrieved using the domain big model.
[0021] The present invention is particularly applicable to the following fields: (1) Social media data analysis and management Application: The large-scale model-based social media text denoising method of the present invention can be used to classify and clean text content on social media, effectively distinguishing valid data from invalid data in different business scenarios, and improving data utilization efficiency and the accuracy of analysis results.
[0022] (2) Sentiment Analysis and Public Opinion Monitoring Application: Preprocessing social media data to improve the accuracy of sentiment analysis models for public opinion monitoring and market research.
[0023] (3) Customer experience analysis Application: In the field of customer service, this invention can be used to filter out irrelevant social media data, such as advertising and marketing information, and focus on users' real feedback and experience sharing.
[0024] The present invention also discloses an electronic device, comprising: one or more processors; as well as A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are enabled to perform the above-mentioned large model-based social media text denoising method.
[0025] The present invention also discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the above-mentioned large model-based social media text denoising method.
[0026] This embodiment does not impose any formal restrictions on the shape, material, structure, etc. of the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are within the scope of protection of the technical solution of the present invention.
Claims
1. A social media text denoising method based on a large model, characterized by: The following steps are involved: S1. Collect social media text data as initial social media data, and then formulate atomic category labels for the initial social media data; S2: Use several large models to perform multi-label classification and annotation on the initial social media data in S1, and then use a voting mechanism to obtain the annotation results; S3. After manual proofreading, the proofread data is used to perform LoRA fine-tuning on the large model, enabling it to accurately identify and classify social media text content. The labeled data results are stored in the vector database to obtain the domain large model and the initial vector database. S4. The social media texts that need to be denoised are passed through a cleaning pipeline to filter out data that does not meet the requirements; S5: Perform similarity retrieval on the social media text and the data in the initial vector database. If the retrieval is successful, the label category is directly output. If the retrieval fails, the process proceeds to S6. S6. Perform multi-label classification on social media texts that failed to be retrieved using the domain model, and then output the classification results.
2. The social media text denoising method based on a large model according to claim 1, characterized in that: The step S6 also includes storing the classification result in an initial vector database.
3. The social media text denoising method based on a large model according to claim 1, characterized in that: The atomized category labels in S1 are category labels of the smallest units in social media data that cannot be further subdivided.
4. The social media text denoising method based on a large model according to claim 1, characterized in that: In S4, the social media text that needs to be denoised is passed through a cleaning pipeline to remove one or more of emoticons, HTML, and irregular character content.
5. An electronic device, characterized in that include: one or more processors; as well as A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to perform the large model-based social media text denoising method according to any one of claims 1 to 4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the social media text denoising method based on a large model according to any one of claims 1 to 4 is implemented.
Citation Information
Cited By
User internet log processing method and device based on large model, and computer program product
CN121743300A