Message Digest Generation Using Label Distribution Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating message digests from social media messages, which are characterized by short text, loud noise, and informal language, fail to produce accurate results due to their reliance on content-based multi-article summarization techniques.
Innovation Solution
A method and apparatus that generate a message digest by creating function label, sentiment label, word category label, and word sentiment polarity distribution models to determine the probability of a word being a subject content word, allowing for a more accurate extraction of important content from associated messages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If content-based multi-article summarization method is used to extract message digest, then the method can be applied to messages, but the extraction accuracy is poor due to short text, loud noise, and informal language characteristics of social media messages
Solution Approach 1:
The patent transforms the digest extraction problem from traditional text summarization parameters to probabilistic parameters by introducing function label distribution models, sentiment label distribution models, and word category label distribution models. These models compute probabilities for different message types, sentiments, and word categories, fundamentally changing how message importance is evaluated and enabling accurate extraction despite social media's informal characteristics
Solution Approach 2:
The patent introduces multiple intermediary models (function label distribution model, sentiment label distribution model, word category label distribution model) that act as mediators between the raw social media messages and the final digest. These intermediaries process and transform the noisy input messages into structured probabilistic representations, enabling accurate digest extraction by filtering out noise and highlighting important content
2Ease of manufacture
If traditional summarization methods are used, then the processing is simple, but the digest quality and informativeness are insufficient
Solution Approach 1:
The patent segments the digest extraction process into distinct functional modules: function label distribution model for message type classification, sentiment label distribution model for sentiment analysis, word category label distribution model for important word identification, and digest generation module. This segmentation allows each component to specialize in one aspect of analysis, improving overall digest quality while maintaining systematic processing
Solution Approach 2:
The patent creates a multi-functional processing system where the same framework handles multiple tasks: classifying message functions, analyzing sentiments, identifying important words, and generating digests. This universal approach improves digest reliability by comprehensively analyzing messages from multiple dimensions rather than relying on a single simple summarization method
Data Source
AI summary
Embodiments of this application provide a message digest generation method and apparatus, and a storage medium. The generation method is performed by an electronic device, and includes: obtaining a plurality of associated messages from a to-be-processed message set; generating a function label distribution model, a sentiment label distribution model, a word category label distribution model, and a word sentiment polarity label distribution model corresponding to each of the plurality of associated messages; determining, based on the function label distribution model, the sentiment label distribution model, the word category label distribution model, and the word sentiment polarity label distribution model, a distribution probability that a category of a word included in the plurality of associated messages is a subject content word; and generating a digest of the plurality of associated messages according to the distribution probability of the subject content word.


