User Population Formation Using Keyword Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in maintaining a constant attribute ratio in user populations over time, particularly when users change attributes, such as from students to employed persons, making it difficult to conduct precise market trend surveys and product development using Web-based user information.
Innovation Solution
A method involving a data collection apparatus that extracts keywords from blog posts to form and replenish user groups based on attribute values, using duplicate keywords to determine attribute values and maintain a consistent attribute ratio by selecting new users with similar posting trends, thereby ensuring accurate representation of user populations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If users are continuously added to the population without considering attribute changes, then the population size increases, but the attribute ratio becomes inaccurate and changes over time
Solution Approach 1:
The system continuously monitors attribute changes in users by extracting keywords from their blog posts and compares current attribute distributions against target ratios. When deviations are detected, the system automatically adjusts the population by adding or removing users to maintain accurate attribute representation, creating a closed-loop feedback mechanism that resolves the contradiction between population growth and attribute accuracy.
Solution Approach 2:
The system dynamically changes the selection criteria for adding users based on current attribute distributions. By calculating the difference between actual and target attribute ratios and adjusting the weights of different attributes in the selection process, the system maintains accurate attribute representation while allowing population growth, effectively resolving the contradiction through adaptive parameter adjustment.
2Reliability
If duplicate keywords are used to determine attribute values, then attribute determination becomes more reliable, but the complexity of keyword processing increases
Solution Approach 1:
The system performs preliminary processing by pre-defining keyword sets associated with specific attributes and pre-calculating the relationships between keywords and attributes. This preliminary preparation allows the system to reliably determine user attributes using duplicate keywords without requiring complex real-time processing, thus improving reliability while managing complexity through advance preparation.
Solution Approach 2:
The system introduces an intermediary layer of keyword normalization and mapping that translates various forms of keyword expressions into standardized attribute representations. This intermediary processing layer simplifies the relationship between raw keyword data and attribute determinations, making the system more reliable while reducing the apparent complexity of keyword processing through abstraction.
3Measurement precision
If the system processes all blog posts to extract keywords, then the accuracy of attribute extraction improves, but the processing time and computational resources increase
Solution Approach 1:
The system extracts only the most relevant keywords from blog posts by using pre-defined keyword sets and filtering mechanisms. Instead of processing all text content equally, the system selectively extracts keywords that match predefined attribute categories, significantly reducing processing time while maintaining high accuracy in attribute extraction through targeted extraction rather than comprehensive analysis.
Solution Approach 2:
The system applies partial processing by focusing on the most influential keywords and attributes rather than uniformly processing all content. By identifying and processing only the critical keyword-attribute relationships needed for accurate population representation, the system achieves sufficient accuracy without the computational burden of exhaustive processing, effectively balancing precision and time efficiency.
Data Source
AI summary
A population formation method is disclosed. Keywords are extracted from public information of providers included as elements in a first provider group. Each element is calculated based on a predetermined attribute value. A first attribute is for the providers of the public information. The attribute value is changed with time. Each of rules set for duplicate keywords is to determine one of the attributes by using one of the duplicate keywords. Provider groups are formed for new public information based on the duplicate keywords and the rules. A provider group having a similar relationship with a first provider group is specified by a distribution of the attribute value of a different attribute from the first attribute. A new provider group corresponding to the first provider group is formed by the providers, for whom the attribute value of the first attribute corresponds to the predetermined attribute value.


