Social Network Skill Standardization via Collaborative Taxonomy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Social networking systems face challenges in standardizing user-submitted skills due to varying descriptions, leading to inconsistencies in skill identification and classification.
Innovation Solution
A method and system for extracting and standardizing skills from social networking profiles by extracting seed phrases, disambiguating their meanings, and grouping similar skills together through a collaborative taxonomy process, utilizing machine learning and user voting to establish standardized entities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If user-submitted skills are stored as-is without standardization, then data diversity and user freedom are preserved, but data consistency and classification accuracy deteriorate
Solution Approach 1:
The patent introduces an intermediary standardization system that sits between user skill submission and storage. This system includes components for extracting seed phrases, determining entity types, resolving ambiguities, and grouping similar skills. The intermediary process transforms diverse user inputs into standardized classifications while preserving the original data for reference, thus maintaining both data diversity and classification accuracy.
Solution Approach 2:
The standardization process is segmented into distinct modular steps: extraction of seed phrases from user profiles, determination of entity types (skill, tool, technology, etc.), ambiguity resolution through multiple data sources, and grouping of similar entities. This segmentation allows each component to handle specific aspects of the standardization challenge independently, maintaining flexibility while achieving consistency.
2Manufacturing precision
If a comprehensive standardization process is implemented, then data consistency is improved, but processing time and system complexity increase
Solution Approach 1:
The system performs preliminary actions by extracting and storing seed phrases from user profiles before the actual standardization process. These seed phrases serve as pre-processed input that reduces the complexity of subsequent entity type determination and grouping operations. The preliminary extraction organizes raw data into manageable units that can be efficiently processed through the standardization pipeline.
Solution Approach 2:
The system incorporates feedback mechanisms where the standardization results are used to refine and improve the standardization process itself. The grouping of similar skills and entities creates a growing knowledge base that feeds back into the entity type determination and ambiguity resolution stages, reducing system complexity over time as the standardized knowledge base expands.
3Measurement precision
If multiple data sources are queried for ambiguity resolution, then classification accuracy is improved, but processing time increases
Solution Approach 1:
The system applies local quality by querying different data sources with different levels of detail and reliability based on the specific context of each skill entity. Not all entities require the same depth of ambiguity resolution - common skills with clear meanings receive minimal processing while ambiguous or niche skills receive more extensive verification across multiple data sources. This localized approach to quality control reduces overall processing time while maintaining high classification accuracy where needed.
Data Source
AI summary
A system extracts data from profiles on a social networking system. The system writes the data to a database when the data exceeds a first threshold. The system then determines a degree of similarity between the data and other similar data, and writes the data and a first portion of the other similar data to the database when the degree of similarity between the data and the first portion of the other similar data exceeds a second threshold. The system then receives into the computer processor input from a plurality of users. The input relates to an agreement or disagreement regarding the degree of similarity between the data and the first portion of the other similar data. The system writes the data and a second portion of the other similar data to the database as a function of the agreement or disagreement of the plurality of users.


