Expertise Identification via Anonymous Word Frequency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large organizations face challenges in identifying employees' expertise within their networks without inadvertently disclosing sensitive information, and existing methods lack effective privacy protection and relevance filtering.
Innovation Solution
A method that analyzes data from emails, messages, and electronic communications using natural language processing, compares word frequencies with those of collaborators, and provides a list of relevant phrases while ensuring privacy through anonymous frequency word counts and optional redaction, using a processor and memory-based apparatus or computer-readable medium.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If word frequency analysis is performed on employee communications to identify expertise, then expertise identification accuracy is improved, but employee privacy and sensitive information protection deteriorate
Solution Approach 1:
The patent extracts only the essential expertise-indicative features (word frequencies, topic distributions) from employee communications while leaving out all personally identifiable information and sensitive content. This extraction process separates the useful expertise signals from the privacy-sensitive data, allowing accurate expertise identification without privacy violation.
Solution Approach 2:
The patent introduces an intermediary processing layer that analyzes communications through aggregated statistical measures rather than direct individual content inspection. This intermediary approach uses word frequency counts and topic modeling as mediators between the raw communication data and expertise determination, preventing direct exposure of sensitive information while maintaining identification accuracy.
2Loss of information
If all words and phrases from employee communications are analyzed, then completeness of expertise identification is improved, but data processing complexity and time consumption worsen
Solution Approach 1:
The patent extracts and retains only those words and phrases that are statistically significant indicators of expertise, discarding common stop words and irrelevant terminology. This selective extraction maintains completeness of expertise information while dramatically reducing the volume of data requiring processing.
Solution Approach 2:
The patent transforms the raw communication data into standardized parameters such as word frequency counts, topic distribution vectors, and co-occurrence statistics. This parameter transformation converts unstructured text into structured numerical data that is more efficient to process and compare, reducing overall system complexity.
3Object-affected harmful factors
If anonymous frequency word counts are used to protect privacy, then privacy protection is improved, but the ability to trace expertise to specific employees deteriorates
Solution Approach 1:
The patent segments the identification process into two independent stages: first, aggregate expertise patterns are identified through anonymous word frequency analysis across the organization; second, these patterns are matched against individual employee profiles to trace specific expertise locations. This segmentation allows privacy protection in the analysis phase while enabling traceability in the matching phase.
Solution Approach 2:
The patent adds a new dimension of aggregation to the analysis, working with frequency counts and statistical patterns rather than direct text content. This dimensional shift from content-space to frequency-space maintains traceability through statistical signatures while protecting privacy by removing identifiable context.
4Loss of information
If common phrases are included in expertise identification, then recall of all potential skills is improved, but precision of relevant expertise identification deteriorates
Solution Approach 1:
The patent applies different quality thresholds and filtering criteria to different types of words and phrases. Common industry terminology and standard technical terms are treated differently from unique, specific expertise indicators. This local quality differentiation allows common phrases to contribute to recall while precision is maintained through stricter criteria for identifying definitive expertise markers.
Solution Approach 2:
The patent initially performs analysis on all words and phrases (excessive action) to ensure complete skill recall, then applies progressive filtering and thresholding to remove common phrases that do not provide discriminative value. This partial action approach starts comprehensive and refines to precise, achieving both recall and precision through multi-stage processing.
Data Source
AI summary
Embodiments of the present disclosure provide a method, apparatus and computer-readable medium for determining expertise. An exemplary method includes analyzing, by a processor, a plurality of data of a user, wherein the data comprises at least words and phrases from at least one of emails, messages, and electronic communications, and receiving, by the processor, a second data, the second data comprising anonymous frequency word counts of a plurality of users. The method further includes determining, by the processor, a correspondence between the analyzed plurality of data and the second data, wherein the correspondence includes words and phrases that are interesting, and providing, by the processor, a list of words and phrases to the user based on the determined correspondence, wherein the list is selectable by the user.


