Expertise Identification via Anonymous Word Frequency Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large organizations face challenges in identifying employees' expertise within their networks without inadvertently disclosing sensitive information, and existing methods lack effective privacy protection and relevance filtering.

Innovation Solution

A method that analyzes data from emails, messages, and electronic communications using natural language processing, compares word frequencies with those of collaborators, and provides a list of relevant phrases while ensuring privacy through anonymous frequency word counts and optional redaction, using a processor and memory-based apparatus or computer-readable medium.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If word frequency analysis is performed on employee communications to identify expertise, then expertise identification accuracy is improved, but employee privacy and sensitive information protection deteriorate

Engineering Contradiction:
Improveexpertise identification accuracyVSAvoidprivacy violation
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the essential expertise-indicative features (word frequencies, topic distributions) from employee communications while leaving out all personally identifiable information and sensitive content. This extraction process separates the useful expertise signals from the privacy-sensitive data, allowing accurate expertise identification without privacy violation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary processing layer that analyzes communications through aggregated statistical measures rather than direct individual content inspection. This intermediary approach uses word frequency counts and topic modeling as mediators between the raw communication data and expertise determination, preventing direct exposure of sensitive information while maintaining identification accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If all words and phrases from employee communications are analyzed, then completeness of expertise identification is improved, but data processing complexity and time consumption worsen

Engineering Contradiction:
Improveexpertise information completenessVSAvoiddata processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts and retains only those words and phrases that are statistically significant indicators of expertise, discarding common stop words and irrelevant terminology. This selective extraction maintains completeness of expertise information while dramatically reducing the volume of data requiring processing.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the raw communication data into standardized parameters such as word frequency counts, topic distribution vectors, and co-occurrence statistics. This parameter transformation converts unstructured text into structured numerical data that is more efficient to process and compare, reducing overall system complexity.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If anonymous frequency word counts are used to protect privacy, then privacy protection is improved, but the ability to trace expertise to specific employees deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoidexpertise traceability
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent segments the identification process into two independent stages: first, aggregate expertise patterns are identified through anonymous word frequency analysis across the organization; second, these patterns are matched against individual employee profiles to trace specific expertise locations. This segmentation allows privacy protection in the analysis phase while enabling traceability in the matching phase.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a new dimension of aggregation to the analysis, working with frequency counts and statistical patterns rather than direct text content. This dimensional shift from content-space to frequency-space maintains traceability through statistical signatures while protecting privacy by removing identifiable context.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Loss of information

If common phrases are included in expertise identification, then recall of all potential skills is improved, but precision of relevant expertise identification deteriorates

Engineering Contradiction:
Improveskill recallVSAvoidexpertise identification precision
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent applies different quality thresholds and filtering criteria to different types of words and phrases. Common industry terminology and standard technical terms are treated differently from unique, specific expertise indicators. This local quality differentiation allows common phrases to contribute to recall while precision is maintained through stricter criteria for identifying definitive expertise markers.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent initially performs analysis on all words and phrases (excessive action) to ensure complete skill recall, then applies progressive filtering and thresholding to remove common phrases that do not provide discriminative value. This partial action approach starts comprehensive and refines to precise, achieving both recall and precision through multi-stage processing.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20200111172A1Method, apparatus, and computer-readable medium for expertise
Publication Date: 2020.04.09 DOTALIGN INC
  • US20200111172A1 patent drawing
  • US20200111172A1 patent drawing
  • US20200111172A1 patent drawing

AI summary

Embodiments of the present disclosure provide a method, apparatus and computer-readable medium for determining expertise. An exemplary method includes analyzing, by a processor, a plurality of data of a user, wherein the data comprises at least words and phrases from at least one of emails, messages, and electronic communications, and receiving, by the processor, a second data, the second data comprising anonymous frequency word counts of a plurality of users. The method further includes determining, by the processor, a correspondence between the analyzed plurality of data and the second data, wherein the correspondence includes words and phrases that are interesting, and providing, by the processor, a list of words and phrases to the user based on the determined correspondence, wherein the list is selectable by the user.