Organizational Language Model Generation via Cloud Data Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generic speech recognition models trained on web data struggle to recognize organization-specific terminology, leading to poor performance in identifying unique terms and phrases used within a particular organization.

Innovation Solution

An organizational-specific language model is generated using company-specific data from cloud-based applications, filtering out private data and aggregating public data to create a tailored model for improved speech recognition within the organization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a generic language model is trained using web data, then the model covers a wide array of generalized speech topics, but the model fails to recognize organization-specific terminology and unique terms

Engineering Contradiction:
Improvecoverage of generalized speech topicsVSAvoidrecognition accuracy of organization-specific terminology
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The language model is segmented into two components: a generic language model trained on web data for general speech coverage, and an organizational language model trained on organization-specific data for domain-specific terminology. This segmentation allows each model to specialize in its respective domain, resolving the contradiction between general coverage and specific accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent combines the generic language model and the organizational language model into a unified speech recognition system. The models work together, with the generic model handling general speech and the organizational model handling domain-specific terms, thereby achieving both broad coverage and high precision for organization-specific terminology.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If organization-specific data is collected and used for training, then the language model recognizes unique organizational terms, but the model requires specialized training data and increased complexity

Engineering Contradiction:
Improverecognition accuracy of organization-specific terminologyVSAvoidcomplexity of data collection and model training system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system automatically collects organization-specific data from cloud-based applications and platforms without requiring manual data gathering. The language model is automatically trained on this collected data, reducing the need for manual intervention and simplifying the overall process while maintaining high recognition accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses cloud-based applications and platforms as intermediaries to collect and provide organization-specific data. These existing platforms serve as data sources, eliminating the need for separate complex data collection systems and reducing overall system complexity while enabling accurate organizational term recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Object-affected harmful factors

If private organizational data is filtered out, then data security and privacy are maintained, but the amount of available training data is reduced

Engineering Contradiction:
Improvedata security and privacy protectionVSAvoidvolume of training data
Core Design Contradiction:
Object-affected harmful factorsVSQuantity of substance

Solution Approach 1:

The system extracts only the necessary public and semi-public data from cloud-based applications for training the language model, while leaving private and confidential data behind. This selective extraction maintains data security and privacy while providing sufficient training data for accurate organizational term recognition.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different data access policies to different types of data: public data is freely used for training, semi-public data is used with appropriate permissions, and private data is excluded. This local quality approach ensures that each data type is handled according to its sensitivity level, maintaining security while maximizing available training data.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11676576B2Organizational-based language model generation
Publication Date: 2023.06.13 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11676576B2 patent drawing
  • US11676576B2 patent drawing
  • US11676576B2 patent drawing

AI summary

Systems and methods are provided for acquiring training data and building an organizational-based language model based on the training data. In organizational data is generated via one or more applications associated with an organization, the collected organizational data is aggregated and filtered into training data that is used for training an organizational-based language model for speech processing based on the training data.