Organizational Language Model Generation via Cloud Data Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generic speech recognition models trained on web data struggle to recognize organization-specific terminology, leading to poor performance in identifying unique terms and phrases used within a particular organization.
Innovation Solution
An organizational-specific language model is generated using company-specific data from cloud-based applications, filtering out private data and aggregating public data to create a tailored model for improved speech recognition within the organization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a generic language model is trained using web data, then the model covers a wide array of generalized speech topics, but the model fails to recognize organization-specific terminology and unique terms
Solution Approach 1:
The language model is segmented into two components: a generic language model trained on web data for general speech coverage, and an organizational language model trained on organization-specific data for domain-specific terminology. This segmentation allows each model to specialize in its respective domain, resolving the contradiction between general coverage and specific accuracy.
Solution Approach 2:
The patent combines the generic language model and the organizational language model into a unified speech recognition system. The models work together, with the generic model handling general speech and the organizational model handling domain-specific terms, thereby achieving both broad coverage and high precision for organization-specific terminology.
2Measurement precision
If organization-specific data is collected and used for training, then the language model recognizes unique organizational terms, but the model requires specialized training data and increased complexity
Solution Approach 1:
The system automatically collects organization-specific data from cloud-based applications and platforms without requiring manual data gathering. The language model is automatically trained on this collected data, reducing the need for manual intervention and simplifying the overall process while maintaining high recognition accuracy.
Solution Approach 2:
The patent uses cloud-based applications and platforms as intermediaries to collect and provide organization-specific data. These existing platforms serve as data sources, eliminating the need for separate complex data collection systems and reducing overall system complexity while enabling accurate organizational term recognition.
3Object-affected harmful factors
If private organizational data is filtered out, then data security and privacy are maintained, but the amount of available training data is reduced
Solution Approach 1:
The system extracts only the necessary public and semi-public data from cloud-based applications for training the language model, while leaving private and confidential data behind. This selective extraction maintains data security and privacy while providing sufficient training data for accurate organizational term recognition.
Solution Approach 2:
The patent applies different data access policies to different types of data: public data is freely used for training, semi-public data is used with appropriate permissions, and private data is excluded. This local quality approach ensures that each data type is handled according to its sensitivity level, maintaining security while maximizing available training data.
Data Source
AI summary
Systems and methods are provided for acquiring training data and building an organizational-based language model based on the training data. In organizational data is generated via one or more applications associated with an organization, the collected organizational data is aggregated and filtered into training data that is used for training an organizational-based language model for speech processing based on the training data.


