Personalizing Speech-Controlled Services via Language Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for personalizing telecommunication services are inadequate for users of speech-controlled devices, as they rely primarily on written and read data, failing to effectively utilize voice interactions and are not well-suited for a variety of services, including those provided by external providers.
Innovation Solution
A method and system that generates user-dependent language models through speech recognition, stores these models, and makes them available to both local devices and external service providers, enabling personalized services on speech-controlled devices and creating a multimodal interaction space that adapts to user preferences and habits.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If user profiles are built from questionnaires and written behavior data, then information about user interests and preferences can be collected, but the process becomes cumbersome for users and fails to capture voice interactions
Solution Approach 1:
The patent replaces mechanical/text-based input methods (questionnaires, typing) with acoustic/voice-based interaction. The speech recognition system captures user interactions through voice commands and conversations, automatically generating language models that reflect user interests and preferences without requiring manual data entry or filling out forms.
Solution Approach 2:
The system automatically collects and processes user interaction data through speech recognition without requiring explicit user input about their preferences. The speech recognition system continuously learns from natural conversations and voice commands to build accurate user profiles and language models autonomously.
2Adaptability or versatility
If speech recognition systems generate user-dependent language models, then personalization for speech-controlled devices is improved, but the system complexity increases
Solution Approach 1:
The speech recognition system performs multiple functions: it recognizes speech commands, generates language models, creates user profiles, and provides personalized service recommendations. By making the speech recognition system multi-functional, the patent avoids the need for separate systems for each function, thereby managing complexity while achieving versatile personalization.
Solution Approach 2:
The language model acts as an intermediary between the speech recognition system and the service personalization layer. It translates raw speech data into structured user profile information that can be used by various services, decoupling the complexity of speech processing from the complexity of service personalization logic.
3Productivity
If user profiles are made available to external service providers, then commercial efficiency is improved, but data privacy and security concerns arise
Solution Approach 1:
The patent applies different data sharing strategies to different types of information. Sensitive personal data remains localized with the user, while aggregated, anonymized language models and service recommendations are shared with external providers. This selective sharing approach maintains privacy while enabling commercial utilization of user data.
Solution Approach 2:
Instead of sharing raw user data, the system creates and shares copies in the form of language models and derived profiles. These copies contain the necessary information for service personalization but obscure the original user data, reducing privacy risks while maintaining analytical value for service providers.
Data Source
AI summary
Method for building a multimodal business channel between users, service providers and network operators. The service provided to the users is personalized with a user's profile derived from language and speech models delivered by a speech recognition system. The language and speech models are synchronized with user dependent language models stored in a central platform made accessible to various value added service providers. They may also be copied into various devices of the user. Natural language processing algorithms may be used for extracting topics from user's dialogues.


