Server Voice Recognition Model Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition systems face performance deterioration when a single voice recognition model is updated for all users simultaneously, as it fails to account for individual differences in voice patterns based on gender, age, and region, leading to decreased accuracy for specific users.
Innovation Solution
A server-based system that provides customized voice recognition models to each device by acquiring use-related information, including user feedback, meta information, and recognition performance, and dynamically updates the model to optimize performance for individual users based on their characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single voice recognition model is updated for all users simultaneously, then the overall voice recognition service can be maintained with a unified model, but the recognition accuracy deteriorates for specific users with unique voice patterns
Solution Approach 1:
The patent segments the unified voice recognition model into multiple customized models based on user characteristics such as gender, age, and region. Each user group receives a dedicated model trained on their specific voice patterns, thereby maintaining high recognition accuracy while managing complexity through systematic segmentation of the user base and corresponding model distribution.
Solution Approach 2:
The patent applies local quality by providing different voice recognition models tailored to specific user groups rather than a uniform model for all. Each localized model is optimized for its target demographic's voice characteristics, ensuring high accuracy for each segment while the server manages multiple specialized models instead of one generic model.
2Measurement precision
If customized voice recognition models are provided for each user, then recognition accuracy is improved, but the system complexity increases
Solution Approach 1:
The patent implements universality through a centralized server that performs multiple functions: storing voice data, training customized models, managing user profiles, and distributing appropriate models to devices. This multi-functional server handles the complexity centrally, allowing individual user devices to benefit from customized models without each device needing to possess full model training and management capabilities.
Solution Approach 2:
The patent introduces a server as an intermediary between voice data collection and model deployment. The server acts as a mediator that processes raw voice data, trains customized models, and delivers them to appropriate user devices. This intermediary absorbs the computational and organizational complexity, enabling accurate customized recognition without requiring complex infrastructure at each user endpoint.
3Reliability
If voice data from all users is collected and processed, then the voice recognition model can be continuously improved, but the time required for model updates increases
Solution Approach 1:
The patent applies preliminary action by continuously collecting and preprocessing voice data in the background, and proactively training updated customized models before users experience performance degradation. The server maintains a pipeline of data collection, model training, and validation that prepares improved models in advance, reducing the perceived update time for users while continuously improving model performance through accumulated voice data.
Data Source
AI summary
A server may provide a voice recognition service. The server may include a memory configured for storing a plurality of voice recognition models, a communication device configured for communicating a plurality of voice recognition devices, and an artificial intelligence device configured for providing a voice recognition service to the plurality of voice recognition devices, acquiring use-related information regarding a first voice recognition device (from among the plurality of voice recognition devices), and changing a voice recognition model corresponding to the first voice recognition device from a first voice recognition model to a second voice recognition model based on the use-related information.


