Natural Language Request Categorization via Word Vector Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for processing insurance claims, such as those described in WO 01/13295 and US 2014/0058763, are inadequate for handling regular requests and are overly complex, lacking efficient means for determining the category of a request or requiring extensive human intervention.
Innovation Solution
A computer-implemented method and system that allows users to input requests as natural language text strings, utilizing natural language processing and machine learning to quickly determine the category of an insurance claim, with the option for human operator intervention when necessary, and enabling submission of requests as images using OCR and pattern detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing methods for processing insurance claims are used, then fraud detection capability is improved, but device complexity and processing time increase significantly
Solution Approach 1:
The system segments claim processing into distinct categories (e.g., property damage, personal injury, liability) and processes each category using specialized algorithms and knowledge bases. This segmentation allows the system to handle different types of claims efficiently without requiring a single complex processing path for all claims.
Solution Approach 2:
The system performs preliminary categorization of claims based on key indicators and patterns before full processing. By pre-classifying claims into categories and identifying obvious cases that can be handled automatically, the system reduces the complexity of subsequent processing steps and minimizes human intervention requirements.
2Reliability
If existing fraud detection methods are applied to regular claims, then potential fraud can be identified, but processing speed and efficiency decrease
Solution Approach 1:
The system applies different processing qualities and levels of scrutiny to different claims based on their category and risk indicators. Routine claims with clear categories receive streamlined processing, while claims with ambiguous categories or high-risk indicators receive more thorough analysis. This local differentiation maintains processing speed for routine claims while ensuring thorough fraud detection when needed.
Solution Approach 2:
The system applies fraud detection algorithms selectively rather than uniformly to all claims. By using automated categorization to identify claims that require detailed fraud analysis versus those that can be processed routinely, the system achieves adequate fraud detection for high-risk claims while maintaining high processing speed for low-risk claims.
3Productivity
If automated processing is implemented for all claims, then productivity increases, but measurement precision and decision accuracy may worsen
Solution Approach 1:
The system incorporates feedback mechanisms where automated categorization decisions are continuously evaluated and refined. When automated processing encounters ambiguous cases or potential errors, the system can flag these for human review, and the outcomes feed back into improving the automated categorization algorithms. This feedback loop maintains high productivity while progressively improving measurement precision through learning from actual cases.
Solution Approach 2:
The system introduces an intermediary human review layer for ambiguous or high-stakes claims. Rather than relying solely on automated processing, the system acts as an intermediary that handles clear-cut cases automatically while routing uncertain cases to human experts. This hybrid approach maintains high overall productivity while ensuring high precision for difficult categorization cases.
Data Source
AI summary
A computer-implemented method determines a category of a request provided by a user by means of a user device. The user device includes connection means and means for receiving a request description relating to said request from said user. The method includes receiving, from the user, the request description, by means of the device, and uploading the request description to a server. The server has access to a database which includes a number of previously categorized requests each including a category and a vocabulary, which includes a number of word vector representations. The method further includes identifying, by the server, a number of component words belonging to a natural language text string included in the request description; obtaining, for at least one of the component words, an associated word vector representation from the vocabulary, and determining a request vector, based on at least one obtained word vector representation.


