Automated Empathy Dataset Labeling Using DSM-5 Morpheme Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge is to efficiently label empathy datasets for depressive disorder diagnostic criteria, as existing methods like manual labeling are costly and time-consuming, while semi-supervised learning may result in incorrect labeling or omission.
Innovation Solution
A labeling method that involves receiving consultation dataset sentences, classifying them into morpheme units, extracting target data, and searching for matches in a word dictionary generated based on DSM-5 depressive disorder diagnostic criteria, with the option to label based on similarity if exact matches are not found.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling is used to label empathy datasets, then labeling accuracy is improved, but labeling cost and time consumption increase
Solution Approach 1:
The patent introduces an automatic labeling system as an intermediary between raw consultation data and the final labeled dataset. This system uses natural language processing and machine learning models to automatically extract emotions, intents, and key information from consultation sentences, serving as a mediator that reduces the burden on manual labelers while maintaining acceptable accuracy through subsequent verification steps
Solution Approach 2:
The labeling process is segmented into multiple independent stages: automatic extraction of target data (emotions, intents, keywords), similarity matching against predefined categories, confidence score calculation, and selective manual verification. This segmentation allows the system to handle different types of data with appropriate methods, reducing overall time consumption while maintaining accuracy for critical labels
2Ease of manufacture
If semi-supervised learning is used to reduce labeling costs, then labeling cost decreases, but labeling accuracy deteriorates due to incorrect labeling or omission
Solution Approach 1:
The system implements feedback mechanisms where automatically labeled data is evaluated against confidence thresholds, and low-confidence labels are flagged for manual review. The model continuously learns from corrected labels, improving its performance over time. This feedback loop ensures that cost reductions from automated labeling do not compromise overall labeling accuracy
Solution Approach 2:
Instead of attempting to automatically label all data points, the system applies automatic labeling selectively to high-confidence cases while using manual labeling for ambiguous or critical instances. This partial automation approach balances cost efficiency with accuracy requirements, avoiding the pitfalls of fully automated semi-supervised learning
3Measurement precision
If a comprehensive word dictionary based on DSM-5 criteria is created, then diagnostic accuracy is improved, but system complexity increases
Solution Approach 1:
The comprehensive DSM-5 diagnostic criteria are segmented into discrete, searchable categories and keywords organized in a structured word dictionary. Each diagnostic criterion is broken down into specific emotional states, behavioral patterns, and key phrases that can be independently identified and labeled, making the complex diagnostic framework manageable and implementable in the automated labeling system
Solution Approach 2:
The word dictionary and categorization system are designed to be universal and multi-functional, serving multiple purposes: automatic label generation, quality control verification, model training data creation, and consistent diagnostic criteria application. This multi-functionality reduces the need for separate systems for each task, thereby managing complexity while maintaining comprehensive diagnostic coverage
Data Source
AI summary
A labeling method according to an embodiment may include: receiving a consultation dataset sentence; classifying the consultation dataset sentence into morpheme units and extracting target data; searching whether the target data is included in each of a plurality of categories of a word dictionary; and labeling the consultation dataset sentence in the word dictionary based on a result of the searching.


