Voice-Based Trauma Screening Using Deep Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for diagnosing trauma, such as post-traumatic stress disorder (PTSD), face challenges due to social prejudice and the difficulty in recognizing mental illness, and existing voice analysis studies are limited, with no effective trauma screening using voice data available.
Innovation Solution
A device and method utilizing deep learning to screen for trauma by converting voice data into image data, recognizing emotions like fear and sadness, and estimating trauma probability through a deep learning model, improving accuracy with post-processing techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice data is used for trauma screening, then non-contact screening is enabled without patient rejection, but existing voice analysis studies are limited and no effective trauma screening is available
Solution Approach 1:
The patent introduces image data as an intermediary representation of voice data. Voice signals are converted into visual spectrograms, allowing deep learning models trained on image data to analyze emotional states. This intermediary transformation enables effective trauma screening by bridging the gap between voice-based input and proven image-based deep learning techniques for emotion recognition.
Solution Approach 2:
The patent replaces traditional mechanical or direct human assessment methods with deep learning-based automated analysis. Instead of requiring direct patient-doctor interaction or contact-based assessments, the system uses automated deep learning models to analyze voice patterns and detect trauma indicators, thereby enabling non-contact screening while maintaining or improving reliability.
2Measurement precision
If deep learning is applied to voice data for trauma screening, then emotion recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The patent creates a visual copy or representation of voice data in the form of spectrograms. Instead of directly analyzing raw voice signals, the system generates image-based copies that capture the essential emotional characteristics. This copying approach allows the use of established image processing deep learning models, improving accuracy while managing complexity by leveraging existing proven techniques rather than developing new voice-specific models from scratch.
3Measurement precision
If voice data is converted into image data for deep learning, then trauma screening accuracy is improved, but processing complexity increases
Solution Approach 1:
The patent transforms one-dimensional voice signal data into two-dimensional image data (spectrograms). This dimensional transformation converts temporal voice patterns into spatial visual representations, enabling the application of powerful 2D image processing deep learning architectures. The dimensionality change improves accuracy by allowing the model to simultaneously analyze multiple voice characteristics across different time and frequency dimensions, while the conversion process itself is achieved through standard signal processing techniques.
Data Source
AI summary
This application relates to a device and a method for voice-based trauma screening using deep learning. The device and method for voice-based trauma screening using deep learning screen for trauma through voices that may be obtained in a non-contact manner without limitations of space or situation. In one aspect, the device includes a memory configured to store at least one program and a processor configured to perform an operation by executing the at least one program. The processor can obtain voice data, pre-process the voice data, convert pre-processed voice data into image data, and input the image data to a deep learning model and obtain a trauma result value as an output value of the deep learning model.


