Speech Emotion Recognition for Job Interview Assessment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current recruitment processes rely heavily on human interaction and lack emotional reaction analysis, making them subjective and incomplete in assessing candidates during interviews.

Innovation Solution

A speech emotion recognition system using convolutional neural networks (CNN) and recurrent neural networks (RNN) to analyze audio clips, generating sentiment classifiers that combine with assessment results to provide a comprehensive and objective evaluation of candidate emotions, adjusting question tone and selection dynamically based on emotion responses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional human interaction interview methods are used, then emotional reaction information can be observed directly, but the assessment process becomes subjective and time-consuming

Engineering Contradiction:
Improveemotional reaction analysis accuracyVSAvoidinterview process time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual human observation and analysis of emotional reactions during interviews with an automated speech emotion recognition system using CNN and RNN models. The system processes audio clips of candidate responses to automatically classify emotions and sentiments, eliminating the need for human interviewers to manually observe and interpret emotional cues, thereby reducing subjectivity and time consumption while maintaining or improving measurement precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If automated speech analysis is implemented, then objective assessment is achieved, but emotional reaction information is lost

Engineering Contradiction:
Improveassessment objectivityVSAvoidemotional reaction information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent introduces speech emotion recognition technology as an intermediary between automated speech analysis and emotional reaction information extraction. The system uses CNN models to extract acoustic features from speech signals, then applies RNN-based sentiment analysis to interpret emotional content. This intermediary processing chain enables automated systems to objectively capture and classify emotional reactions (such as excitement, confidence, or hesitation) that would otherwise be lost in purely text-based or traditional automated assessments.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If comprehensive emotional and technical assessment is performed, then complete candidate evaluation is achieved, but system complexity increases

Engineering Contradiction:
Improvecandidate evaluation completenessVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the comprehensive assessment task into distinct functional modules: a CNN-based speech emotion recognition module that processes audio clips to extract emotional features, an RNN-based sentiment analysis module that evaluates technical content, and an integration layer that combines both assessments. This segmentation allows each module to specialize in specific aspects (emotional vs. technical evaluation) while working together to provide complete candidate assessment, managing system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

4Ease of operation

If real-time emotion feedback is provided, then interactive assessment is improved, but processing time increases

Engineering Contradiction:
Improveassessment interactivityVSAvoidreal-time processing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-training the CNN and RNN models on extensive speech and sentiment data before deployment. During actual interviews, the pre-trained models can rapidly process audio clips and provide real-time emotion feedback without requiring complex computations during the assessment itself. The system also uses efficient feature extraction and optimized inference pipelines to minimize processing time, enabling interactive assessment where candidates receive immediate feedback on their emotional expressions while maintaining fast processing speeds.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10937446B1Emotion recognition in speech chatbot job interview system
Publication Date: 2021.03.02 LTD LUOKESHI
  • US10937446B1 patent drawing
  • US10937446B1 patent drawing
  • US10937446B1 patent drawing

AI summary

Methods and systems are provided for speech emotion recognition interview process. In one novel aspect, in addition to contents assessment to an answer audio clip, the concurrent sentiment classifier is generated based on emotion classifier of the answer audio clip. In one embodiment, the computer system obtains a sentiment classifier of an audio clip of a first answer to the first question, wherein the sentiment classifier is derived from an emotion classifier resulting from a convolutional neural network (CNN) model analysis of the audio clip; obtains an assessment result to the first question by analyzing the audio clip of the first answer to the first question using a recurrent neural network (RNN) model; and generates a first emotion response result to the first question based on the sentiment classifier and the assessment result, wherein the first emotion response result presents a sampling experience factor to the response assessment result.