Voice Feature Vector Purchase Intention Estimation Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing consumer behavior prediction models, such as the PAD model, struggle to estimate purchase intention generated by voice stimuli, with limited study on voice features and their impact on purchasing behavior, despite voice stimuli being a significant external factor.
Innovation Solution
A consumer behavior prediction method and device that acquire voice feature quantity vectors, emotion expression vectors, and purchase intention vectors to generate a model estimating purchase intention using these inputs, leveraging techniques like neural networks and OpenSMILE for voice feature extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional PAD model is used with conventional external stimuli (store congestion, product arrangement, BGM), then purchasing behavior can be influenced by emotion changes, but voice stimulus purchase intention cannot be estimated
Solution Approach 1:
The patent extends the PAD model to handle multiple types of external stimuli including voice stimuli, not just traditional stimuli like store congestion or music. The emotion recognition system is designed to process various input modalities (visual, auditory, textual) and generate unified emotion representations that can predict purchase intention across different stimulus types, making the model universally applicable.
Solution Approach 2:
The patent introduces an intermediary emotion recognition system that bridges the gap between voice stimulus input and purchase intention prediction. This intermediary layer processes voice features (pitch, rhythm, tone) and converts them into emotion representations, which then feed into the purchase intention prediction model, enabling indirect but accurate prediction of voice-induced purchase behavior.
2Loss of information
If limited feature quantities are used (tempo of BGM), then the model is simple, but information from five senses and other factors is not captured
Solution Approach 1:
The patent segments the voice stimulus into multiple feature dimensions (pitch, rhythm, tone, volume) and processes each dimension separately through specialized extraction modules. This segmentation allows comprehensive capture of voice information while maintaining manageable complexity through modular architecture, preventing information loss from any single feature dimension.
Solution Approach 2:
The patent transitions from low-dimensional simple features (like BGM tempo) to high-dimensional multi-modal feature representations by incorporating voice-specific dimensions (pitch contours, rhythmic patterns, tonal characteristics). This dimensional expansion enables the model to capture rich information from voice stimuli and other sensory inputs without overwhelming complexity.
Data Source
AI summary
An acquisition unit acquires a voice feature quantity vector representing a feature of input voice data, an emotion expression vector representing a customer's emotion corresponding to the voice data, and a purchase intention vector representing a purchase intention of the customer corresponding to the voice data. A learning unit generates, by learning, a purchase intention estimation model for estimating a purchase intention of a customer corresponding to the voice data by using the voice feature quantity vector, the emotion expression vector, and the purchase intention vector.


