Joint Speech Text Sentiment Detection Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition and sentiment detection systems face inefficiencies due to separate processing of textual and acoustic data, leading to increased processing load and latency, and reduced battery life in communication devices.
Innovation Solution
A method and apparatus that combine textual and acoustic data to detect sentiment in speech content, utilizing a processor and memory to analyze speech data, determine locations and timestamps, and generate reviews based on predefined sentiments, thereby integrating sentiment detection into communication devices efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If separate processing is used for textual and speech-based sentiment tasks, then processing independence is maintained, but processing load and latency increase and battery life decreases
Solution Approach 1:
The patent combines separate textual sentiment analysis and acoustic sentiment analysis into a unified joint processing framework. The system processes text and speech features simultaneously through integrated neural network models, merging previously independent processing pipelines into a single coordinated system that shares computational resources and reduces overall processing load.
Solution Approach 2:
The patent creates a universal sentiment detection system that handles multiple types of input (textual data, acoustic data, speech features) through a single multi-functional processing architecture. This unified system can process different modalities of sentiment information simultaneously, making the processing framework adaptable to various input types without requiring separate specialized processors.
2Device complexity
If separate processing is used for textual and speech-based sentiment tasks, then modular processing is maintained, but battery life is reduced
Solution Approach 1:
The patent merges separate processing modules for textual and acoustic sentiment analysis into an integrated joint processing system. By combining these modules, the system eliminates redundant computations and shared memory access operations that would occur if separate modules processed the same speech input independently, thereby reducing overall energy consumption and extending battery life.
3Productivity
If joint processing of textual and acoustic data is used, then processing efficiency is improved, but system complexity increases
Solution Approach 1:
The patent segments the joint processing system into distinct functional components: a feature extraction module that processes raw speech and text inputs, a sentiment analysis module that applies neural network models, and an output generation module that produces sentiment classifications. This segmentation allows each component to be optimized independently while maintaining overall system efficiency, managing complexity through modular functional decomposition.
4Measurement precision
If joint processing of textual and acoustic data is used, then sentiment detection accuracy is improved, but processing resources increase
Solution Approach 1:
The patent dynamically adjusts processing parameters such as feature extraction depth, neural network model complexity, and processing batch sizes based on available computational resources and input characteristics. This allows the system to maintain high sentiment detection accuracy by using more sophisticated processing when resources are abundant, while automatically reducing resource consumption when processing capacity is constrained, thus balancing accuracy with resource efficiency.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus for generating a review based in part on detected sentiment may include a processor and memory storing executable computer code causing the apparatus to at least perform operations including determining a location(s) of the apparatus and a time(s) that the location(s) was determined responsive to capturing voice data of speech content associated with spoken reviews of entities. The computer program code may further cause the apparatus to analyze textual and acoustic data corresponding to the voice data to detect whether the textual or acoustic data includes words indicating a sentiment(s) of a user speaking the speech content. The computer program code may further cause the apparatus to generate a review of an entity corresponding to a spoken review(s) based on assigning a predefined sentiment to a word(s) responsive to detecting that the word indicates the sentiment of the user. Corresponding methods and computer program products are also provided.