Multi-Modality Soft-Agent Query Population for Enterprise Virtual Assistants
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in interpreting and responding to multi-modality user requests, such as speech, text, and touch inputs, as they require complex menu-driven processes and lack the ability to determine user intent and focus effectively, especially in enterprise applications and data stores.
Innovation Solution
A method and apparatus that utilize sensor data to generate soft-queries in real-time, automatically determine user intention, and execute queries across multiple applications and data stores, providing a multi-mode response, including audio signals, to offer an interactive dashboard with relevant results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If multi-modality inputs (speech, text, touch) are used to improve user interaction convenience, then ease of operation is improved, but device complexity increases due to the need to process and integrate multiple input modalities
Solution Approach 1:
The system segments the complex multi-modality processing into distinct modules: sensor data acquisition module, soft-query generation module using object encoders, user intention determination module, and query execution module. Each module handles a specific aspect of the processing chain, making the overall complex system manageable and maintainable while preserving multi-modality input capabilities
Solution Approach 2:
The patent introduces soft-queries as an intermediary representation between raw multi-modality inputs and application-specific queries. The object encoders transform diverse sensor inputs into unified soft-query objects that capture user intent without requiring the system to directly process and integrate all modalities simultaneously, thereby reducing computational complexity
2Ease of operation
If pre-defined query formats are used to simplify query processing, then ease of operation is improved, but adaptability deteriorates because queries cannot be easily refined or augmented
Solution Approach 1:
The system employs dynamic query refinement where soft-queries are initially generated from multi-modality inputs and then adaptively adjusted based on user feedback and context. The object encoders continuously update soft-query representations to reflect user intention changes, allowing queries to be refined and augmented while maintaining a structured processing framework
Solution Approach 2:
The system incorporates feedback mechanisms where user interactions with initial query results feed back into the soft-query generation process. The object encoders use this feedback to adjust and refine subsequent soft-queries, enabling continuous query improvement while maintaining ease of operation through the standardized soft-query interface
3Measurement precision
If context-aware processing is implemented to improve response relevance, then measurement precision is improved, but loss of time increases due to additional processing requirements
Solution Approach 1:
The system performs preliminary processing of sensor data through object encoders that transform raw inputs into structured soft-queries before full context analysis is required. This preliminary encoding captures essential user intent information in advance, allowing faster context-aware processing without sacrificing intention accuracy when queries are executed
Solution Approach 2:
The system applies context-aware processing selectively rather than uniformly to all queries. The object encoders identify which aspects of sensor data require intensive context analysis and which can be processed more quickly, applying different processing depths to different parts of the input data based on their informational value
Data Source
AI summary
Methods and systems for multi-modality soft-agents for an enterprise virtual assistant tool are disclosed. An exemplary method comprises capturing, with a computing device, one or more user requests based on at least one multi-modality interaction, populating, with a computing device, soft-queries to access associated data sources and applications, and mining information retrieved by executing at least one populated soft-query. A soft-query is created from user requests. A multi-modality user interface engine annotates the focus of user requests received via text, speech, touch, image, video, or object scanning. A query engine populates queries by identifying the sequence of multi-modal interactions, executes queries and provides results by mining the query results. The multi-modality interactions identify specific inputs for query building and specific parameters associated with the query. A query is populated and used to generate micro-queries associated with the applications involved. Micro-query instances are executed to obtain results.


