Audio Intent Summarization for Low-Resource Interface Controls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Communication operations in existing systems consume significant network resources, including power, memory, and processing resources, particularly in lengthy or large data exchanges, leading to inefficient use of resources.

Innovation Solution

A system and method that dynamically generates interface controls based on audio data exchanged between devices, using machine learning algorithms to transcribe and summarize the data, determining intent, and generating visual representations in real time to reduce resource consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If traditional communication operations are used for data exchange between devices, then complete data transmission is achieved, but network resources (power, memory, processing) are significantly consumed

Engineering Contradiction:
Improvenetwork resource consumptionVSAvoiddata transmission completeness
Core Design Contradiction:
Loss of energyVSLoss of information

Solution Approach 1:

The patent extracts only the essential information from audio data by performing transcription and summarization to determine intent, rather than transmitting or processing the complete audio data. This extraction approach reduces network resource consumption while preserving the critical information needed for communication operations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary transcription and summarization of audio data to determine intent before proceeding with full communication operations. This preliminary action allows the system to identify and process only the necessary information, reducing subsequent resource consumption during actual data exchange operations.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If real-time transcription and summarization of audio data is performed using machine learning algorithms, then intent determination accuracy is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveintent determination accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by performing transcription and summarization only when necessary for intent determination, rather than processing all audio data in real-time. The system selectively applies machine learning algorithms to extract meaningful intent information, reducing overall processing time while maintaining accuracy for critical communication operations.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If interface controls are dynamically generated based on transcribed and summarized audio data, then user interaction capability is enhanced, but device complexity increases

Engineering Contradiction:
Improveinterface adaptabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal intent determination mechanism that processes various audio data types (conversations, instructions, requests) through a common machine learning pipeline. This multi-functional approach allows the system to handle diverse communication scenarios with a unified interface control generation process, managing complexity while enhancing adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250390200A1System and method to transform audio data
Publication Date: 2025.12.25 BANK OF AMERICA CORP
  • US20250390200A1 patent drawing
  • US20250390200A1 patent drawing
  • US20250390200A1 patent drawing

AI summary

A system comprises a memory communicatively coupled to at least one processor. The at least one processor is configured to obtain audio data from a user device. Further, in response to receiving the audio data, the processor is configured to execute a machine learning algorithm to transcribe the audio data into text data and summarize the text data into a data summary. The data summary is representative of a predicted intent associated with the audio data. The processor is configured to determine an interface property based on the data summary in response to summarizing the text data. The interface property is one or more communication commands to interact with the data summary. The processor is configured to determine an interface control based on the data summary and the interface property, bind the interface property to a rendered interface control, and present the rendered interface control to a workspace device.