Terminal Sound Classification for Network Load Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing action estimation techniques in network environments, such as cloud networks, face challenges in reducing network load and achieving accurate action estimation due to the high volume of data from ultrasonic sounds, which are more susceptible to noise and require wider frequency bands with shorter sampling periods.
Innovation Solution
A two-phase structure comprising a terminal and a computer connected via a network, where the terminal identifies and outputs only non-steady sounds to the computer, which then estimates actions using a trained model, reducing data transmission and load by focusing on sound information in specific frequency bands and utilizing machine learning for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all sound information including ultrasonic sounds is transmitted to the computer for action estimation, then the action estimation accuracy can be improved, but the network load increases significantly
Solution Approach 1:
The terminal extracts only non-steady sound components from the collected sound information using a first trained model, and transmits only these extracted components to the computer. This extraction principle reduces the transmitted data volume while preserving the essential information needed for accurate action estimation, as steady sounds (background noise) are filtered out locally.
Solution Approach 2:
The terminal performs preliminary processing of sound information by classifying it into steady and non-steady components before transmission. The first trained model pre-processes the audio data, identifying and separating non-steady sounds that contain action-related information, thereby reducing the burden on the network and the computer while maintaining estimation accuracy.
2Difficulty of detecting and measuring
If ultrasonic frequency bands are used for sound collection, then the sound detection capability is improved, but the sampling rate requirement increases leading to larger data volume
Solution Approach 1:
The sound frequency spectrum is segmented into steady and non-steady components using machine learning models. By segmenting the audio data in the frequency domain and processing only the non-steady portions, the system maintains high detection capability for ultrasonic sounds while reducing the overall data volume that requires high-speed processing.
Solution Approach 2:
The first trained model acts as an intermediary between the sound collector and the action estimation process. It processes the raw audio data locally, extracting relevant non-steady sound features before transmission, thereby mediating between the high sampling rate requirements of ultrasonic detection and the data processing efficiency of the network system.
3Quantity of substance
If steady sound filtering is performed at the terminal before transmission, then the network load is reduced, but the device complexity at the terminal increases
Solution Approach 1:
The terminal performs self-service by autonomously classifying and filtering sound information using locally deployed trained models. The terminal independently identifies non-steady sounds and prepares the filtered data for transmission without requiring external assistance, thereby reducing network load while accepting the trade-off of increased local processing capability.
Solution Approach 2:
The system changes the processing parameters by applying machine learning models that analyze sound characteristics in the frequency domain. The first trained model transforms raw audio data into classified sound information, changing the data representation from time-domain waveforms to frequency-domain features, which enables efficient filtering and reduces transmitted data volume.
Data Source
AI summary
In an information processing system, an estimation as to whether a sound collected by a microphone is steady sound or non-steady sound is executed, sound information estimated to indicate the non-steady sound is transmitted to a server as output sound information when it is estimated that it is the non-steady sound, and the server acquires the output sound information, and estimates an action of a person from a resulting output obtained by inputting the output sound information to a second trained model indicative of a relevance between the output sound information and action information on an action of a user.


