Systems, methods and computer-readable media for providing explainable
selective attention in multi-source or multi-speaker environments. A
selective attention module receives multimodal sensor data including audio, video,
gaze, text, and physiological signals from a plurality of sources. An attention
inference engine generates attention distributions over the sources and fuses them into a probabilistic belief state. An explainability module produces interpretable outputs corresponding to the fused belief, including attention matrices, confidence scores, reliability measures, margin-based differentiators, and
natural language rationales. The explainability outputs are rendered through visual, auditory, or augmented /
virtual reality interfaces to indicate the attended source, suppressed sources, and reasoning for the selection. The
system enables user interaction by providing justifications in real time,
logging explanations for
retrospective analysis, and supporting
adaptation of thresholds and model weights based on feedback. The disclosed technology improves transparency,
interpretability, and trust in
selective attention systems, while maintaining real-time performance in dynamic multi-speaker environments.