Title: AI, Innovation and Global Transformation: Interdisciplinary Perspectives on Technology, Business and Society
Editors: Dr. J. Preetha, and Dr. Siddhartha Mehrotra
ISBN: 978-81-69857-64-2
Chapter: 20
DOI: https://doi.org/10.59646/809/20
Author: Partha Shankar Nayak
Abstract
Next-generation Human–Computer Interaction (HCI) demands perceptual computing interfaces capable of interpreting human communicative intent across linguistic, acoustic, visual, and spatial channels with low latency and contextual robustness. Classical unimodal paradigms often fail in unconstrained, real-world environments characterized by non-stationary acoustic noise, dynamic visual occlusions, disparate sensor clock rates, and asynchronous data arrival. This chapter presents an asymmetric, reliability-gated cross-modal transformer architecture designed for real-time natural user interaction in mixed-reality environments. The framework integrates continuous 120 Hz binocular corneal-reflection eye-gaze vectors, 60 Hz 3D hand skeleton keypoints, 16 kHz beamformed directional audio, and 30 Hz facial action units into a unified latent metric space. Evaluated on a 40-participant spatial computing testbed executing 12,000 manipulation trials, the system achieves a 94.82% overall intent recognition accuracy and an end-to-end compute latency of 22.80 milliseconds on an edge-compute envelope. The empirical findings confirm the viability of dynamic reliability-gated cross-attention in mitigating sensory dropouts during complex human-machine spatial interactions.
Keywords: Multimodal Interaction, Asymmetric Cross-Attention, Spatial Computing, Reliability Gating, Affective Computing, Latent Synchronization.