This study examines the practical generalization limits of cross-subject EEG emotion recognition through a multimodal baseline. It focuses on the central deployment question of whether models trained on one group of participants can retain reliable performance when applied to previously unseen individuals, where physiological variation creates a substantial distribution shift. The work frames multimodal evaluation as a way to distinguish gains that are robust across subjects from results that depend on subject-specific characteristics.