PrivHSD: A Reference Architecture for Privacy-Preserving Hate Speech Detection with Multi-Level Adversarial Disentanglement

Nur Akbar

Abstract


Hate speech detection (HSD) models pose an inherent privacy risk: by processing text, they inevitably encode stylistic fingerprints that enable authorship attribution, membership inference, and attribute inference, transforming protective moderation tools into potential surveillance infrastructure. We present PrivHSD, a reference architecture and validation harness for privacy-preserving HSD that operates agnostically to the author's identity. The architecture combines three complementary mechanisms—differential privacy via DP-SGD with formal (ε,δ) guarantees, multi-level adversarial identity disentanglement (MLAD) with gradient reversal at the pooler, token, and attention-head levels, and mutual information minimization via the MINE estimator, to strip identity signals from the representation space while preserving hate-relevant features. We provide a fully typed ModelConfig dataclass, a shape-annotated reference implementation (∼15.3M parameters for the ALBERT-base-v2 backbone), and a comprehensive validation harness covering shape correctness (9 tests), gradient flow (8 tests), numerical stability in bf16 (6 tests), and a privacy attack suite spanning membership inference, attribute inference, and stylometry-based re-identification risk. All 45+ unit tests pass on synthetic data, and the ablation framework systematically characterises 11 single-field config variants. This paper does not report trained-model results on real datasets—the contribution is the design artifact itself and the validation framework, enabling reproducible follow-up work on the privacy–utility Pareto frontier for content moderation

Full Text:

PDF

Refbacks

  • There are currently no refbacks.


Publisher:

Department of Electrical Engineering
Universitas Padjadjaran

Jl. Ir. Soekarno km.21, Jatinangor, Sumedang, Jawa Barat 45363