SGENLOG: SUPERVISED CONTRASTIVE BERT FINE-TUNING FOR LOG-BASED ANOMALY DETECTION
DOI:
https://doi.org/10.56651/lqdtu.jst.v15.n1.1250.ictKeywords:
Log analysis, anomaly detection, contrastive learning, representation learningAbstract
Log-based anomaly detection remains challenging due to extreme class imbalance and the “shallow” representations of current models. This paper proposes a sequence-level generative log framework (SGENLog). SGENLog is driven by a novel supervised contrastive finetuning strategy, named supervised BERT (SBERT). Unlike traditional methods, SBERT leverages Supervised Contrastive Loss (SupCon) on log sequences joined via [SEP] tokens to capture deep, discriminative context even when anomalous data is nearly non-existent. These enhanced embeddings are then integrated into a multi-task Transformer encoder-decoder (called GENLog), where the reconstruction of input acts as a regularizer to mitigate overspecialization. Extensive evaluations on the Blue Gene/L (BGL), Hadoop Distributed File System (HDFS), and Thunderbird (TDB) datasets under extremely low anomaly ratios (e.g., 0.2%) demonstrate that SBERT is the primary catalyst for performance, enabling SGENLog to significantly surpass state-of-the-art (SOTA) models like NeuralLog and CLDTLog. The results confirm that SBERT provides the robust, context-aware representations necessary for real-world, data-scarce system monitoring.










