Analytics · Information Systems · Data Science
I am an undergraduate researcher working on natural language inference and interdisciplinary AI, with a focus on the interpretability, reliability, and robustness of machine learning models within explainable AI and contexts relating to autonomous systems under distribution shift. My interests extend to human-AI interaction, cross-domain adaptation, scientific claim verification, and representation learning for unstructured data, particularly NLP and NLI, using transfer learning and embeddings.
Profile Snapshot

Education
Korea University Business School
Undergraduate, BBA
Concentration: IS · Analytics
Fields of Study
IS, DS/AI, Machine Learning, NLP, Business Analytics
Current Work (Sept 2026)
Research — Finishing up some work on cross-domain ML and scientific AI under supervision.
Preprints and Write-ups — Write-ups for research on SciFact leakage correction and interpretable DL on molecular representations (See. Selected Projects).
Graduate School Applications — Programmes in AI, Data Science, and Information Systems
About

Korea University Business School (KUBS)
Bachelor of Business Administration
Expected Graduation: February 2027 (Early Graduation)
Concentration: Information Systems & Analytics
Completed Tracks:
- Business Analytics (비즈니스애널리틱스)
- Artificial Intelligence for Business (AI와경영)
Relevant Coursework
Academic Experience
KUBS Official Undergraduate Tutor
Mar 2026 - Present
Tutoring Subjects:
- Business Analytics II
- Business Statistics
- Management Information Systems
Awards & Honors
- Dean's Award
- Admission/Excellence/Top Scholarships, Korea University - 2023, 2024, 2025, 2026
- Best in AI Award Category @ 2026 KU AI Forum Poster Session, Korea University DHUSS
Technical Skills
Programming: Python, R, SQL (MySQL, PostgreSQL), C++
Machine Learning & AI: PyTorch, TensorFlow, scikit-learn, Hugging Face Transformers & Hub, Captum, MLflow, AWS, LangChain, LangGraph
Data & Statistics: pandas, NumPy, SciPy, Excel, statsmodels
Visualization: matplotlib, seaborn, ggplot2, Tableau
Web & Misc: FastAPI, Docker, HTML/CSS/JS, TypeScript, SeleniumBase, Git, GitHub Actions
Research Tools: LaTeX, Overleaf, Zotero
Research Interests
NLP & NLI
- Multimodal Scientific Claim Verification
- Unstructured Data & Text Analytics
- Information Retrieval
- Contextual Representation & Geometry
ML
- Cross-domain Adaptation
- Representation Learning
- Transfer Learning
- Robustness & Learning Theories
- Semantic Collapse & LLM Hallucination
X-Autonomous Systems & AI+X
- Explainable AI (XAI)
- Interpretability of ML Models in AI
- Feature Attribution and Other Methods
- Reliable Autonomous Systems under Distribution Shift
Showcase
clAIm
An end-to-end scientific claim verification system built on a two-stage fine-tuned DeBERTa-v3-base with retrieval augmentation and explainability. Identified and corrected bilateral data leakage affecting 38.4% of the SciFact development split. Macro-F1: 0.9043 (MultiNLI), 0.8640 (SciFact, oracle), 0.7713 (SciTail, zero-shot). Deployed with FastAPI, Docker, and Vercel; integrates Semantic Scholar retrieval, Integrated Gradients (Captum), and GPT-OSS-20B explanations.
Best in AI Award Category w/ Research Grant
2026 KU AI Forum Poster Session
“A Tale of Two Splits: Detecting and Correcting Bilateral Leakage in SciFact”
Interpretable Deep Learning of Structure–Toxicity Relationships
With S. T. Hmuu, School of Biosystems and Biomedical Sciences, Korea University.
A pre-registered, causal faithfulness study comparing sequence (ChemBERTa) and graph (GIN) neural network representations for interpretable molecular toxicity prediction, using perturbation-based metrics (comprehensiveness, sufficiency, deletion and insertion AUC) validated directly against curated ground-truth toxicophore substructures. Diagnosed a class-direction confound in metric aggregation that, left uncorrected, would have reversed one of the four primary findings, then resolved it through subgroup analysis and recomputed the full result set before drawing conclusions. Found ChemBERTa's attributions substantially and consistently more faithful than GIN's across all four metrics (n = 1,457, p < 10⁻⁸⁰), with the effect stable under molecular distribution shift and replicated on an independent, differently imbalanced endpoint (Tox21/SR-ARE).
VeriScite
An autonomous agentic system for citation-faithfulness verification, orchestrating dual independent verification (fine-tuned DeBERTa-v3 NLI + zero-shot LLM) with a LangGraph ReAct planner that resolves disagreement via tool use — no human intervention required. Caught a shared-bias case where both verifiers agreed on an incorrect SUPPORT verdict at 0.9931 confidence, and corrected it to NOT_ENOUGH_INFO through autonomous escalation across three bounded agent actions, streamed live via Server-Sent Events.
Transformer Encoder from First Principles
Diagnosed representation anisotropy in transformer encoder outputs — tracing the cause through recent literature (Godey et al., 2024) — and corrected it via GloVe-based embedding initialization, validated through PCA geometric analysis. Built the full encoder from scratch in NumPy (multi-head self-attention, sinusoidal positional encoding, GELU, layer norm, residual connections, following Vaswani et al. 2017) to get direct inspection and control over the representation space itself.
Referees
Prof. Kyuhan Lee
Assistant Professor of Information Systems
Prof. Byungwan Koh
Professor of Information Systems · Area Chair in Information Systems
Prof. Angela Aerry Choi
Associate Professor of Information Systems
Prof. Gunwoong Lee
Associate Professor of Information Systems
