About
Welcome to Fenriar Tech!
This is a personal technical blog focused on Large Language Models (LLM), GPU High-Performance Computing (CUDA), Deep Learning Architecture, Systems Programming, and Applied Algorithms.
About the Author
Hi! I’m Sean Wang, a software engineer and researcher passionate about AI systems, low-level performance optimization, and deep learning engineering.
Focus Areas
- LLM Systems & Inference Optimization: Quantization (GPTQ/AWQ), Operator Fusion (FlashAttention), Instruction Tuning, and KV Cache optimizations.
- GPU High-Performance Computing: CUDA Kernel engineering, parallel algorithms, shared memory, and hardware architecture.
- Inference Engines: TensorRT custom plugin development and real-time model deployment.
- Core Algorithms & Type Theory: Functional programming paradigms, Monads, and mathematics behind machine learning.
Stay tuned for upcoming technical articles and deep dives!