Besides, I find immense joy in pushing my boundaries by tackling complex issues
that have the potential to positively impact the public. My fervor
converges with enhancing computer systems,
networks, databases, and their corresponding
security aspects, which allows me to have impacts
on computing for both the environment (greener while efficient)
and people (faster while reliable). I am also a passionate free software developer,
I maintain several open-source projects and contribute to
many others.
|
05/2026 - 07/2026
Deep Learning Algorithm Engineering Intern, NVIDIA
(NeMo and Megatron (HQ))
• Add and optimize delta-compressed weight and GPU Direct RDMA P2P refit support for NeMo-RL
• Add rollout real quant aware RL support for NVFP4 W4A16 and W4A4 with NeMo-RL
• Add cuDNN fused DSA kernel support with THD packing and Context Parallelism for Megatron
|
|
08/2025 - 11/2025
Reinforcement Learning Infrastructure Engineer, Fellou
(world’s first agentic browser)
• Building infrastructures for Agentic RL training of Vision Language Models (VLMs) on top of verl
• Verl is Volcano Engine & ByteDance Seed’s Reinforcement Learning framework for LLMs
|
|
04/2025 - 08/2025
Reinforcement Learning Algorithm Engineer, xDAN AI
(Agentic RL training with verl)
• Developing solutions for Agentic LLMs training using reinforcement learning
• Targeting areas include coding, math, finance, deep researcher, and RAG enhancement
|
|
01/2025
AI Prompt Engineer, Outlier
(create RLHF datasets for code generation)
• Crafting and answering questions related to computer science in order to help train AI models
• Evaluating and ranking code generated by AI models
|
|
06/2024 - 10/2024
OSPP Summer, OceanBase
(add support for the IVFPQ vector index algorithm)
• Enable more vector indexing capabilities for OceanBase - a multi-cloud distributed SQL database
• OceanBase is developed by Ant Group (Alibaba), supports the peak traffic in Double 11 Shopping Festivals
|
|
05/2024 - 09/2024
Mentor for Google Summer of Code 2024, MIT CSAIL & OpenSUSE
(simultaneously)
• Mentor Chang Min Bark for Updating the Multi-select Plugin in Blockly and Cross-testing Other Plugins
• Mentor Jiamin Wang for Refactor Customize IBus with Modern Interface and Extensibility
|
|
01/2024 - 07/2024
Junior Application Specialist, CSC - IT Center for Science
(master’s thesis)
• Implement the real-time GPU usage alert service for HPC clusters
• Deploy Large Language Models (LLM) inference service in HPC environment with vLLM
|
|
06/2023 - Present
Founder, Hollow Software
(company registered and established in Finland, business ID: 3369577-7)
• Freelance for full-stack computer programming tasks
|
|
06/2023 - 05/2024
Developer Relations Engineer, Enclaive
(confidential computing, with Prof. Sebastian Gajek)
• Create technical documentation, tutorials and content
• Manage channel to interact with the community and grow the customer base
|
|
06/2023 - 12/2023
Coding Experience 2023 I & II, Igalia
(Wolvic, ex. Mozilla Firefox Reality)
• Develop VR Browser, refactor Android deprecated methods
• Address user issues. Implement UI, graphics, browser, and openXR related features
• Contribute to the majority of the features available from v1.4.2 - v1.6
|
|
05/2023 - 08/2023
HPC Developer Internship, CSC - IT Center for Science
• Develop the GPU monitoring service for HPC clusters (NVIDIA and AMD GPUs, including LUMI)
• Develop the Documentation GPT chatbot service using llamaIndex, LangChain and Redis/Qdrant
• Make improvements to the HPC web operating interface project Open OnDemand using Ruby on Rails
• LUMI ranked 3rd around the world and 1st in Europe at HPC supercomputer TOP500 list by 06/2023
|
|
03/2023 - 05/2023
CNCF LFX Mentorship, Kubescape
(Kubernetes security platform)
• Add packaging support for various package managers in Linux, macOS, and Windows
• Improve release CI by adding ARM64 binaries building and testing support (QEMU and cross compiling)
• Add support for Kubescape auto version bumping in other Kubescape-related repositories
• Enable Kubescape GitHub action to support code review based on generated SARIF report
• A blog post by Cloud Native Computing Foundation (CNCF) that introduces my work
|
|
07/2022 - 09/2022
Rhino-Bird Open-Source Training Program, Tencent
(Kona JDK)
• Investigate keypairs generation for SM2 with JDK, compare performance and security of algorithms
|
|
05/2022 - 09/2022
Google Summer of Code, MIT CSAIL & Google
(Blockly, web-based visual program editor)
• Create a plugin to allow selecting, dragging and manipulating multiple blocks at once for Google Blockly
• Resolve the issue that had troubled the Google Engineers for a decade (since the Blockly project started)
• @mit-app-inventor/blockly-plugin-workspace-multiselect now has an average of 250+ downloads per week
|
|
05/2022 - 08/2022
Summer of Bitcoin, Cryptoanarchy Debian
(Bitcoin related Debian packages)
• Add Continuous Integration (CI) by investigating systemd support inside Docker with GitHub Actions
• The chance to get into this internship in 2022: 83/20317
|
|
05/2022 - 08/2022
RL Open Source Fest, Microsoft Research NYC
(Vowpal Wabbit, ML library)
• Add native CSV parser support for the Vowpal Wabbit project in C++, 100% Test & Code Coverage
• Manage to make 10x performance improvement and is comparable to speed of original VW text format
|
|
01/2022 - 07/2022
Open Source Internship, Institute of Software, CAS
(openEuler, Linux distribution)
• First to obtain 150 intern credits. Help solve 6 issues, related technologies include D-Bus, Glib, building Cloud Native Container images, Rust programming, Chrome DevTools Protocol, test case, RPM packaging
• One of the achievements mdbook-pdf now has an average of 500+ downloads per week
|
|
06/2021 - 08/2021
Google Summer of Code, openSUSE
(IBus, input method framework)
• Add customizing themes feature for IBus in desktop environments, improve IBus settings support in GNOME
• A blog post at openSUSE News that introduces my work
• GNOME extension Customize IBus now has an average of 300+ downloads per week
|
|
07/2021 - 09/2021
OSPP Summer, openEuler
(Developed an ISO writer using QT and C++)
|
|
07/2020 - 09/2020
OSPP Summer, Emacs Application Framework
(EAF, work with Yong Wang)
• Add new features and address all kinds of issues using Emacs Lisp, Python, JavaScript, SQLite
• A blog post (in Chinese) about my work posted by Yong Wang (ManateeLazyCat)
|
|
07/2020 - 08/2020
Summer of Code, Alibaba
(Arthas, Alibaba’s most starred project on GitHub, Java diagnostic tool)
• Create Chinese and English interactive online tutorials for Arthas using Katacoda and Vue.js
• Coding part highlight: Support directly searching Class Loader by name instead of Hash Code in Java
• A Wechat post (in Chinese) about my work posted by hengyunabc
|
1.7k GitHub followers.
198.6k GitHub stars across all of the following selected repositories:
|
1.
|
Nereus: Adaptive Parallelism for LLM Post-Training
[abstract]
Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang, Mario Di Francesco, and Bo Zhao
preprint arXiv:2609.34645 2026
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job’s distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14-7.27× over OpenRLHF and by 1.10-1.47× over Verl across diverse clusters.
|
|
2.
|
Accelerating GLM-5.2 Long-Context Training with cuDNN Fused DSA Kernels and THD-Packed Context Parallelism using Megatron Bridge
[abstract]
Songlin Jiang, Yu Yao, and Wenwen Gao
Announcements at NVIDIA NeMo Megatron-Bridge 2026
Long-context training hits a wall: standard attention costs O(S²), so doubling the sequence length quadruples the attention work. DeepSeek Sparse Attention (DSA), used by GLM-5.2, breaks that curve by using a lightweight indexer that selects the top-k most relevant key/value tokens per query, and the main attention runs only over that top-k set, turning the dominant term near-linear. But a sparse-attention layer is only as good as its kernels: a naive implementation materializes several S × S intermediates that are memory-prohibitive past 32K tokens. NVIDIA closes that gap with fused DSA kernels in NVIDIA cuDNN, wired through cuDNN frontend, Megatron Core, and NeMo Megatron Bridge with support for context parallelism (CP), packed sequences (THD), and GLM-5.2’s new IndexShare. On GLM-5.2 743B training, the cuDNN backend delivers ~3.4× higher throughput than the TileLang reference on NVIDIA GB200 (~2× on NVIDIA H100), at an equivalent memory footprint.
|
|
3.
|
Router Replay R3: Why It Failed and How We Fixed It
[abstract]
Songlin Jiang, Yiwen Lu, Qihan Liu, Andrew Chen, Pony Ma, and Mind Lab
Blog Post at Mind Lab: A Lab for Experiential Intelligence 2026
Training-inference mismatch (TIM) is the silent failure mode for MoE: the training path and the inference path diverge, gradients stop matching the policy you deploy, and optimization drifts. In this blog post, we walk through a real case where TIM is much worse on a DeepSeekV3-architecture model (Moonlight-16B-A3B) than on Qwen3 MoE models (Qwen3-30B-A3B). We investigate why common mitigations failed, and how we fixed Router Replay R3 across vLLM and veRL to remove the mismatch without discarding samples.
|
|
4.
|
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
[abstract]
Mind Lab, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Songlin Jiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Aaron Guan, Jun Gao, Pyke Han, Nolan Ho, and others
preprint arXiv:2608.09819 2026
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti (748B) combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-35B-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned Harness Context Protocol contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.
|
|
5.
|
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
[abstract]
Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Songlin Jiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, and Cheng Jin
preprint arXiv:2607.14952 2026
Long-context RL post-training is constrained by the lifetime of state and gradients, not attention cost alone. In GRPO, one multi-million-token prompt must serve old-policy and reference scoring plus multiple policy responses, while conventional autograd keeps the prompt graph and all response graphs live alongside model weights, caches, and distributed communication buffers. We present LongStraw, an objective-aware, architecture-aware system for resident-state virtualization, response replay, and distributed-gradient execution. Its transaction captures the shared prompt without autograd, retains only the architecture-required state on explicitly owned pages, restores that state for each group member, scores old/reference branches without a graph, replays one policy response at a time with autograd, and accumulates the resulting gradients before one distributed finalization and optimizer step. This schedule bounds the live training graph by the response suffix while reusing the expensive prompt computation across the complete GRPO group. We instantiate this design for two incompatible model structures. Qwen3.6-27B combines 48 recurrent GDN layers with 16 full-attention layers; LongStraw keeps the compact recurrent state and physically CP8-sharded KV pages, composes global attention through cross-rank LSE/output merging, and performs blockwise response replay. GLM-5.2 combines a 78-layer MLA/DSA attention stack with a 256-expert, top-8 MoE tail. Its implementation keeps CP-sharded MLA latent pages and DSA indexer-key pages in CPU memory, stages one layer at a time, reconstructs IndexShare-aware global sparse selection over CP32, and dispatches routed response tokens over EP32. The two paths share one transaction contract while specializing the retained state, replay operator, and collective communication to the architecture.
|
|
6.
|
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
[abstract]
Mind Lab, Vin Bo, Song Cao, Vic Cao, Andrew Chen, Kaijie Chen, Cleon Cheng, Songlin Jiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Hongquan Gu, Aaron Guan, Nolan Ho, and others
preprint arXiv:2606.02437 2026
Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapters as persistent local state on top of strong shared foundation models. In this framing, the base model provides shared competence while adapters carry instance-specific behavior such as preferences, skills, tool habits, and memory-like updates. We organize the problem around three scaling axes: Scale Up, where stronger shared priors make small local updates more useful; Scale Down, where we study how small adapters can be while remaining reliable; and Scale Out, where many persistent adapted instances coexist. MinT provides one infrastructure example for managing adapter identity, revision, provenance, evaluation, and serving residency. Together, the results suggest that PEFT can be a compact substrate for persistent personal models rather than only a budget substitute for full fine-tuning.
|
|
7.
|
MinT: Managed Infrastructure for Training and Serving Millions of LLMs
[abstract]
Mind Lab, Song Cao, Vic Cao, Andrew Chen, Kaijie Chen, Cleon Cheng, Songlin Jiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Hongquan Gu, Aaron Guan, Nolan Ho, and others
preprint arXiv:2605.13779 2026
We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a setting where many trained policies are produced over a small number of expensive base-model deployments. Instead of materializing each policy as a merged full checkpoint, MinT keeps the base model resident and moves exported LoRA adapter revisions through rollout, update, export, evaluation, serving, and rollback, hiding distributed training, serving, scheduling, and data movement behind a service interface. MinT scales this path along three axes. Scale Up extends LoRA RL to frontier-scale dense and MoE architectures, including MLA and DSA attention paths, with training and serving validated beyond 1T total parameters. Scale Down moves only the exported LoRA adapter, which can be under 1% of base-model size in rank-1 settings; adapter-only handoff reduces the measured step by 18.3x on a 4B dense model and 2.85x on a 30B MoE, while concurrent multi-policy GRPO shortens wall time by 1.77x and 1.45x without raising peak memory. Scale Out separates durable policy addressability from CPU/GPU working sets: a tensor-parallel deployment supports 10^6-scale addressable catalogs (measured single-engine sweeps through 100K) and thousand-adapter active waves at cluster scale, with cold loading treated as scheduled service work and packed MoE LoRA tensors improving live engine loading by 8.5-8.7x. MinT thus manages million-scale LoRA policy catalogs while training and serving selected adapter revisions over shared 1T-class base models.
|
|
8.
|
OpenRLHF: A Ray-based Easy-to-use, Scalable and High-performance RLHF Framework
[abstract]
Jian Hu, Xibin Wu, Wei Shen, Jason Klein Liu, Weixun Wang, Songlin Jiang, Haoran Wang, Hao Chen, Bin Chen, Wenkai Fang, Xianyu, Yu Cao, Haotian Xu, and Yiming Liu
Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP): System Demonstrations 2025
Large Language Models (LLMs) fine-tuned via Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning with Verifiable Rewards (RLVR) significantly improve the alignment of human-AI values and further raise the upper bound of AI capabilities, particularly in reasoning-intensive, long-context Chain-of-Thought (long-CoT) tasks. However, existing RLHF (or RLVR) frameworks commonly face challenges such as inference bottlenecks and complexity barriers, restricting their accessibility for newcomers. To bridge this gap, we introduce OpenRLHF, a user-friendly, scalable, and easy-to-learn open-source RLHF framework built upon Ray, vLLM, DeepSpeed, and HuggingFace Transformers, featuring a simplified design, clear code structure, and comprehensive documentation to facilitate entry for researchers and practitioners. Experimental results show that OpenRLHF achieves superior training efficiency with speedups ranging from 1.22texttimes to 1.68texttimes across different model sizes compared to state-of-the-art frameworks, while requiring significantly fewer lines of code for implementation. OpenRLHF is publicly available at https://github.com/OpenRLHF/OpenRLHF, and has already been adopted by leading institutions to accelerate RLHF research and learning.
|
|
9.
|
Accelerating RLHF with vLLM, Best Practice from OpenRLHF
[abstract]
Jian Hu, Songlin Jiang, Zilin Zhu, Xibin Wu, Kaichao You, Cody Yu, and Rui Qiao
Blog Post at vLLM 2025
As demand grows for training reasoning-capable large language models (LLMs), Reinforcement Learning from Human Feedback (RLHF) has emerged as a cornerstone technique. However, conventional RLHF pipelines—especially those using Proximal Policy Optimization (PPO)—are often hindered by substantial computational overhead. This challenge is particularly pronounced with models that excel at complex reasoning tasks (such as OpenAI-o1 and DeepSeek-R1), where generating long chain-of-thought (CoT) outputs can account for up to 90% of total training time. These models must produce detailed, step-by-step reasoning that can span thousands of tokens, making inference significantly more time-consuming than the training phase itself. As a pioneering inference framework, vLLM provides a user-friendly interface for generating RLHF samples and updating model weights.
|
|
10.
|
Building trillion-parameter reasoning RL with 10% GPUs
[abstract]
Qihan Liu, Songlin Jiang, Rio Yang, Alex Yin, Pony Ma, Andrew Chen, and Mind Lab
Blog Post at Mind Lab: A Lab for Experiential Intelligence 2025
We present what we believe is the first end-to-end Reinforcement Learning (RL) with Low-Rank Adaptor (LoRA) on a trillion-parameter reasoning model. Our system runs on large Mixture-of-Experts (MoE) models with 10% GPUs compared to conventional full-parameter RL. Our solutions have also been contributed to major open-source projects: NVIDIA’s Megatron-Bridge and Volcengine’s verl.
|
|
11.
|
QUASAR: Quantum Assembly Code Generation Using Tool-Augmented LLMs via Agentic RL
[abstract]
Cong Yu, Valter Uotila, Shilong Deng, Qingyuan Wu, Tuo Shi, Songlin Jiang, Lei You, and Bo Zhao
preprint arXiv:2510.00967 2025
Designing and optimizing task-specific quantum circuits are crucial to leverage the advantage of quantum computing. Recent large language model (LLM)-based quantum circuit generation has emerged as a promising automatic solution. However, the fundamental challenges remain unaddressed: (i) parameterized quantum gates require precise numerical values for optimal performance, which also depend on multiple aspects, including the number of quantum gates, their parameters, and the layout/depth of the circuits. (ii) LLMs often generate low-quality or incorrect quantum circuits due to the lack of quantum domain-specific knowledge. We propose QUASAR, an agentic reinforcement learning (RL) framework for quantum circuits generation and optimization based on tool-augmented LLMs. To align the LLM with quantum-specific knowledge and improve the generated quantum circuits, QUASAR designs (i) a quantum circuit verification approach with external quantum simulators and (ii) a sophisticated hierarchical reward mechanism in RL training. Extensive evaluation shows improvements in both syntax and semantic performance of the generated quantum circuits. When augmenting a 4B LLM, QUASAR has achieved the validity of 99.31% in Pass@1 and 100% in Pass@10, outperforming industrial LLMs of GPT-4o, GPT-5 and DeepSeek-V3 and several supervised-fine-tuning (SFT)-only and RL-only baselines.
|
|
2024 - 2025
|
|
2023 - 2024
Aalto University School of Science Dean's Incentive Scholarship
|
|
2022 - 2024
SECCLO Erasmus Mundus Scholarship
(26/767) Full-ride scholarship, includes full tuition fee waiver, covers all living costs and travel expenses
|
|
2024
|
|
2024
Travel Fund for The Linux Foundation Open Source Summit Europe 2024
To speak at the LFX Mentorship Showcase
|
|
2024
Certified Kubernetes Security Specialist ( CKS), CNCF
|
|
2023
Black Hat Europe 2023 Student Scholarship
|
|
2022
|
|
2022
Huawei Cloud Certified Developer Associate in Internet of Things (HCCDA - IoT)
|
|
2022
Certified Kubernetes Administrator ( CKA), CNCF
|
|
2022
Recognized Tencent Open Source Contributor
Only 30+ were issued globally by the end of 2022
|
|
2022
|
|
2022
|
|
2021
|
|
2021
|
|
2020
|
|
2020
CATTI Level 3 (English & Chinese) Translator
|
|
2020
Best Paper Award in Lanzhou University Summer Social Practice Report
Paper: Research on the Functionality and Optimization of Urban Communities in China During COVID-19 in 2020
|
|
2020
|
|
2019
Lanzhou University Outstanding Undergraduate International Exchange Scholarship
(Only 1) Full covering of visa fees, traveling fees, and living costs, plus the host university exchange program tuition fee waiver agreement
|
|
2019
|
|
2019
|
|
2018 - 2022
Nine Lanzhou University Undergraduate Scholarships
3 x scholarship for academic excellence,
5 x scholarship for innovation and entrepreneurship projects,
2020 scholarship for IELTS test fee
|