RL
post-training, distributed learning, and model analysis
- PyTorch
- TensorFlow
- GRPO
- PPO
- verl
- FSDP
- vLLM
- transformers
- sparse autoencoders
- scikit-learn
- reward engineering
- Qwen
Click the dark cells
post-training, distributed learning, and model analysis
tool-using systems with observable, verifiable behavior
simulation, stochastic systems, and chemical evidence
the infrastructure underneath training and live products

Alkera AI
GRPO policy optimization across verifiable spreadsheet workflows
LangAlpha
Tool schemas compacted per LLM call; interaction latency also fell 140 ms → 40 ms
SPEQTRO
Cross-modal candidate search across NMR, IR, and MS evidence
Water · unsupervised ML
Density denoising before the two-component mixture classifier
Chem-ICL · Ersilia
Molecular-property prediction from a small labelled context

Working notes from classes, papers, and ideas I am still trying to make precise.
31 notes · 4 folders