Notes
IV. SIGGRAPH ’26 interesting papers & Coding Agents
SIGGRAPH ’26 interesting papers
Geometric processing: Point clouds are easy to collect, but are hard to work with. It is inevitable that current geometries of all kind (e.g., surfaces) needs to work with the rising prevelance of point clouds.
- Numerical Geometry a la Mode
- Winding Numbers & Surface Representations
- Creating & Iso-Surfacing SDFs
- Parametric Surfaces
Geometric synthesis: Geometry is not limited to record the state of our physical world, it can also be used to generate new geometrically interesting ways for human consumption.
Radiances: Light is fast, is everywhere, is directional, etc. It transmits information very well and how most of us perceive the world, the reality, before we interact.
- Real-Time Rendering
- Light Transport
- Shading, Relighting, and Volumes
- Efficient Sampling: ReSTIR and More
- Walk This Way: Monte Carlo Geometry
Reconstruction: Check my first note.
- Fluid Reconstruction & Optimization
- Dynamic Scenes
- i.e., Articulated (Interactable) Scenes
- Reconstruction
- Reconstruction & Sampling
- Differentiable Geometry Processing & Flow Maps
- Inverse Rendering
Every so often, siggraph hosts some paper on terrain modeling.
- Stochastic geomorphological transport for terrain erosion simulation
- This is a link to an actual paper, not a general session
Coding Agents
Currently, GLM-5.2 is the best open-weight coding model out there in-practice (according to internet) but plagued with cost issues, while the quality is being compared to Opus-level even. What does it take to run the model locally?
Officially supported inference platform: SGLang, vLLM, ktransformers, unsloth.
Actually, don’t.
For a GPU, you won’t get your money back even if you use the $200 monthly plan for at least 5 years. More models like GLM-5.2 are using more reasoning and agentic workflows that even makes CPU-based runtime unworkable. Don’t even try dreaming about your next own Opus-level coding agent workstation. GPU Cloud services does not seem too good either, especially when the utility of building your workstation is at play here also.
Something more useful is probably analyzing the cost of various hardware models at different (V)RAM capacity per unit that equates to 4TB (V)RAM for comparison sake.
@ 4TB=4000GB (V)RAM
| Processor | (Video) Memory Bandwith | Features | Power Draw | Very Rough Cost (Amazon) |
|---|---|---|---|---|
| L4 24GB | 575K | |||
| L40 48GB | 500K | |||
| H100 94GB | 1,423K | |||
| A100 80GB | 820K | |||
| A100 40GB | 472K | |||
| A40 48GB | 500K | |||
| A16 4x16GB | 271K | |||
| A2 16GB | 10kW-15kW | 190K | ||
| T4 16GB | 17.5kW | 200K | ||
| V100 32GB | 31kW-33kW | 113K | ||
| V100 16GB | 62.5kW-75kW | 150K | ||
| P100 16GB | 62.5kW-75kW | 68K | ||
| P40 24GB | 41.7kW | 84K | ||
| M60 2x8GB | 50K | |||
| M40 24GB | 50K | |||
| M10 4x8GB | 44K | |||
| K80 2x12GB | 30K | |||
| 6000 BW 96GB | 584K | |||
| 5000 BW 72GB | 612K | |||
| 5000 BW 48GB | 625K | |||
| 4000 BW 24GB | 417K | |||
| 2000 BW 16GB | 17.5kW | 275K | ||
| 6000 Ada 48GB | 625K | |||
| 5000 Ada 32GB | 575K | |||
| 4500 Ada 24GB | 500K | |||
| 4000 Ada 20GB | 500K | |||
| 2000 Ada 16GB | 17.5kW | 188K | ||
| A6000 48GB | 459K | |||
| A5500 24GB | 500K | |||
| A5000 24GB | 38.4kW | 467K | ||
| A4500 20GB | 40kW | 260K | ||
| A4000 16GB | 25kW | 325K | ||
| Quadro 8000 48GB | 290K | |||
| Quadro 6000 24GB | 250K | |||
| Quadro 5000 16GB | 175K | |||
| Quadro GV100 32GB | 250K | |||
| Quadro GP100 16GB | 175K | |||
| Quadro P6000 24GB | 150K | |||
| Quadro P5000 16GB | 100K | |||
| Quadro M6000 24GB | 100K | |||
| Quadro M6000 12GB | 133K | |||
| 5090 32GB | 562K | |||
| 5080 16GB | 375K | |||
| 5070 Ti 16GB | 75kW | 250K | ||
| 5060 Ti 16GB | 45kW | 150K | ||
| 4090 24GB | 600K | |||
| 4080 (Super) 16GB | 80kW | 250K-375K | ||
| 4070 Ti Super 16GB | 71.3kW | 213K | ||
| 3090 Ti 24GB | 75kW | 267K | ||
| 3090 24GB | 58.4kW | 267K | ||
| Titan RTX 24GB | 217K | |||
| DDR4 64GB | 35K | |||
| DDR5 64GB | 69K | |||
| DDR4 128GB | 38K | |||
| DDR5 128GB | 44K | |||
| DDR4 256GB | 22K-50K | |||
| DDR5 256GB | 78K-141K | |||
| DDR4 512GB | 20K-40K | |||
| DDR5 512GB | 50K-100K | |||
| Radeon VII 16GB | 250K | |||
| Radeon RX 6800 16GB | 113K | |||
| Radeon RX 6800 XT 16GB | 150K | |||
| Radeon RX 6900 XT 16GB | 175K-350K | |||
| Radeon RX 6950 XT 16GB | 275K | |||
| Radeon RX 7600 XT 16GB | 125K | |||
| Radeon RX 7800 XT 16GB | 225K | |||
| Radeon RX 7900 GRE 16GB | 300K | |||
| Radeon RX 7900 XT 20GB | 220K | |||
| Radeon RX 7900 XTX 24GB | 234K | |||
| Radeon RX 9060 XT 16GB | 125K | |||
| Radeon RX 9070 (XT) 16GB | 175K |
Shortlisted version
| Processor | Internal Memory Bandwith (overall bitrate) | Features | Power Draw | Very Rough Cost (Amazon) | W/K$ | Overall bitrate (TB/s) / Power Draw (kW) |
|---|---|---|---|---|---|---|
| A2 16GB | 50TB/s | 10kW-15kW | 190K | 53-79 | 3.3-5 | |
| T4 16GB | 80TB/s | 17.5kW | 200K | 88 | 4.57 | |
| V100 32GB | 112TB/s | 31kW-33kW | 113K | 274-292 | 3.39-3.61 | |
| V100 16GB | 207-225TB/s | 62.5kW-75kW | 150K | 417-500 | 2.76-3.6 | |
| P100 16GB | 183TB/s | 62.5kW-75kW | 68K | 919-1102 | 2.44-2.93 | |
| P40 24GB | 57.6TB/s | 41.7kW | 84K | 496 | 1.38 | |
| 2000 BW 16GB | 72TB/s | 17.5kW | 275K | 64 | 4.11 | |
| 2000 Ada 16GB | 56TB/s | 17.5kW | 188K | 93 | 3.2 | |
| A5000 24GB | 74.7TB/s | 38.4kW | 467K | 82 | 1.95 | |
| A4500 20GB | 102TB/s | 40kW | 260K | 154 | 2.55 | |
| A4000 16GB | 96TB/s | 25kW | 325K | 77 | 3.84 | |
| 5070 Ti 16GB | 224TB/s | 75kW | 250K | 300 | 2.99 | |
| 5060 Ti 16GB | 112TB/s | 45kW | 150K | 300 | 2.49 | |
| 4080 (Super) 16GB | 179.2-184TB/s | 80kW | 250K-375K | 213-320 | 2.24-2.3 | |
| 4070 Ti Super 16GB | 168TB/s | 71.3kW | 213K | 335 | 2.36 | |
| 3090 Ti 24GB | 168TB/s | 75kW | 267K | 281 | 2.24 | |
| 3090 24GB | 156TB/s | 58.4kW | 267K | 219 | 2.67 |
Short-shortlisted version
Unless we’re renting out an office space, most household has very limited wattage rating (for the whole house.) Even if we have unlimited money, we are still limited by power, which means power-efficient processors (basically GPUs) becomes very important. We want the lowest W/K$ (power efficiency in cost) and highest (TB/s)/kW (bitrate efficiency over power).
| Processor | Internal Memory Bandwith (overall bitrate) | Features | Power Draw | Very Rough Cost (Amazon) | W/K$ | Overall bitrate (TB/s) / Power Draw (kW) |
|---|---|---|---|---|---|---|
| A2 16GB | 50TB/s | 10kW-15kW | 190K | 53-79 | 3.3-5 | |
| T4 16GB | 80TB/s | 17.5kW | 200K | 88 | 4.57 | |
| 2000 BW 16GB | 72TB/s | 17.5kW | 275K | 64 | 4.11 | |
| 2000 Ada 16GB | 56TB/s | 17.5kW | 188K | 93 | 3.2 | |
| A4000 16GB | 96TB/s | 25kW | 325K | 77 | 3.84 |
As expected, with prices as of today, workstation and datacenter cards are power efficient in investment1 and higher bitrate for the same power. Interestingly, they are all usually 16GB cards, but the prices are volatile so it could change also.
1 Per cost as a proxy to the quality of the processor. In itself, it’s useless or sometimes bad (if it get’s suspicously low,) but useful when comparing it relatively against other processors.
Unfortunately, running GLM-5.2 requires us to rent an external office/commercial space; it’s not even possible to do it in a standard household. Let’s consider the x16 of A2 16GB, around ~1kW that is well under the hard power limit in standard household wiring, we would get around 256GB VRAM. If we follow the standard rule of 2:1 ratio of system to video RAM, we could get up 97.5% accuracy to the original GLM-5.2 (FP8) via quantization (per the inference provider unsloth.) There’s the MOE, context activation, agentic workflows, etc. that makes the numbers muddier, so let’s just say x16 A2 16GB GPUs. How would one even build a homelab homeserver of this quantity?
Designing this theoretical homeserver will be left as future work.
Future Work
Another interesting aspect regarding coding agents are people are beginning to explore other areas where a local GPU can save costs. Instead of re-hosting LLMs on slow consumer GPUs with even the most quantized/lobotomized models, reranker from RAG are being explored on top of vector embedding similarity score that has been done many times with OK quality.