DevQuasar

community

Verified

https://devquasar.com/

Activity Feed

AI & ML interests

Open-Source LLMs, Local AI Projects: https://pypi.org/project/llm-predictive-router/

Recent Activity

csabakecskemeti updated a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-350m-GGUF

csabakecskemeti published a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-350m-GGUF

csabakecskemeti updated a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-1b-GGUF

View all activity

csabakecskemeti

updated a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-350m-GGUF

Text Generation • 0.4B • Updated about 4 hours ago

csabakecskemeti

published a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-350m-GGUF

Text Generation • 0.4B • Updated about 4 hours ago

csabakecskemeti

updated a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-1b-GGUF

Text Generation • 2B • Updated about 4 hours ago

csabakecskemeti

published a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-1b-GGUF

Text Generation • 2B • Updated about 4 hours ago

csabakecskemeti

updated a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-h-1b-GGUF

Text Generation • 1B • Updated about 4 hours ago

csabakecskemeti

published a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-h-1b-GGUF

Text Generation • 1B • Updated about 4 hours ago

csabakecskemeti

updated a model about 4 hours ago

DevQuasar/ibm-granite.granite-4.0-h-350m-GGUF

Text Generation • 0.3B • Updated about 4 hours ago • 21

csabakecskemeti

published a model about 5 hours ago

DevQuasar/ibm-granite.granite-4.0-h-350m-GGUF

Text Generation • 0.3B • Updated about 4 hours ago • 21

csabakecskemeti

updated a model about 17 hours ago

DevQuasar/cerebras.GLM-4.5-Air-REAP-82B-A12B-GGUF

Text Generation • 85B • Updated about 17 hours ago • 131

csabakecskemeti

posted an update 12 days ago

Post

2490

Christmas came early this year

3 replies

csabakecskemeti

posted an update 5 months ago

Post

3032

Has anyone ever backed up a model to a sequential tape drive, or I'm the world first? :D
Just played around with my retro PC that has got a tape drive—did it just because I can.

5 replies

csabakecskemeti

posted an update 5 months ago

Post

464

Deepseek R1 0528 Q2 locally.
(I believe it has overthinking it a bit :) )
https://youtu.be/Iqu5s9aFaXA?si=QWZe293iTKf_3ELU

DevQuasar/deepseek-ai.DeepSeek-R1-0528-GGUF

csabakecskemeti

posted an update 7 months ago

Post

2110

Local Llama4 Maverick Q2
https://youtu.be/4F8g_LThli0?si=MGba2SUTHt6xYw3T
Quants uploading now

Big thanks to @ngxson !

csabakecskemeti

posted an update 7 months ago

Post

1744

Why the 'how many r's in strawberry' prompt "breaks" llama4? :D

Quants DevQuasar/meta-llama.Llama-4-Scout-17B-16E-Instruct-GGUF

3 replies

csabakecskemeti

posted an update 7 months ago

Post

3407

I'm collecting llama-bench results for inference with a llama 3.1 8B q4 and q8 reference models on varoius GPUs. The results are average of 5 executions.
The system varies (different motherboard and CPU ... but that probably that has little effect on the inference performance).

https://devquasar.com/gpu-gguf-inference-comparison/
the exact models user are in the page

I'd welcome results from other GPUs is you have access do anything else you've need in the post. Hopefully this is useful information everyone.

csabakecskemeti

posted an update 7 months ago

Post

2408

Managed to get my hands on a 5090FE, it's beefy

| llama 8B Q8_0 | 7.95 GiB | 8.03 B | CUDA | 99 | pp512 | 12207.44 ± 481.67 |
| llama 8B Q8_0 | 7.95 GiB | 8.03 B | CUDA | 99 | tg128 | 143.18 ± 0.18 |

Comparison with others GPUs
http://devquasar.com/gpu-gguf-inference-comparison/

csabakecskemeti

posted an update 7 months ago

Post

1847

GTC new model announcement now from Nvidia
nvidia/Llama-3_3-Nemotron-Super-49B-v1

GGUFs:
DevQuasar/nvidia.Llama-3_3-Nemotron-Super-49B-v1-GGUF

Enjoy!

csabakecskemeti

posted an update 8 months ago

Post

593

Cohere Command-a Q2 quant
DevQuasar/CohereForAI.c4ai-command-a-03-2025-GGUF

6.7t/s on a 3gpu setup (4080 + 2x3090)

(q3, q4 currently uploading)

csabakecskemeti

posted an update 8 months ago

Post

843

Fine tuning on the edge. Pushing the MI100 to it's limits.
QWQ-32B 4bit QLORA fine tuning
VRAM usage 31.498G/31.984G :D

4 replies

csabakecskemeti

posted an update 8 months ago

Post

1998

-UPDATED-
4bit inference is working! The blogpost is updated with code snippet and requirements.txt
https://devquasar.com/uncategorized/all-about-amd-and-rocm/
-UPDATED-
I've played around with an MI100 and ROCm and collected my experience in a blogpost:
https://devquasar.com/uncategorized/all-about-amd-and-rocm/
Unfortunately I've could not make inference or training work with model loaded in 8bit or use BnB, but did everything else and documented my findings.

4 replies

AI & ML interests

Recent Activity

Team members 1

DevQuasar's activity