安装方式
命令行安装
在项目根目录执行以下命令,完成 Skill 安装。
npx bzskills add luongnv89/skills --skill ollama-optimizer Optimize Ollama configuration for the current machine's hardware. Use when asked to speed up Ollama, tune local LLM performance, or pick models that fit available GPU/RAM. Don't use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers.
5
下载量
命令行安装
在项目根目录执行以下命令,完成 Skill 安装。
npx bzskills add luongnv89/skills --skill ollama-optimizer name: ollama-optimizer
description: Optimize Ollama configuration for the current machine's hardware. Use when asked to speed up Ollama, tune local LLM performance, or pick models that fit available GPU/RAM. Don't use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers.
license: MIT
effort: medium
metadata:
version: 1.2.0
author: "Luong NGUYEN <luongnv89@gmail.com>"Optimize Ollama configuration based on system hardware analysis.
Use this skill when the user asks to optimize Ollama, configure Ollama, speed up Ollama, fix Ollama running slow, set up a local LLM, tune inference speed, reduce memory usage, or select models that fit their GPU/RAM. The skill analyzes hardware (GPU, VRAM, RAM, CPU) and produces tailored recommendations.
Do not use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers (OpenAI, Anthropic) — those use different runtimes and tuning surfaces.
Fast path (opt-in only): only skip full hardware analysis if the user explicitly asks to. Otherwise always run Phases 1-4 and follow the tier-based recommendation — do not apply shortcuts by default, and do not let them override a tier decision already made. For the per-platform shortcut commands and env vars, see [Platform-Specific Setup](references/platform_specific.md) and [Environment Variables](references/environment_variables.md).
Run the detection script to gather hardware information:
python3 scripts/detect_system.py
Parse the JSON output to identify:
hardware_tier — the script's computed category, max_model_size, and recommended_quantUse hardware_tier from Phase 1 as the tier decision. Do not re-derive it; the table below explains what each tier means and which optimizations it implies. Override the script only with an explicit reason (e.g. VRAM shared with a display), and state that reason in the report.
Hardware Tier Classification:
Tier (category) | Script band | Max Model | Key Optimizations |
|---|---|---|---|
cpu_only | No GPU detected | 3B | num_thread tuning, Q4_K_M quant |
low_vram | <6GB VRAM | 3B | Flash attention, KV cache q4_0 |
entry | 6-10GB VRAM | 8B | Flash attention, KV cache q8_0 |
prosumer | 10-16GB VRAM | 14B | Flash attention, full offload |
workstation | 16-48GB VRAM | 32B | Standard config, Q5_K_M option |
high_end | 48GB+ VRAM | 70B+ | Multiple models, Q5/Q6 quants |
Apple Silicon Special Case:
entryprosumerworkstation; 64GB+ Mac → high_endCreate a structured optimization guide with these sections:
#### 1. System Overview
Present detected hardware specs and highlight constraints (e.g., "8GB unified memory limits to 8B models").
#### 2. Dependency Assessment
List what's needed based on the platform:
#### 3. Configuration Recommendations
Essential environment variables:
# Always recommended
export OLLAMA_FLASH_ATTENTION=1
# Memory-constrained systems (<12GB)
export OLLAMA_KV_CACHE_TYPE=q8_0 # or q4_0 for severe constraints
Model selection guidance:
ollama list outputModelfile tuning (when needed):
PARAMETER num_gpu <layers> # Partial offload for limited VRAM
PARAMETER num_thread <cores> # CPU threads (physical cores, not hyperthreads)
PARAMETER num_ctx <size> # Reduce context for memory savings
#### 4. Execution Checklist
Provide copy-paste commands in order:
$SHELL decides: ~/.zshrc, ~/.bashrc, or ~/.bash_profile) and append the env vars: RC=~/.zshrc # or ~/.bashrc / ~/.bash_profile, matching $SHELL
cp "$RC" "$RC.ollama-bak"
printf '\n# ollama-optimizer start\nexport OLLAMA_FLASH_ATTENTION=1\n<KV cache + other export lines from section 3, per tier>\n# ollama-optimizer end\n' >> "$RC"
ollama run <model> --verbosecp ~/.zshrc.ollama-bak ~/.zshrc — then restart Ollama.# Benchmark current performance
python3 scripts/benchmark_ollama.py --model <model>
# Expected output: tokens/s and generation latency — record as the post-tuning baseline.
# Check GPU memory usage (NVIDIA)
nvidia-smi
# Verify config is applied
ollama run <model> "test" --verbose 2>&1 | head -20
A run passes when all of the following are true:
OLLAMA_FLASH_ATTENTION, KV-cache quantisation) are written to a shell init file the user actually uses, with a backup of the prior file.ollama run <model> with --verbose and captures the actual offload/cache numbers.After completing each major step, output a status report in this format:
◆ [Step Name] ([step N of M] — [context])
··································································
[Check 1]: √ pass
[Check 2]: √ pass (note if relevant)
[Check 3]: × fail — [reason]
[Check 4]: √ pass
[Criteria]: √ N/M met
____________________________
Result: PASS | FAIL | PARTIAL
Adapt the check names to match what the step actually validates. Use √ for pass, × for fail, and — to add brief context. The "Criteria" line summarizes how many acceptance criteria were met. The "Result" line gives the overall verdict.
◆ Detection (step 1 of 4 — hardware profiling)
··································································
Hardware detected: √ pass — macOS 14, Apple M2
GPU identified: √ pass — Apple Metal (unified memory)
RAM measured: √ pass — 16GB unified memory
[Criteria]: √ 3/3 met
____________________________
Result: PASS
◆ Analysis (step 2 of 4 — profile selection)
··································································
Tier classified: √ pass — Prosumer (16GB unified)
Profile selected: √ pass — Flash attention, full offload
Bottlenecks identified: √ pass — memory bandwidth primary constraint
[Criteria]: √ 3/3 met
____________________________
Result: PASS
◆ Plan (step 3 of 4 — optimization guide)
··································································
Guide generated: √ pass — ollama-optimization-guide.md written
Parameters tuned: √ pass — OLLAMA_FLASH_ATTENTION=1, KV_CACHE_TYPE=q8_0
Model recommendations ready: √ pass — llama3.1:14b-instruct-q4_K_M suggested
[Criteria]: √ 3/3 met
____________________________
Result: PASS
◆ Verification (step 4 of 4 — config validation)
··································································
Benchmark commands listed: √ pass — python3 scripts/benchmark_ollama.py
Config verified: √ pass — ollama run --verbose output checked
[Criteria]: √ 2/2 met
____________________________
Result: PASS
Generate an ollama-optimization-guide.md file. Ask the user where to save it (suggest ~/.config/ollama/optimization-guide.md or current directory). Contents:
# Ollama Optimization Guide
**Generated:** <timestamp>
**System:** <OS> | <CPU> | <RAM>GB RAM | <GPU>
## System Overview
<hardware summary and constraints>
## Current Configuration
<existing Ollama setup and env vars>
## Recommendations
### Environment Variables
<shell commands to set vars>
### Model Selection
<recommended models with rationale>
### Performance Tuning
<Modelfile adjustments if needed>
## Execution Checklist
- [ ] <step 1>
- [ ] <step 2>
...
## Verification
<benchmark commands and expected results>
## Rollback
<commands to revert changes if needed>