Membangun Local API Proxy dengan Token Bucket untuk Menangani Rate Limit

Membangun Local API Proxy dengan Token Bucket untuk Menangani Rate Limit

Banyak layanan API menerapkan rate limit untuk melindungi server dari beban berlebih. Ketika aplikasi melebihi batas, respons 429 Too Many Requests muncul dan proses terhenti.

Solusi klasik adalah retry dengan exponential backoff, namun pendekatan itu tetap membiarkan burst request sampai ke upstream. Pendekatan yang lebih baik: local proxy yang menampung request, mengatur keluar secara ritmis, dan menyimpan cache respons.

Artikel ini menunjukkan arsitektur dan implementasi minimal menggunakan Python (FastAPI + httpx).

Arsitektur Proxy Lokal

<svg width="500" height="250" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 500 250">
  <!-- Client -->
  <rect x="20" y="80" width="100" height="50" rx="5" fill="#2563eb"/>
  <text x="70" y="110" text-anchor="middle" fill="white" font-size="14" font-family="system-ui, sans-serif">Client</text>
  <!-- Arrow 1 -->
  <path d="M120 105 L180 105" stroke="#475569" stroke-width="2" marker-end="url(#arrow)"/>
  <!-- Local Proxy -->
  <rect x="180" y="40" width="140" height="170" rx="8" fill="#fef3c7" stroke="#f59e0b" stroke-width="2"/>
  <text x="250" y="65" text-anchor="middle" fill="#92400e" font-size="14" font-weight="600" font-family="system-ui, sans-serif">Local Proxy</text>
  <text x="250" y="85" text-anchor="middle" fill="#92400e" font-size="11" font-family="system-ui, sans-serif">Token Bucket</text>
  <text x="250" y="105" text-anchor="middle" fill="#92400e" font-size="11" font-family="system-ui, sans-serif">Request Queue</text>
  <text x="250" y="125" text-anchor="middle" fill="#92400e" font-size="11" font-family="system-ui, sans-serif">Retry Manager</text>
  <text x="250" y="145" text-anchor="middle" fill="#92400e" font-size="11" font-family="system-ui, sans-serif">Cache Store</text>
  <!-- Arrow 2 -->
  <path d="M320 105 L380 105" stroke="#475569" stroke-width="2" marker-end="url(#arrow)"/>
  <!-- Upstream API -->
  <rect x="380" y="80" width="100" height="50" rx="5" fill="#16a34a"/>
  <text x="430" y="110" text-anchor="middle" fill="white" font-size="14" font-family="system-ui, sans-serif">Upstream API</text>
  <!-- Arrow definition -->
  <defs>
    <marker id="arrow" markerWidth="10" markerHeight="10" refX="8" refY="3" orient="auto" markerUnits="strokeWidth">
      <path d="M0,0 L0,6 L9,3 z" fill="#475569"/>
    </marker>
  </defs>
</svg>

Penjelasan Alur

  1. Client mengirim request ke Local Proxy (bukan langsung ke upstream).
  2. Proxy memeriksa Token Bucket: apakah token tersedia? Jika ya, request dilepaskan ke upstream; jika tidak, request masuk Request Queue.
  3. Retry Manager menangani respons 429 dari upstream dengan membaca header Retry-After dan menambahkan jitter.
  4. Cache Store menyimpan respons GET yang berhasil; request identik berikutnya disajikan dari cache tanpa menghitung ke bucket.

Implementasi Minimal (FastAPI + httpx)

# proxy.py
import asyncio
import time
import hashlib
from collections import deque
from fastapi import FastAPI, Request, Response
import httpx

app = FastAPI()
UPSTREAM = "https://api.example.com"
RATE = 10          # request per second
BURST = 20         # bucket capacity
CACHE_TTL = 300    # seconds


class TokenBucket:
    def __init__(self, rate: float, burst: int):
        self.rate = rate
        self.burst = burst
        self.tokens = float(burst)
        self.last = time.monotonic()
        self.lock = asyncio.Lock()

    async def take(self) -> bool:
        async with self.lock:
            now = time.monotonic()
            self.tokens = min(self.burst, self.tokens + (now - self.last) * self.rate)
            self.last = now
            if self.tokens >= 1:
                self.tokens -= 1
                return True
            return False


bucket = TokenBucket(RATE, BURST)
cache: dict[str, tuple[bytes, float]] = {}


async def forward(req: Request) -> Response:
    key = f"{req.method}:{req.url.path}:{req.query_params}"

    # Cache hit untuk GET
    if req.method == "GET" and key in cache:
        cached, ts = cache[key]
        if time.time() - ts < CACHE_TTL:
            return Response(content=cached, media_type="application/json")

    async with httpx.AsyncClient(base_url=UPSTREAM, timeout=10) as client:
        while True:
            if await bucket.take():
                break
            await asyncio.sleep(0.1)

        upstream_req = client.build_request(
            req.method,
            req.url.path,
            params=req.query_params,
            headers=dict(req.headers),
            content=await req.body()
        )
        resp = await client.send(upstream_req)

        if resp.status_code == 429:
            retry = int(resp.headers.get("Retry-After", "1"))
            await asyncio.sleep(retry + 0.1 * (hash(key) % 10))
            continue

        if req.method == "GET" and resp.status_code == 200:
            cache[key] = (resp.content, time.time())

        return Response(
            content=resp.content,
            status_code=resp.status_code,
            media_type=resp.headers.get("content-type")
        )


@app.api_route("/{path:path}", methods=["GET", "POST", "PUT", "DELETE"])
async def proxy(path: str, request: Request):
    return await forward(request)


if __name__ == "__main__":
    import uvicorn
    uvicorn.run(app, host="0.0.0.0", port=8000)

Menjalankan Proxy

pip install fastapi httpx uvicorn
uvicorn proxy:app --port 8000

Sekarang arahkan client ke http://localhost:8000/... sebagai pengganti upstream asli.

Checklist Produksi

  • [ ] Tambahkan metrics (Prometheus) untuk mengamati bucket & queue
  • [ ] Implementasikan circuit breaker jika upstream terus gagal
  • [ ] Gunakan Redis untuk cache & bucket terdistribusi
  • [ ] Tambahkan authentication (API key) pada proxy
  • [ ] Tulis unit & integration test untuk retry logic

Kesimpulan

Dengan local proxy token-bucket, aplikasi tidak lagi perlu khawatir soal 429. Proxy menyerap burst, mengatur throughput, dan mengurangi beban upstream lewat caching. Pola ini cocok untuk microservices, scraper, dan integrasi pihak ketiga.

Artikel ini dipublikasikan otomatis melalui pipeline Writer → Designer → Reviewer → Publisher.

Story originally reported by Dev.to. View at Dev.to →
← Back to all news