latest update:
[apexdb]

Integration guides · six common stacks

The snapshot ships in four formats (JSON, CSV, SQLite, Parquet), compatible with most analytical stacks. Examples assume the variants table lives at ./variants.parquet and the combined SQLite DB at ./apex.sqlite.

Parquet + DuckDB

DuckDB is the fastest path to interactive queries over the full dataset. No load step, no schema declaration.

duckdbsql
-- Open the snapshot and query directly
SELECT brand, model, model_year, power_hp, co2_combined_g_km
FROM 'variants.parquet'
WHERE powertrain IN ('ICE', 'HEV', 'PHEV')
  AND model_year >= 2020
ORDER BY co2_combined_g_km
LIMIT 20;

Parquet + polars (Python)

python
import polars as pl

df = pl.read_parquet("variants.parquet")
print(df.shape)              # (43077, ...)  # one row per variant
print(df["powertrain"].value_counts())

Parquet + pandas / pyarrow

python
import pyarrow.parquet as pq

table = pq.read_table("variants.parquet")
df = table.to_pandas()
# Lazy column access. Avoid full materialization when you only need a few cols:
df = pq.read_table("variants.parquet",
                   columns=["brand", "model", "model_year", "co2_combined_g_km"]).to_pandas()

CSV + Excel / Power BI

UTF-8 CSV with a header row and comma delimiter. Power BI and Excel auto-detect both. For BI tools, Parquet via the ODBC connector is preferred; CSV is ~10× larger and slower to refresh.

SQLite

Zero-dependency embedded use. The SQLite file is the same row set, one table per shipped CSV (variants, provenance, sources, variants_pcaf, coverage_by_pcaf_co2, generation_ratings).

sh
sqlite3 apex.sqlite \
  "SELECT brand, model, mpg_combined_us
   FROM variants
   WHERE powertrain = 'HEV' AND model_year = 2024
   LIMIT 5"

REST API · Python

The REST API is live at https://api.apex-db.org/v1. The examples below work today.

python
import httpx

client = httpx.Client(
    base_url="https://api.apex-db.org/v1",
    headers={"Authorization": f"Bearer {APEX_TOKEN}"},
    timeout=10.0,
)

r = client.get("/apex", params={"brand": "toyota", "model": "rav4-hybrid"})
r.raise_for_status()
for v in r.json()["data"]:
    print(v["id"], v["variant_name"], v["mpg_combined_us"])

REST API · TypeScript

Available alongside the Python client above.

ts
const res = await fetch(
  "https://api.apex-db.org/v1/apex?brand=bmw&model=5%20Series&model_year=2024",
  { headers: { Authorization: `Bearer ${APEX_TOKEN}` } },
);
if (!res.ok) throw new Error(`apex ${res.status}`);
const { data } = (await res.json()) as { data: Variant[] };

R + arrow

r
library(arrow)
library(dplyr)

ds <- open_dataset("variants.parquet")
ds |>
  filter(powertrain == "BEV", model_year >= 2022) |>
  select(brand, model, battery_capacity_gross_kwh, range_wltp_km) |>
  collect()

Other stacks

Stack requests to data@apex-db.org receive a working example by reply. Common requests are published on this page.