Threads vs Processes for PDAL Workloads
TL;DR: Use a ProcessPoolExecutor (or separate pdal pipeline subprocesses) for tile-level parallelism — one tile per worker, OMP_NUM_THREADS=1 in each — because processes isolate memory, crashes and global library state. Reach for threads only for I/O-bound orchestration such as uploads, and measure: the right answer depends on your stages, your storage and your core count.
# Context and Motivation
This guide is part of Parallel Execution in PDAL. When a batch of tiles is too slow on one core, Python offers three levers: threads within one process, a pool of processes, and the threads that some PDAL stages start themselves through OpenMP. They interact, and choosing badly either wastes cores or oversubscribes them until the machine thrashes.
The usual intuition from pure-Python work — “threads are useless because of the GIL” — is only half true for PDAL. The heavy work happens in C++, and whether other Python threads can run during it depends on whether the bindings release the GIL during execution, which has varied between binding versions. More importantly, even when threads can run in parallel, they share one process: one tile’s out-of-memory failure kills every tile in flight, and global state in GDAL and PROJ is shared. Processes avoid both problems at a modest start-up cost.
# Prerequisites and Assumptions
- PDAL 2.x with Python bindings, on a multi-core Linux machine or VM.
- A batch of independent tiles — the usual case for LiDAR production.
- Enough memory for N simultaneous tiles, where N is the number of workers; see estimating PDAL memory from point layout.
# Step-by-Step Implementation
# Step 1 — Fix the per-process thread count
Set OMP_NUM_THREADS=1 (and GDAL_NUM_THREADS=1 if writers compress with multiple threads) in the environment of every worker. Otherwise each process starts one OpenMP thread per core and N processes oversubscribe the machine N-fold.
# Step 2 — Write the per-tile function at module level
Process pools pickle the function and its arguments. Define the worker at module top level, take a path and return a small result — never the point arrays.
# Step 3 — Benchmark three configurations
Run the same set of tiles with a thread pool, a process pool, and a single process with OpenMP threads enabled, keeping total cores constant.
# Step 4 — Read the results together with memory
Wall time alone is not the decision; peak memory per configuration decides how many workers fit.
# Step 5 — Choose and document
Record the choice, the worker count and the thread settings in the batch configuration so the next person does not re-derive them.
# Complete Working Example
"""Benchmark thread pool, process pool and OpenMP-only execution for PDAL tiles."""
from __future__ import annotations
import json
import os
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
from pathlib import Path
TILES = sorted(Path("tiles").glob("*.laz"))[:16]
CORES = os.cpu_count() or 8
def pipeline_for(tile: Path) -> str:
return json.dumps({"pipeline": [
str(tile),
{"type": "filters.range", "limits": "Classification![7:7],Classification![18:18]"},
{"type": "filters.smrf", "slope": 0.15, "window": 18, "threshold": 0.5},
{"type": "filters.range", "limits": "Classification[2:2]"},
{"type": "writers.gdal", "filename": f"out/{tile.stem}_dtm.tif", "resolution": 1.0,
"output_type": "idw", "data_type": "float32"},
]})
def run_tile(tile: str) -> int:
import pdal # import inside the worker: fresh library state per process
return pdal.Pipeline(pipeline_for(Path(tile))).execute()
def init_single_thread() -> None:
os.environ["OMP_NUM_THREADS"] = "1"
os.environ["GDAL_NUM_THREADS"] = "1"
def bench(name: str, fn) -> None:
t0 = time.perf_counter()
fn()
print(f"{name:<28} {time.perf_counter() - t0:7.1f} s")
if __name__ == "__main__":
Path("out").mkdir(exist_ok=True)
tiles = [str(t) for t in TILES]
os.environ["OMP_NUM_THREADS"] = "1"
bench(f"threads x{CORES}", lambda: list(ThreadPoolExecutor(CORES).map(run_tile, tiles)))
bench(f"processes x{CORES}", lambda: list(
ProcessPoolExecutor(CORES, initializer=init_single_thread).map(run_tile, tiles)))
os.environ["OMP_NUM_THREADS"] = str(CORES)
bench(f"serial, OMP x{CORES}", lambda: [run_tile(t) for t in tiles])Illustrative results for 16 tiles of about 20 million points on a 16-core worker:
threads x16 412.6 s
processes x16 188.3 s
serial, OMP x16 905.9 sThe ordering, not the numbers, is the lesson: on this workload processes win comfortably, and relying on in-stage OpenMP alone leaves most cores idle because only some stages parallelize internally.
# Key Parameter Table
| Setting | Recommended | Why |
|---|---|---|
| Pool type for tiles | ProcessPoolExecutor |
Memory and failure isolation, independent library state |
| Workers | min(cores, memory ÷ peak per tile) | Memory is usually the binding limit |
OMP_NUM_THREADS |
1 per process | Avoids N × cores oversubscription |
GDAL_NUM_THREADS |
1 per process | Same, for multithreaded compression in writers |
max_tasks_per_child |
10–50 (Python 3.11+) | Recycles workers to release fragmented memory |
| Thread pools | I/O only | Uploads, downloads, API calls around PDAL work |
# Verification
- CPU utilization.
htopduring a process-pool run should show all cores near 100 percent and load average near the core count. A load average far above it means oversubscription. - Output equality. Rasters from the three configurations should be byte-identical or identical after reading; parallelism must not change results.
- Failure isolation. Deliberately feed one corrupt tile; the process pool should report one failure and finish the rest.
# Gotchas and Edge Cases
Oversubscription is silent. Sixteen processes each starting sixteen OpenMP threads run 256 threads on 16 cores. Everything still works, just slower than serial in some stages, and nothing in the logs says why.
Fork and library state. On Linux, ProcessPoolExecutor forks by default, and a parent that has already initialized GDAL or PDAL passes that state to children. Import pdal inside the worker function, or use the spawn start method, to give each worker a clean start.
Returning arrays is expensive. Results cross process boundaries by pickling. Return counts, paths and small summaries; write point data to files from inside the worker.
Threads for orchestration are fine. A thread pool that uploads finished rasters to S3 while the process pool computes the next tiles is a good use of threads — the work is I/O-bound and releases the GIL.
# Frequently Asked Questions
Does the GIL stop PDAL from running in parallel threads?
PDAL’s work runs in C++, so the GIL matters only for whether other Python threads can run during it, which depends on the bindings. Regardless, threads share one process’s memory and library state, so processes are the more robust choice for tile-level parallelism.
How many worker processes should I use?
The smaller of the core count and the number of tiles that fit in memory at once. Memory is usually the tighter limit for PDAL; measure peak memory for your largest tile and divide available RAM by it.
Should I set OMP_NUM_THREADS to 1?
Yes, when running several PDAL processes at once. Each process would otherwise start one OpenMP thread per core, oversubscribing the machine. Raise it only when running one large tile at a time.
When are threads the right choice?
For I/O-bound work around PDAL: downloading inputs, uploading outputs, polling APIs. Those tasks spend their time waiting, and threads let many waits overlap cheaply.
# Related
- Parallel Execution in PDAL — file-level versus stage-level parallelism
- Parallel Tile Processing with ProcessPoolExecutor — the production pattern
- Load Balancing Uneven LiDAR Tiles — keeping every worker busy
- Optimizing PDAL for Multi-Core Processing — OpenMP and chunk tuning
- Dask vs ProcessPoolExecutor for PDAL — when to go beyond one machine