/* git log --author=dwsmith1983 --merged */

Featured

Apache DataFusion Comet 17 merged

#6297

key each scan's planning data by its plan node in a native block

fix 2026-09-30
#6280

check Iceberg DPP pruning by planned file tasks

test 2026-09-30
#6116

reject a file without field ids at any depth whether or not id matching is on

matches Spark's missing field-id check at any depth
fix 2026-09-27
#5880

report task input metrics after the native iterator closes and add to Spark's counters

fix 2026-09-27
#5654

match Spark's duplicate field and field id semantics in parquet field lookup

matches Spark's field-id semantics for nested fields
fix 2026-09-25
#6041

let decimal SUM recover from an intermediate overflow like Spark

matches Spark when an intermediate sum overflows
fix 2026-09-22
#6042

gate the regr_r2 degenerate-case swap on the Spark patch release

fix 2026-09-19
#6040

guard page skipping in the native Iceberg scan

test 2026-09-19
#5874

run length on binary input natively

feat 2026-09-19
#5873

drop the redundant width_bucket shim registrations and guard serde uniqueness

chore 2026-09-19
#5615

cache parsed plan data across a stage's tasks

~5x lower per-task planning cost
perf 2026-09-16
#5565

reuse zstd compression contexts across shuffle blocks

~6% faster shuffle encode
perf 2026-09-16
#5866

decline structs with duplicate field names before they reach Java Arrow

fix 2026-09-13
#5612

compile user regex patterns once per planned expression

2.1x faster on small batches
perf 2026-09-13
#5762

label pull requests by changed paths and title prefix

ci 2026-09-08
#5653

expand object store option references, uniquify constant metadata names, drop dead parquet JNI

306 lines of dead JNI removed
fix 2026-09-04
#5568

reuse per-partition scratch in the shuffle write path

27% faster writes at 2,000 partitions
perf 2026-09-04

Apache DataFusion 1 merged

Apache Spark 2 merged

dataconf 5 merged

Libraries

Reports