Compacts small adjacent Parquet files into larger ones. Frequent small inserts each write their own file; merging keeps file counts down and scans fast.
Usage
merge_adjacent_files(
ducklake_name = NULL,
table_name = NULL,
schema_name = NULL,
max_compacted_files = NULL,
min_file_size = NULL,
max_file_size = NULL
)Arguments
- ducklake_name
Name of the attached DuckLake catalog. If
NULL, the current database is used.- table_name
Optional table name, optionally qualified as
"schema.table". When provided, only that table is compacted.- schema_name
Optional schema name. When provided, only tables in that schema are compacted.
- max_compacted_files
Optional cap on the number of compaction operations per table in a single call.
- min_file_size
Optional minimum file size in bytes; smaller files are excluded from merging.
- max_file_size
Optional maximum file size in bytes; files at or above this size are excluded. Defaults to the lake's target file size.
Value
A data frame with one row per output file (columns
schema_name, table_name, files_processed, files_created).
Details
Merging does not delete the original small files – they may still be
referenced by older snapshots. They are scheduled for deletion once no
snapshot references them; run cleanup_old_files() to remove them.
Examples
lake_dir <- tempfile("merge_files_lake_")
dir.create(lake_dir)
attach_ducklake("merge_files_lake", lake_path = lake_dir)
create_table(mtcars, "cars")
rows_insert(
get_ducklake_table("cars"),
data.frame(mpg = 30, cyl = 4),
by = "mpg"
)
# Compact the whole lake
merge_adjacent_files()
#> No adjacent files to merge.
#> [1] schema_name table_name files_processed files_created
#> <0 rows> (or 0-length row.names)
# Compact one table, only touching files under 10 MB
merge_adjacent_files(table_name = "cars", max_file_size = 10e6)
#> No adjacent files to merge.
#> [1] schema_name table_name files_processed files_created
#> <0 rows> (or 0-length row.names)
detach_ducklake("merge_files_lake", shutdown = TRUE)
unlink(lake_dir, recursive = TRUE)
