Declares how newly written data files for a table should be split up. Partitioning lets DuckLake prune files during query planning, which can speed up filtered reads on large tables considerably.
Arguments
- table_name
The name of the table to partition.
- partition_by
Character vector of partition expressions. Each entry must be one of:
a column name, e.g.
"region"(identity transform)"year(col)","month(col)","day(col)", or"hour(col)"for timestamp columns"bucket(n, col)"for hash bucketing intonbuckets
Details
Partitioning only affects data written after the keys are set;
previously written files keep their layout. To re-partition existing
data, set the keys and then rewrite the table with replace_table():
the rewrite carries the keys over to the new table and writes every row
through them.
Runs ALTER TABLE ... SET PARTITIONED BY (...). The expressions are
validated against the transforms DuckLake supports before any SQL is
built.
See also
reset_table_partitioning(), get_table_partitions()
Other partitioning:
get_table_partitions(),
reset_table_partitioning()
Examples
lake_dir <- tempfile("part_lake_")
dir.create(lake_dir)
attach_ducklake("part_lake", lake_path = lake_dir)
create_table(mtcars, "cars")
# Plain column partitioning
set_table_partitioning("cars", "cyl")
#> Table "cars" is now partitioned by "cyl".
#> ℹ Only newly written data is partitioned; existing files keep their layout.
# Compound key
set_table_partitioning("cars", c("gear", "cyl"))
#> Table "cars" is now partitioned by "gear" and "cyl".
#> ℹ Only newly written data is partitioned; existing files keep their layout.
# Timestamp columns can be split by year(), month(), day(), or hour()
detach_ducklake("part_lake", shutdown = TRUE)
unlink(lake_dir, recursive = TRUE)
