Create a DuckLake table
Arguments
- data_source
Raw data source. Can be:
A URL (http:// or https://)
A file path (e.g., "data.csv", "data.parquet")
An R data.frame or tibble
A lazy table (tbl_duckdb_connection or tbl_lazy). A lazy table on the package's connection, such as a dplyr pipeline built on
get_ducklake_table(), is written withCREATE TABLE ... ASinside DuckDB, so its rows never pass through R. A lazy table on another connection is collected first.
- table_name
Name of the new table
- labels
When
TRUE(the default), store variable labels in the lake as column comments, in the same transaction as the table creation, so both land as one snapshot. For a data frame the labels are its haven/labelledlabelattributes; for a lazy table, each output column keeps the comment of the same-named column in the tables the query reads, so labels follow the data through a pipeline (a renamed or derived column starts without one). Collecting the table later restores the labels (seeget_table_comments()), and every other client of the lake can read them too. Set toFALSEto skip.
Examples
lake_dir <- tempfile("create_lake_")
dir.create(lake_dir)
attach_ducklake("create_lake", lake_path = lake_dir)
# From data.frame
create_table(mtcars, "cars")
# From a local file
csv_path <- tempfile(fileext = ".csv")
utils::write.csv(mtcars, csv_path, row.names = FALSE)
create_table(csv_path, "cars_from_csv")
# From a lazy table: the query runs inside DuckDB and writes straight
# into the lake, without collecting into R
get_ducklake_table("cars") |>
dplyr::filter(cyl > 4) |>
create_table("big_cars")
# From a URL -- needs network access and the httpfs extension
if (FALSE) { # \dontrun{
create_table("https://example.com/data.csv", "remote_table")
} # }
unlink(csv_path)
detach_ducklake("create_lake", shutdown = TRUE)
unlink(lake_dir, recursive = TRUE)
