Opens the change feed from get_table_changes() in an interactive
viewer. A sidebar lists each part of the diff with a count: the schema
changes, the columns and cells that changed, the rows inserted, updated,
and deleted, and the snapshots in range. The main panel shows the
selected part as a table. An update is one row, shown from its old or its
new side, with the changed cells highlighted and both values on hover; a
cells view lists every change as snapshot, rowid, column, old, new.
Usage
view_table_changes(
changes,
max_rows = 10000,
table_name = NULL,
ducklake_name = NULL,
conn = NULL,
width = NULL,
height = NULL,
elementId = NULL
)Arguments
- changes
A change feed from
get_table_changes(): the lazy table it returns, possibly narrowed with dplyr verbs (the filters run in DuckDB before anything is collected), or that table collected into a data frame. The columnssnapshot_id,rowid, andchange_typemust still be present.- max_rows
The most feed rows to collect. A longer feed is cut at a snapshot and row boundary, never between the two images of one update, and the viewer says how many rows it holds out of how many there are. A dplyr filter on the feed narrows it.
- table_name
The table the feed describes, as
"table"or"schema.table". A lazy feed carries it; a collected data frame does not, and without it the viewer shows no snapshot authors and messages and no schema panel.- ducklake_name
Optional name of the attached DuckLake catalog. A lazy feed carries it; otherwise the current database is used.
- conn
Optional DuckDB connection object. A lazy feed carries it; otherwise the default ducklake connection is used.
- width, height
The widget's size, in CSS units or pixels.
NULLfills the viewer.- elementId
An id for the widget's HTML element.
Value
An htmlwidget of class ducklake_changes, which prints in the
RStudio Viewer, a browser, or an HTML document.
Details
Requires the htmlwidgets package (listed in Suggests).
The layout follows Hadley Wickham's
data-diff, a command-line tool
with a browser interface for comparing Parquet files, with thanks.
DuckLake's change feed makes the job simpler than comparing two files:
every row carries its identity (rowid) and the kind of change, so the
viewer pairs an update's two images by snapshot and rowid and compares
them cell by cell, with no key guessing.
A dplyr filter on a data column can keep one image of an update and drop
the other: filter(status == "closed") keeps the image in which the
status is closed. Such a row is shown from the side that is present,
marked as one-sided, with no changed cells. Filters on rowid,
snapshot_id, or change_type never split a pair. DuckLake records an
update even when the new values equal the old, and that shows as an
update with no changed cells.
The schema panel reads the column history from the lake's catalog for
the table's current id, so a rebuild by replace_table() or
restore_table_version() inside the range appears as the table's
creation. Values are compared before they are formatted for display:
numbers show 15 significant digits, timestamps are in UTC, and a missing
value shows as NA. NA and NaN compare equal.
Examples
lake_dir <- tempfile("view_lake_")
dir.create(lake_dir)
attach_ducklake("view_lake", lake_path = lake_dir)
create_table(data.frame(id = 1:3, amount = c(10, 20, 30)), "orders")
rows_update(
get_ducklake_table("orders"),
data.frame(id = 2L, amount = 25),
by = "id"
)
rows_delete(get_ducklake_table("orders"), data.frame(id = 3L), by = "id")
# The table's whole history
view_table_changes(get_table_changes("orders"))
# Narrow the feed first; the filter runs in DuckDB
get_table_changes("orders") |>
dplyr::filter(change_type != "insert") |>
view_table_changes()
detach_ducklake("view_lake", shutdown = TRUE)
unlink(lake_dir, recursive = TRUE)
