delta.sharing connects R to tables exposed through Delta Sharing. A typical workflow
is to create a client, select a table, describe a read, and choose the
form of the result.
Try the public example data
The Delta Sharing project hosts an open server that can be used without registering or creating a private credential:
library(delta.sharing)
client <- sharing_client(demo_profile())
housing <- client$table("delta_sharing.default.boston-housing")
housing$snapshot(
columns = c("chas", "medv"),
limit = 5
)$to_tibble()demo_profile() fetches the public profile maintained by
the Delta Sharing project. The same client and table methods work with a
private share.
Connect to your share
Pass the path to a .share profile:
client <- sharing_client("~/config.share")sharing_client() also accepts a parsed profile list. See
?sharing_client for the supported profile fields and
authentication methods.
Find a table
Use the client to discover the data available through its profile:
client$list_shares()
client$list_schemas("sales")
client$list_tables("sales", "default")These methods follow pagination automatically and return printable lists of records. Create a reusable table handle from a listed table:
orders <- client$table("sales.default.orders")
orders
orders$version()
orders$schema()Creating a table handle does not read table rows. Its
protocol() and metadata() methods expose
additional table details.
Read a snapshot
Configure the read with snapshot(), then materialize it.
to_tibble() is the usual choice for R analysis:
snapshot <- orders$snapshot(
columns = c("order_id", "status", "amount"),
limit = 1000
)
orders_tbl <- snapshot$to_tibble()Snapshots can also target a table version or point in time:
orders$snapshot(version = 42)$to_tibble()
orders$snapshot(timestamp = "2024-01-01T00:00:00Z")$to_tibble()Choose the materializer that matches the next consumer:
| Result | Method |
|---|---|
| Tibble for ordinary R analysis | to_tibble() |
| Base data frame | to_data_frame() |
| In-memory Arrow table | to_arrow() |
| Lazy Arrow reader, including for DuckDB | to_arrow_reader() |
| Low-level Arrow C Stream | to_arrow_stream() |
Arrow is a required dependency. The tibble and data-frame methods
automatically convert BIGINT columns to bit64::integer64,
including empty and nested results. A valid -9223372036854775808 raises
a conversion error because bit64 reserves it for NA; Arrow
materializers retain that value. Each materializer call performs a new
read, so reuse the result when the same data is needed more than
once.
columns selects the returned columns and
limit caps the number of returned rows.
predicate accepts nested R lists that are sent to the
sharing server as best-effort hints; predicates are not exact row
filters. See ?SharingTable for the complete snapshot
options and predicate structure.
Read changes
For a table with change data feed enabled, changes()
reads an inclusive version or timestamp range:
changes_tbl <- orders$changes(
starting_version = 120,
ending_version = 125,
columns = c(
"order_id",
"status",
"_change_type",
"_commit_version",
"_commit_timestamp"
)
)$to_tibble()_change_type distinguishes inserts, deletes, and the
before and after rows of updates. The commit columns identify when each
change was recorded. Change reads expose the same materializers as
snapshots.
Downloads and caching
Selected files are downloaded concurrently and cached for the R session. New handles for the same table reuse files already present in the cache, so overlapping reads may avoid downloading them again.
The cache is stored under R’s session temporary directory and normally needs no user management.
See vignette("performance-caching") for cache identity
and lifetime, cold and repeated reads, download concurrency,
materializer costs, and Arrow batching.
Next steps
See ?SharingTable, ?SharingSnapshot, and
?SharingChanges for the complete read options. The README
shows how to query a shared table with DuckDB without first creating an
R data frame.