How to get
This section describes how to build the Lakehouse source connector. You can get the Lakehouse source connector using one of the following methods:- Download the NAR package from the download page.
- Build it from the source code.
-
Clone the source code to your machine.
-
Build the connector in the
pulsar-io-lakehousedirectory.-
Build the NAR package for your local file system.
-
Build the NAR package for your cloud storage (Including AWS, GCS, and Azure-related package dependency).
-
Build the NAR package for your local file system.
How to configure
Before using the Lakehouse source connector, you need to configure it. This table lists the properties and the descriptions.- Delta Lake
The Lakehouse source connector uses the Hadoop file system to read and write data to and from cloud objects, such as AWS, GCS, and Azure. If you want to configure Hadoop related properties, you should use the prefix
hadoop..Examples
You can create a configuration file (JSON or YAML) to set the properties if you use Pulsar Function Worker to run connectors in a cluster.- Delta Lake
-
The Delta table that is stored in the file system
-
The Delta table that is stored in cloud storage (AWS S3, GCS, or Azure)
Data format types
Currently, The Lakehouse source connector only supports reading Delta table changelogs, which adopt aparquet storage format.
How to use
You can use the Lakehouse source connector with Function Worker. You can use the Lakehouse source connector as a non built-in connector or a built-in connector.- Use it as a non built-in connector
- Use it as a built-in connector
If you already have a Pulsar cluster, you can use the Lakehouse source connector as a non built-in connector directly.This example shows how to create a Lakehouse source connector on a Pulsar cluster using the
pulsar-admin sources create command.