Getting Started
Introduction
There are several options for getting started with Apache Gravitino.
Installing and configuring Hive and Trino can be a little complex. If you are unfamiliar with the technologies, using Docker might be a good choice. There are pre-packaged containers for Gravitino, Apache Hive, Apache Hadoop, Trino, MySQL, PostgreSQL, and others. Check installing Gravitino playground for more details.
This page guides you through the process of downloading and installing Gravitino from source.
- Prepare environment
- Deploy and run Gravitino on Amazon Web Service (AWS)
- Deploy and run Gravitino on Google Compute Platform (GCP)
- Run Gravitino on your own machine
- Install Gravitino
- Start Gravitino
- Install Apache Hive
- Interact with Apache Gravitino API
If you want to access the instance remotely, be sure to read Accessing Gravitino on AWS externally.
Environment Preparation
AWS
To work in an AWS environment, follow these steps:
-
In the AWS console, launch a new instance. Select
Ubuntuas the operating system andt2.xlargeas the instance type. Create a key pair named Gravitino.pem for SSH access and download it. Allow HTTP and HTTPS traffic if you want to connect to the instance remotely. Set the Elastic Block Store storage to 20GiB. Leave all other settings at their defaults. Other operating systems and instance types may work but have not been fully tested. -
Start the instance and connect to it via SSH using the downloaded
.pemfile:ssh ubuntu@<IP_address> -i ~/Downloads/Gravitino.pemNote: you may need to adjust the permissions on your
.pemfile usingchmod 400to enable SSH connections. -
Update the Ubuntu OS to ensure it's up-to-date:
sudo apt update
sudo apt upgradeYou may need to reboot the instance for all changes to take effect.
-
Install the Java Development Kit (JDK). Java 17 is supported.
sudo apt install openjdk-<version>-jdk-headlessVerify the Java version with:
java -versionYou should see information about the OpenJDK version.
GCP
To work on the GCP platform, follow these steps:
-
In the Google Cloud console, launch a new instance. Select
e2-standard-4as the instance type and 20 GB for the boot disk size. Allow HTTP and HTTPS traffic if you want to connect to the instance remotely. Leave all other settings as their defaults. Other operating systems and instance types may work, but are not fully tested. -
Start the instance and connect to it via the SSH-in-browser tool.
-
Update the Debian OS to ensure it's up-to-date:
sudo apt update
sudo apt upgradeYou may need to reboot the instance for all changes to take effect.
-
Install the Java Development Kit (JDK), Java 17 is supported.
wget -O - https://apt.corretto.aws/corretto.key | sudo gpg --dearmor -o /usr/share/keyrings/corretto-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/corretto-keyring.gpg] https://apt.corretto.aws stable main" | sudo tee /etc/apt/sources.list.d/corretto.list
sudo apt-get update
sudo apt-get install -y java-<version>-amazon-corretto-jdkVerify the Java version with:
java -versionYou should see information about the OpenJDK version.
Local Workstation
To build and install Gravitino locally on a macOS or a Linux workstation, follow these steps:
-
Install the Java Development Kit (JDK). Java 17 is supported. This can be done using sdkman, for example:
sdk install java <version>You can also use different package managers to install JDK, for example, Homebrew on macOS,
apton Ubuntu/Debian, andyumon CentOS/RedHat.
Install Gravitino
Install Gravitino from the binary release packages or the container images. Follow how-to-install.
Or you can install Gravitino from scratch. Follow how-to-build and how-to-install.
Start Gravitino
Start Gravitino using the gravitino.sh script:
<path-to-gravitino>/bin/gravitino.sh start
Install Apache Hive
If you already have Apache Hive and Apache Hadoop in your environment, you can skip this step and use the existing service with Gravitino. Or else, you can follow the instructions to install Apache Hive.