CLI Reference
Introduction
Use --help to list all commands, or --help --type <command> for command-specific help.
By default, optimizer CLI loads conf/gravitino-optimizer.conf from the current working
directory. Use --conf-path only when you need a custom config file.
Command Quick Reference
Command (--type) | Required options | Optional options | Purpose |
|---|---|---|---|
submit-strategy-jobs | --identifiers, --strategy-name | --dry-run, --limit | Recommend and optionally submit jobs |
update-statistics | --calculator-name | --identifiers, --statistics-payload, --file-path | Calculate and persist statistics |
append-metrics | --calculator-name | --identifiers, --statistics-payload, --file-path | Calculate and append metrics |
monitor-metrics | --identifiers, --action-time | --range-seconds, --partition-path | Evaluate rules with before/after metrics |
list-table-metrics | --identifiers | --partition-path | Query stored table or partition metrics |
list-job-metrics | --identifiers | None | Query stored job metrics |
submit-update-stats-job | --identifiers | --dry-run, --update-mode, --updater-options, --spark-conf | Submit built-in Iceberg update stats/metrics Spark jobs |
Option Field Meanings
| Option | Meaning | Used by |
|---|---|---|
--identifiers | Comma-separated identifiers. Table format supports catalog.schema.table (or schema.table when default catalog is configured). | Most commands |
--strategy-name | Policy name to evaluate, for example iceberg_compaction_default. | submit-strategy-jobs |
--dry-run | Preview mode. Prints recommendations or job configs without submitting jobs. | submit-strategy-jobs, submit-update-stats-job |
--limit | Maximum number of strategy jobs to process. Must be > 0. | submit-strategy-jobs |
--calculator-name | Statistics/metrics calculator implementation name (for example local-stats-calculator). | update-statistics, append-metrics |
--statistics-payload | Inline JSON Lines content as input. Mutually exclusive with --file-path. | update-statistics, append-metrics |
--file-path | Path to JSON Lines input file. Mutually exclusive with --statistics-payload. | update-statistics, append-metrics |
--action-time | Action timestamp in epoch seconds used as evaluation anchor. | monitor-metrics |
--range-seconds | Time window (seconds) for monitor evaluation. Default is 86400 (24h). | monitor-metrics |
--partition-path | Partition path JSON array, for example '[{"dt":"2026-01-01"}]'. Requires exactly one identifier. | monitor-metrics, list-table-metrics |
--update-mode | Controls what built-in update job updates: stats, metrics, or all (default). | submit-update-stats-job |
--updater-options | Flat JSON map passed to updater logic. For stats/all, include gravitino_uri and metalake. | submit-update-stats-job |
--spark-conf | Flat JSON map of Spark and Iceberg catalog configs used by the job. | submit-update-stats-job |
Global option:
--conf-path: Optional custom config file path. If omitted, CLI usesconf/gravitino-optimizer.conf.
Input Format for local-stats-calculator
local-stats-calculator reads JSON Lines (one JSON object per line).
Reserved Fields
stats-type:table,partition, orjobidentifier: object identifierpartition-path: only for partition data, for example{"dt":"2026-01-01"}timestamp: optional epoch seconds (record-level default timestamp for metric points)
All other fields are treated as metric or statistic values.
Supported Examples by Scope
Use JSON Lines (one JSON object per line). The following examples focus on table, partition, and job scopes with multiple metric/statistic fields:
{"stats-type":"table","identifier":"catalog.db.t1","timestamp":1735689600,"row_count":100}
{"stats-type":"table","identifier":"catalog.db.t1","row_count":100,"total_file_size":1048576}
{"stats-type":"table","identifier":"catalog.db.t1","timestamp":1735689660,"row_count":120,"file_count":24,"avg_file_size":10485.76}
{"stats-type":"partition","identifier":"catalog.db.t1","timestamp":1735689720,"partition-path":{"dt":"2026-01-01"},"row_count":20}
{"stats-type":"partition","identifier":"catalog.db.t1","partition-path":{"dt":"2026-01-01","region":"us"},"row_count":12,"file_count":3}
{"stats-type":"job","identifier":"job-1","timestamp":1735689800,"duration_ms":12500,"rewritten_files":18}
Identifier Rules
- Table and partition records:
catalog.schema.table - If
gravitino.optimizer.gravitinoDefaultCatalogis set,schema.tableis also accepted - Job records: parsed as a regular Gravitino
NameIdentifier
CLI Workflow Examples
Batch Statistics Update
Calculate and persist table or partition statistics from JSONL input.
./bin/gravitino-optimizer.sh \
--type update-statistics \
--calculator-name local-stats-calculator \
--file-path ./table-stats.jsonl
Batch Metrics Append
Calculate and append table or job metrics from JSONL input.
./bin/gravitino-optimizer.sh \
--type append-metrics \
--calculator-name local-stats-calculator \
--file-path ./table-stats.jsonl
Dry-Run Strategy Submission
Preview recommendations without actually submitting jobs.
./bin/gravitino-optimizer.sh \
--type submit-strategy-jobs \
--identifiers rest_catalog.db.t1 \
--strategy-name iceberg_compaction_default \
--dry-run \
--limit 10
Submit Strategy Jobs
Submit jobs for identifiers that match the given policy name.
./bin/gravitino-optimizer.sh \
--type submit-strategy-jobs \
--identifiers rest_catalog.db.t1 \
--strategy-name iceberg_compaction_default \
--limit 10
Monitor Metrics
Evaluate monitor rules around an action time.
./bin/gravitino-optimizer.sh \
--type monitor-metrics \
--identifiers catalog.db.sales \
--action-time 1735689600 \
--range-seconds 86400
Configure evaluator rules in gravitino-optimizer.conf:
gravitino.optimizer.monitor.gravitinoMetricsEvaluator.rules = table:row_count:avg:le,job:duration:latest:le
Rule format is scope:metricName:aggregation:comparison:
scope:tableorjob(tablerules also apply to partition scope)aggregation:max|min|avg|latestcomparison:lt|le|gt|ge|eq|ne
When metrics are produced by submit-update-stats-job --update-mode metrics, metric names are
often custom-* (for example custom-data-file-mse). Use list-table-metrics first and
configure rules with the exact metric names returned by your environment.
Submit Built-In Update Stats Jobs
Submit built-in Iceberg update stats/metrics Spark jobs directly.
./bin/gravitino-optimizer.sh \
--type submit-update-stats-job \
--identifiers rest_catalog.db.t1 \
--update-mode all \
--updater-options '{"gravitino_uri":"http://localhost:8090","metalake":"test"}' \
--spark-conf '{"spark.sql.catalog.rest_catalog.type":"rest","spark.sql.catalog.rest_catalog.uri":"http://localhost:9001/iceberg","spark.hadoop.fs.defaultFS":"file:///"}'
Notes:
--identifierssupportscatalog.schema.tableorschema.table(when default catalog is configured).--update-modesupportsstats|metrics|all(defaultall).- For
statsorall,--updater-optionsmust includegravitino_uriandmetalake. - If
--updater-optionsincludes external JDBC metrics settings (gravitino.optimizer.jdbcMetrics.*), ensure the JDBC driver JAR is available to Spark runtime classpath (for example viaspark.jarsin--spark-conf). --spark-confand--updater-optionsare flat JSON maps.
List Table Metrics
Query stored metrics at table scope.
./bin/gravitino-optimizer.sh \
--type list-table-metrics \
--identifiers catalog.db.sales
For partition scope, provide a partition path JSON array:
./bin/gravitino-optimizer.sh \
--type list-table-metrics \
--identifiers catalog.db.sales \
--partition-path '[{"dt":"2026-01-01"}]'
List Job Metrics
Query stored metrics at job scope.
./bin/gravitino-optimizer.sh \
--type list-job-metrics \
--identifiers catalog.db.optimizer_job
Output Guide
SUMMARY: ...: summary forupdate-statisticsandappend-metricsDRY-RUN: ...: recommendation preview without job submissionSUBMIT: ...: strategy job or built-in update-stats job submitted successfullySUMMARY: submit-update-stats-job ...: summary for built-in update-stats submissionMetricsResult{...}: returned by list commandsEvaluationResult{...}: returned by monitor command
Examples:
SUMMARY: statistics totalRecords=3 tableRecords=2 partitionRecords=1 jobRecords=0
DRY-RUN: strategy=iceberg-data-compaction identifier=rest_catalog.db.t1 score=95 jobTemplate=builtin-iceberg-rewrite-data-files jobOptions={catalog_name=rest_catalog, table_identifier=db.t1}
SUBMIT: strategy=iceberg-data-compaction identifier=rest_catalog.db.t1 score=95 jobTemplate=builtin-iceberg-rewrite-data-files jobOptions={catalog_name=rest_catalog, table_identifier=db.t1} jobId=1f54c6d3-4e27-4cc8-bdfa-b05ecf59a4c2
DRY-RUN: identifier=rest_catalog.db.t1 jobTemplate=builtin-iceberg-update-stats jobConfig={catalog_name=rest_catalog, table_identifier=db.t1, update_mode=all, updater_options={"gravitino_uri":"http://localhost:8090","metalake":"test"}, spark_conf={"spark.master":"local[2]","spark.hadoop.fs.defaultFS":"file:///"}}
SUMMARY: submit-update-stats-job total=1 submitted=1 dryRun=false
MetricsResult{scopeType=TABLE, identifier=rest_catalog.db.t1, partitionPath=<table-or-job-scope>, metrics={row_count=[{timestamp=1735689600, value=100}]}}
EvaluationResult{scopeType=TABLE, identifier=rest_catalog.db.t1, partitionPath=<table-or-job-scope>, evaluation=true, evaluatorName=gravitino-metrics-evaluator, actionTimeSeconds=1735689600, rangeSeconds=86400, beforeMetrics={row_count=[MetricSample{timestampSeconds=1735686000, value=120}]}, afterMetrics={row_count=[MetricSample{timestampSeconds=1735689600, value=100}]}}
Built-in Job Templates
Three job templates ship with the service, and they are complementary rather than alternatives. A full maintenance pass collects statistics, compacts data files, and then expires the snapshot history that compaction just created.
| Job template | What it does |
|---|---|
builtin-iceberg-update-stats | Collects file statistics and metrics |
builtin-iceberg-rewrite-data-files | Compacts small data files |
builtin-iceberg-expire-snapshots | Removes old snapshot metadata |
Each can be submitted directly over REST, and the first two are also what the policy-driven workflow submits on your behalf. See Quick Start for the policy-driven path.
These templates set Iceberg Spark session and catalog classes, but they do not list an Iceberg Spark
runtime in jars. gravitino-jobs also excludes that runtime from its shaded JAR, so the version
that runs with your Spark cluster is yours to supply. Provide a matching
iceberg-spark-runtime-<sparkMajor>_<scala> JAR on the Spark classpath used by the job executor —
commonly through spark.jars in spark_conf, or by installing it into SPARK_HOME. Align the
artifact with the Spark, Scala, and Iceberg versions you actually run. A reference coordinate used
in Gravitino's own jobs tests is org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.11.0.
Without that runtime, built-in Iceberg jobs fail after Spark starts instead of continuing without
Iceberg support.
Optional template arguments are still listed as --flag + {{placeholder}} pairs. If jobConf
omits a key (or leaves the placeholder unresolved), the flag remains on the process command line as
a dangling argument (for example --updater-options with no value before --spark-conf). Callers
and UIs should supply every placeholder they care about with an explicit value, including optional
ones they intentionally disable or leave at a documented default, rather than omitting the key.
Update Statistics
builtin-iceberg-update-stats reads a table and writes back the statistics and metrics that policies evaluate. Compaction policies read custom-data-file-mse and custom-delete-file-number, so nothing else will fire until this job has run at least once.
Its jobConf is documented in Configuration.
Rewrite Data Files
builtin-iceberg-rewrite-data-files performs the compaction itself, merging small data files into larger ones. It is what a compaction policy submits when its thresholds are crossed.
In alpha this works only on Iceberg tables where every partition uses an identity transform. Tables combining identity with a time or bucket transform fail during the rewrite, which is covered in Troubleshooting.
For the policy that drives it, including threshold tuning, see Iceberg Compaction Policy.
Expire Snapshots
builtin-iceberg-expire-snapshots removes old Iceberg snapshots and the metadata files behind them. Without periodic expiration, snapshot JSON files and manifest lists accumulate indefinitely, which slows table operations and wastes storage. Compaction makes this worse, since every rewrite creates a snapshot.
The job calls Iceberg's expire_snapshots stored procedure through Spark SQL.
| Property | Value |
|---|---|
| Name | builtin-iceberg-expire-snapshots |
| Type | Spark |
| Version | v1 |
| Main class | org.apache.gravitino.maintenance.jobs.iceberg.IcebergExpireSnapshotsJob |
Parameters
catalog_name and table_identifier are required. The rest are optional.
| Key | Description | Default |
|---|---|---|
catalog_name | Iceberg catalog name as registered in Spark | Required |
table_identifier | Fully qualified table name, such as db.sample | Required |
older_than | Expire snapshots older than this yyyy-MM-dd HH:mm:ss timestamp | Five days ago |
retain_last | Minimum number of recent snapshots to keep regardless of age | 1 |
stream_results | Streams intermediate delete results when present | Disabled |
spark_conf | JSON map of Spark configuration | None |
older_than and retain_last work together, and retain_last wins. Setting older_than to yesterday with retain_last at 5 keeps five snapshots even if all five are older than yesterday.
Submitting the Job
curl -X POST -H "Accept: application/vnd.gravitino.v1+json" \
-H "Content-Type: application/json" \
-d '{
"jobTemplateName": "builtin-iceberg-expire-snapshots",
"jobConf": {
"catalog_name": "rest_catalog",
"table_identifier": "db.t1",
"older_than": "2024-01-01 00:00:00",
"retain_last": "3",
"spark_master": "local[2]",
"spark_executor_instances": "1",
"spark_executor_cores": "1",
"spark_executor_memory": "1g",
"spark_driver_memory": "1g",
"catalog_type": "rest",
"catalog_uri": "http://localhost:9001/iceberg",
"warehouse_location": ""
}
}' \
http://localhost:8090/api/metalakes/test/jobs
Omitting older_than and passing only retain_last is the safer default for a first run, since it bounds the result by count rather than by a date you have to reason about.
The job builds this statement, including only the optional parameters you supplied:
CALL `rest_catalog`.system.expire_snapshots(
table => 'db.t1',
older_than => TIMESTAMP '2024-01-01 00:00:00',
retain_last => 3,
stream_results => true
)
Verifying the Result
curl -sS "http://localhost:8090/api/metalakes/test/jobs/{job_id}" | jq '.job.state'
cat /tmp/gravitino/jobs/staging/test/builtin-iceberg-expire-snapshots/{job_id}/stdout.log
A successful run reports its state as SUCCEEDED and logs the counts it removed:
Expire Snapshots Results:
Deleted data files: 12
Deleted manifest files: 8
Deleted manifest lists: 3
Related
- Table Maintenance Service for the concepts and the walkthrough
- Configuration for the three configuration layers
- Iceberg Compaction Policy for tuning the built-in strategy
- Manage Jobs for job status and templates