Install with the GEO Knowledge Hub CLI
The GEO Knowledge Hub CLI produces the deployment for you. You describe the instance once, in a
short configuration file, and gkh deploy renders the Helm values, the Kubernetes Secrets and the
post-install sequence from it. Before it renders anything, it checks your configuration against the
failure modes we have hit while running the GEO Knowledge Hub.
The result is reproducible: the configuration is the only file you keep, and regenerating it always produces the same deployment.
1. Install the CLI
Section titled “1. Install the CLI”The GEO Knowledge Hub CLI is the gkh command. Deployment and validation are plugins it discovers,
so both arrive with a single install:
uv tool install \ --with "gkh-deploy[validation] @ git+https://github.com/geo-knowledge-hub/geo-deploy.git" \ git+https://github.com/geo-knowledge-hub/geo-cli.gitConfirm both plugins loaded:
gkh --versionIt is expected to you to see:
gkh-cli 0.1.0 deploy: gkh-deploy 0.1.0 validate: gkh-deploy 0.1.02. Describe the instance
Section titled “2. Describe the instance”gkh deploy init writes the configuration, asking only for what cannot be defaulted. Use the
hostname you settled on when you prepared the cluster:
gkh deploy init \ --profile minimal \ --hostname invenio.local \ --admin-email admin@invenio.orggkh deploy init \ --profile standard \ --hostname invenio.local \ --admin-email admin@invenio.org \ --set datacite.prefix=10.5072The standard profile enables DOI minting, so it requires a DataCite prefix. 10.5072 is
DataCite’s test prefix, you can replace it with your own before this goes anywhere real.
This writes gkh-deploy.yaml, the only file you maintain:
version: 1target: k8shostname: invenio.localrelease: invenionamespace: invenioimage: repository: geoknowledgehub/geo-knowledge-hub tag: v1.7.0.dev17admin: email: admin@invenio.orgdatacite: enabled: false prefix: '10.5072' test_mode: true secret_name: gkh-dataciteingress: enabled: true class: nginx tls_secret: ''scaling: web_replicas: 1 worker_replicas: 1 uwsgi_processes: 2 uwsgi_threads: 2 worker_concurrency: 2storage: enabled: true size: 5G class: standardsecrets: postgresql: gkh-postgresql rabbitmq: gkh-rabbitmqopensearch: heap: 1024m memory_request: 1Gi memory_limit: 2Gi cpu_request: 500m cpu_limit: '1'flower: enabled: falseextra_config: {}extra_values: {}Open it and adjust anything your cluster needs, typically the image tag, the ingress class and the storage class. The equivalent hand-written Helm values file is around two hundred lines, and it is generated from this one.
Anything the chart accepts but this file does not name can still be set through extra_config (application settings, the INVENIO_* variables) and extra_values (raw chart values). The full list of chart parameters is in the parameters reference.
3. Check the configuration
Section titled “3. Check the configuration”Once you finish the configurations of gkh-deploy.yaml file, you can validate it using the GEO Knowledge Hub CLI. In this process, the CLI checks the values defined in the document to ensure configurations are well defined (e.g., DOI provider is defined, OpenSearch memory, Image Tag and many others).
To validate your configuration file, you can use the following command:
gkh deploy checkIf you have everything right, it is expected you to see the following message:
No problems found.Each check covers a way the deployment can go wrong without saying so. Most of them are silent issues: the instance installs, looks healthy, and misbehaves later. So, the CLI implements those checks to support your deploy.
To complement, in the table below, it is presented the rules we use to check the user configuration and the reasons why we check them:
| Check | What goes wrong | What to do |
|---|---|---|
datacite-secret-unreachable | DOI credentials are supplied twice over, so the application receives neither and the pods never start. | Supply them one way only: let the bundle hold them, or point at a secret you created yourself. |
opensearch-heap-oversized | The search engine is allowed more memory than its container has, so it is killed while loading the vocabularies. It looks like intermittent search timeouts, not a memory problem. | Give the container at least twice the memory the search engine is allowed to use. |
ratelimit-storage-url-missing | The rate limiter cannot find where to record request counts, and answers every request with an error. | Point the rate limiter at the same cache the rest of the application uses. |
init-job-enabled | The chart’s own setup job runs before the vocabularies exist and repeats work the post-install sequence already does. | Leave it disabled and let bootstrap.sh seed the instance. |
default-users-not-a-mapping | The accounts you asked for are never created, and nothing reports it. | Write each account as an email and a password, not as a plain list. |
hostname-override-inconsistent | The web interface and the API end up on different addresses, and the browser’s requests are rejected. | Serve both from one hostname. The API address is the site address with /api on the end. |
unknown-chart-key | A setting you wrote has no effect. Helm accepts values it does not recognise and ignores them in silence. | Check the spelling against the parameters reference, or remove the setting. |
image-tag-not-pinned | The deployment is not reproducible: the same configuration installed again gives you a different version of the application. | Pin a known-good release tag instead of a moving one. |
deprecated-extra-config | An older spelling of the application settings still works, but is superseded, and the two disagree wherever both set the same value. | Move those settings to the current spelling, extraConfig. |
Findings are reported as errors or warnings:
warning: [image-tag-not-pinned] image.tag is 'latest'. Pin a known-good release tag so the same configuration always produces the same instance. basis: the GEO Knowledge Hub chart overlay: never `latest`Note that findings are reported as warnings, so the script won’t stop when it hits a minor issue. This prevents users from getting blocked mid-run. By default, validation also skips the DOI provider check. If you want any finding to halt the check, pass the --strict flag, which turns warnings into errors.
4. Generate the bundle
Section titled “4. Generate the bundle”Once your configuration file (i.e., gkh-deploy.yaml) passes checks in the previous step, you can generate a deployment bundle. For this, you can use the following command:
gkh deploy generate -o deploy/This command will produce a directory with a complete deployment bundle, which includes:
Directorydeploy/
- README.md the install commands for this bundle
- values.yaml Helm values
- secrets.sh creates the Secrets
values.yamlrefers to - bootstrap.sh the post-install sequence
- gkh-deploy.yaml a copy of the configuration it was rendered from
Now, you are ready to start configuring the GEO Knowledge Hub in your cluster.
5. Create the Secrets
Section titled “5. Create the Secrets”Once the cluster is up and running, you can start configuring it. For this, first access the generated deploy/ directory:
cd ./deploy/Then create the Secrets used by the services:
./secrets.shThe script creates the namespace if it is missing, followed by the PostgreSQL and RabbitMQ Secrets with generated passwords. Re-running it is safe, since each Secret is created only if it is absent.
6. Install the chart
Section titled “6. Install the chart”Now, in your helm, install the invenio chart repository:
helm repo add inveniordm https://inveniosoftware.github.io/helm-invenio/helm repo updateThen install, using the values.yaml available in the deploy/, you can start the GEO Knowledge Hub required services:
helm install invenio inveniordm/invenio \ --version 0.11.1 \ --namespace invenio \ -f values.yamlWatch the rollout:
kubectl get pods -n invenio -w7. Run the post-install sequence
Section titled “7. Run the post-install sequence”Once all services are starting, you can bootstrap the platform content (e.g., roles, permissions and vocabularies) with the following command:
./bootstrap.shThe script finds the worker pod itself and waits for the database, broker, cache and search cluster to be ready before it starts, so you do not have to watch the rollout first. It seeds the schema, the file location, the GEO-specific roles and access policies, the administrator, the search indices and the controlled vocabularies, then waits for the workers to fall idle. It takes several minutes, mostly in the vocabularies.
None of the steps can be applied twice. If one fails, fix the cause and resume from it rather than starting over:
./bootstrap.sh rolesThe steps are db, files, roles, access, users, index and fixtures, in that order. What each of them does is explained in Post-install setup.
To read the commands executed in each step without running any of them, you can use the following command:
gkh deploy bootstrapNothing is executed and no cluster is contacted, so you can review the sequence, substitute the administrator password, and run it yourself inside the worker container if you want.
Next steps
Section titled “Next steps”1) Open the instance on the address you set up when you prepared the cluster, and log in with the administrator account the sequence created.
2) Validate the deployment with gkh validate, which drives the instance over its API and UI and reports what does not work.
3) Review the parameters reference when you need a chart value that gkh-deploy.yaml does not name directly.