---
metadata:
  - name: generator
    content: Diplodoc Platform v5.50.6
alternate:
  - https://ydb.tech/docs/en/concepts/query_execution/federated_query/greenplum.md?version=main
  - https://ydb.tech/docs/ru/concepts/query_execution/federated_query/greenplum.md?version=main
sourcePath: en/core/concepts/query_execution/federated_query/greenplum.md
---
> **Documentation Index:** Fetch the complete configuration index at https://ydb.tech/docs/en/llms.txt

# Working with Greenplum databases

<!-- source: en/concepts/query_execution/federated_query/_includes/experimental_connectors_warning.md -->
{% note warning %}

External connectors are an experimental feature of YDB. To work with external DBMS through connectors, you need to deploy [fq-connector-go](https://ydb.tech/docs/en/devops/deployment-options/manual/federated-queries/connector-deployment.md?version=main#fq-connector-go) and explicitly enable the corresponding sources in the cluster configuration YDB. The functionality may change, so use in production is not recommended without thorough testing.

{% endnote %}
<!-- endsource: en/concepts/query_execution/federated_query/_includes/experimental_connectors_warning.md -->

This section describes the basic information about working with an external [Greenplum](https://greenplum.org) database. Since Greenplum is based on [PostgreSQL](https://ydb.tech/docs/en/concepts/query_execution/federated_query/postgresql.md?version=main), integrations with them work similarly, and some links below may lead to PostgreSQL documentation.

To work with an external Greenplum database, you need to perform the following steps:

1. Create a [secret](https://ydb.tech/docs/en/concepts/datamodel/secrets.md?version=main) containing the password for connecting to the database.


   ```yql
   CREATE SECRET greenplum_datasource_user_password WITH (value = "<password>");
   ```

2. Create an [external data source](https://ydb.tech/docs/en/concepts/datamodel/external_data_source.md?version=main) that describes a specific database within a Greenplum cluster. Pass the network address of the Greenplum [master node](https://greenplum.org/introduction-to-greenplum-architecture/) in the `LOCATION` parameter. By default, the [namespace](https://docs.vmware.com/en/VMware-Greenplum/6/greenplum-database/ref_guide-system_catalogs-pg_namespace.html) `public` is used for reading, but this value can be changed using the optional `SCHEMA` parameter. You can enable encryption of connections to the external database using the `USE_TLS="TRUE"` parameter.


   ```yql
   CREATE EXTERNAL DATA SOURCE greenplum_datasource WITH (
       SOURCE_TYPE="Greenplum",
       LOCATION="<host>:<port>",
       DATABASE_NAME="<database>",
       AUTH_METHOD="BASIC",
       LOGIN="user",
       PASSWORD_SECRET_PATH="greenplum_datasource_user_password",
       USE_TLS="TRUE",
       SCHEMA="<schema>"
   );
   ```

3. <!-- source: en/concepts/query_execution/federated_query/_includes/connector_deployment.md -->
   Deploy the [connector](https://ydb.tech/docs/en/concepts/query_execution/federated_query/architecture.md?version=main#connectors)  and [configure](https://ydb.tech/docs/en/devops/deployment-options/manual/federated-queries/index.md?version=main)  the YDB dynamic nodes to interact with it. Additionally, ensure network access from the YDB dynamic nodes to the external data source (at the address specified in the `LOCATION` parameter of the `CREATE EXTERNAL DATA SOURCE` request). If network connection encryption to the external source was enabled in the previous step, the connector will use the system's root certificates. More details on TLS configuration can be found in the [guide](https://ydb.tech/docs/en/devops/deployment-options/manual/federated-queries/connector-deployment.md?version=main) on deploying the connector.
   <!-- endsource: en/concepts/query_execution/federated_query/_includes/connector_deployment.md -->
4. [Execute a query](#query) to the database.

## Query syntax {#query}

The following SQL query form is used to work with Greenplum:


```yql
SELECT * FROM greenplum_datasource.<table_name>
```


where:

- `greenplum_datasource`: external data source identifier
- `<table_name>` is the name of the table inside the external data source.

## Limitations {#limitations}

When working with Greenplum clusters, there are a number of limitations:

1. <!-- source: en/concepts/query_execution/federated_query/_includes/supported_requests.md -->
   External sources are available only for reading data through `SELECT` queries. The federated query processing engine currently does not support queries that modify tables in external sources.
   <!-- endsource: en/concepts/query_execution/federated_query/_includes/supported_requests.md -->
2. <!-- source: en/concepts/query_execution/federated_query/_includes/datetime_limits.md -->
   If the date value stored in the external data source is outside the allowed range for YDB (all dates used must be later than 1970-01-01 but earlier than 2105-12-31), such a value in YDB will be converted to `NULL`.
   <!-- endsource: en/concepts/query_execution/federated_query/_includes/datetime_limits.md -->
3. <!-- source: en/concepts/query_execution/federated_query/_includes/predicate_pushdown_preamble.md -->
   The YDB federated query processing system is capable of delegating the execution of certain parts of a query to the system acting as the data source. Query fragments are passed through YDB directly to the external system and processed by them. This optimization, known as "predicate pushdown", significantly reduces the amount of data transferred from the source to the federated query processing engine. This reduces network load and saves computational resources for the federated YDB.

   A specific case of predicate pushdown, when the filtering expressions are specified after the `WHERE` keyword, are passed to the data source, is called "filter pushdown". Filter pushdown is possible when using:
   <!-- endsource: en/concepts/query_execution/federated_query/_includes/predicate_pushdown_preamble.md -->

   <!-- source: en/concepts/query_execution/federated_query/_includes/predicate_pushdown_examples.md -->
   |Description|Example|
   |---|---|
   |`NULL` checks|`WHERE column1 IS NULL` or `WHERE column1 IS NOT NULL`|
   |Logical conditions `OR`, `NOT`, `AND` and parentheses for controlling calculation priority. |`WHERE column1 IS NULL OR (column2 IS NOT NULL AND column3 > 10)`.|
   |[Comparison operators](../../../yql/reference/syntax/expressions.md#comparison-operators) with other columns or constants. |`WHERE column1 > column2 OR column3 <= 10`, `WHERE column1 + column2 > 10`, `WHERE column1 = (10 + 10)`|

   When using other types of filters, pushdown to the data source is not performed: filtering of the external table rows will be executed by the federated YDB, which means that YDB will perform a full scan of the external table when processing the query.
   <!-- endsource: en/concepts/query_execution/federated_query/_includes/predicate_pushdown_examples.md -->

   Supported data types for filter pushdown:

   | Data type YDB |
   | --- |
   | `Bool` |
   | `Int8` |
   | `Int16` |
   | `Int32` |
   | `Int64` |
   | `Float` |
   | `Double` |

## Supported data types

In the Greenplum database, the optionality flag of column values (whether a column is allowed or not allowed to contain `NULL` values) is not part of the data type system. The `NOT NULL` constraint for each column is implemented as the `attnotnull` attribute in the [pg_attribute](https://docs.vmware.com/en/VMware-Greenplum/6/greenplum-database/ref_guide-system_catalogs-pg_attribute.html) system catalog, i.e., at the table metadata level. Consequently, all basic Greenplum types can contain `NULL` values by default, and in the YDB type system they must be mapped to [optional](https://ydb.tech/docs/en/yql/reference/types/optional.md?version=main) types.

Below is a table of correspondence between Greenplum types and YDB. All other data types, except those listed, are not supported.

| Greenplum data type | Data type YDB | Notes |
| --- | --- | --- |
| `boolean` | `Optional<Bool>` |  |
| `smallint` | `Optional<Int16>` |  |
| `int2` | `Optional<Int16>` |  |
| `integer` | `Optional<Int32>` |  |
| `int` | `Optional<Int32>` |  |
| `int4` | `Optional<Int32>` |  |
| `serial` | `Optional<Int32>` |  |
| `serial4` | `Optional<Int32>` |  |
| `bigint` | `Optional<Int64>` |  |
| `int8` | `Optional<Int64>` |  |
| `bigserial` | `Optional<Int64>` |  |
| `serial8` | `Optional<Int64>` |  |
| `real` | `Optional<Float>` |  |
| `float4` | `Optional<Float>` |  |
| `double precision` | `Optional<Double>` |  |
| `float8` | `Optional<Double>` |  |
| `json` | `Optional<Json>` |  |
| `date` | `Optional<Date>` | Valid date range from 1970-01-01 to 2105-12-31. If the value goes beyond the range boundaries, `NULL` is returned. |
| `timestamp` | `Optional<Timestamp>` | Valid time range from 1970-01-01 00:00:00 to 2105-12-31 23:59:59. If the value goes beyond the range boundaries, the value `NULL` is returned. |
| `bytea` | `Optional<String>` |  |
| `character` | `Optional<Utf8>` | Default [collation rules](https://www.postgresql.org/docs/current/collation.html), the string is padded with spaces to the required length. |
| `character varying` | `Optional<Utf8>` | Default [collation rules](https://www.postgresql.org/docs/current/collation.html). |
| `text` | `Optional<Utf8>` | Default [collation rules](https://www.postgresql.org/docs/current/collation.html). |
