DataZen Documentation
DataZen User Guide

Enforce a Strong Schema

Overview

Enforcing a strong schema may be important depending on the target systems being considered. DataZen provides a flexible short-hand notation to apply a schema on the current data pipeline data set and will coerce fields as much as possible to match the desired data type.

Two notations are available for data types: SQL-like and C#-like.

You can use the APPLY SCHEMA operation to enforce a schema at once, or use the ADD OR ALTER COLUMN to change a single column at a time. Both commands operate the same way when adding or editing a field, but different options are available for more advanced scenarios.

  • APPLY SCHEMA: See the APPLY SCHEMA documentation for details on how to use it
  • ADD COLUMN: See the ADD COLUMN documentation for details on how to use it

For example, this command adds or updates columns as needed to the current data pipeline data set, and removes any other columns not found in the list. If the dateFound field exists, it will be coerced into a date/time field if possible; if it doesn't exist it will be added with a NULL default value. If data types are not specified, the string data type will be used.

APPLY SCHEMA (
	datetime dateFound
	string(36) guid
	string(255) link
	_batchId = #rndguid()
	_rowId = @rndguid()
) STRICT_COLUMNS;