A Clojure(Script) library for declarative data description and validation.
One of the difficulties with bringing Clojure into a team is the overhead of understanding the kind of data (e.g., list of strings, nested map from long to string to double) that a function expects and returns. While a full-blown type system is one solution to this problem, we present a lighter weight solution: schemas. (For more details on why we built Schema, check out this post.
- Referenced post: Schema for Clojure(Script) Data Shape Declaration and Validation
Schema is a rich language for describing data shapes, with a variety of features:
- Data validation, with descriptive error messages of failures (targeted at programmers)
- Annotation of function arguments and return values, with optional runtime validation
- Schema-driven data coercion, which can automatically, succinctly, and safely convert complex data types (see the Coercion section below)
- Other
- Schema is also built into our plumbing and fnhouse libraries, which illustrate how we build services and APIs easily and safely with Schema
- Schema also supports experimental clojure.test.check data generation from Schemas, as well as completion of partial datums, features we've found very useful when writing tests as part of the
schema-generatorslibrary
Getting started
The spec library (API docs) specifies the structure of data, validates or conforms it, and can generate data based on the spec.
Clojure is a dynamic language. Among other things this means that type annotations are not required for code to run. While Clojure has some support for type hints, they are not an enforcement mechanism, nor comprehensive, and are limited to communicating information to the compiler to aid in efficient code generation. Clojure gets runtime checking of a richer set of types by the JVM itself.
However it has always been a guiding principle of Clojure, widely valued and practiced by the community, to simply represent information as data. Thus important properties of Clojure systems are represented and conveyed by the shape and other predicative properties of the data, not captured or checked anywhere since the runtime types are indistinguishable heterogeneous maps and vectors.
Documentation strings can be used to communicate with human consumers, but they can’t be leveraged by programs or tests, i.e. they have minimal power. Users have turned to various libraries such as Schema and Herbert to get more powerful specifications.
When it comes to data exchange, JSON Schema stands out as a powerful standard for defining the structure and rules of JSON data. It uses a set of keywords to define the properties of your data.
While JSON Schema provides the language, validating a JSON instance against a schema requires a JSON Schema validator. The JSON validator checks if the JSON documents conform to the schema.
JSON Schema Validators are tools that implement the JSON Schema specification. Such tooling enables easy integration of JSON Schema into projects of any size.
Cerberus provides powerful yet simple and lightweight data validation functionality out of the box and is designed to be easily extensible, allowing for custom validation. It has no dependencies and is thoroughly tested from Python 2.7 up to 3.6, PyPy and PyPy3.
While exploring tooling for Kubernetes I had need for schemas to describe the definition files, and went looking for something that didn't require either kubectl or similar installed or even a working Kubernetes installation.
It turns out that the OpenAPI specification contain this information, but not in a particularly usable format for tools which might just want a raw JSON Schema.
This repository contains a set of schemas for most recent Kubernetes versions. For each specified Kubernetes versions you should find four different flavours:
vX.Y.Z - URL referenced based on the specified GitHub repository
vX.Y.Z-standalone - de-referenced schemas, more useful as standalone documents
vX.Y.Z-local - relative references, useful to avoid the network dependency
vX.Y.Z-strict - prohibits properties not defined in the schema
Note that the Kubernetes API allows additional properties to be submitted, but kubectl acts like the strict flavour above.
jsonschema is an implementation of JSON Schema for Python (supporting 2.7+ including Python 3).
Features
- Full support for Draft 6, Draft 4 and Draft 3
- Lazy validation that can iteratively report all validation errors.
- Small and extensible
- Programmatic querying of which properties or items failed validation.
JSON Schema is a vocabulary that allows you to annotate and validate JSON documents.