Skip to content

[source]

Pipeline#

Pipeline is a Transformer decorator capable of composing an arbitrarily long series of Transformer middleware into a single unit. It fits the stack to a training dataset without mutating it, transforms incoming samples by streaming them through each transformer in order, and — when updated — refines the fitting of Elastic transformers (or lazily fits any Stateful ones that have not yet been seen) while streaming a working copy of the data through the chain.

Interfaces: Transformer, Stateful, Elastic, Persistable

Data Type Compatibility: Categorical, Continuous, Image, Other

Parameters#

# Name Default Type Description
1 transformers array A list of transformers to be composed in order.

Example#

use Rubix\ML\Transformers\Pipeline;
use Rubix\ML\Transformers\HotDeckImputer;
use Rubix\ML\Transformers\OneHotEncoder;
use Rubix\ML\Transformers\ZScaleStandardizer;

$transformer = new Pipeline([
    new HotDeckImputer(5),
    new OneHotEncoder(),
    new ZScaleStandardizer(),
]);

Fitting and Updating#

Because Pipeline is a Stateful transformer itself, it exposes fit() and fitted(). Calling fit() will refit every stateful transformer in the stack to the current dataset, streaming a working copy of the data through the chain — the input dataset is left unaltered.

$transformer = new Pipeline([
    new OneHotEncoder(),
    new ZScaleStandardizer(),
]);

$transformer->fit($dataset);

If any transformer in the stack is Elastic, the pipeline is also Elastic: update() will refine each elastic fitting and lazily fit any stateful transformer that has not yet been seen, again streaming a working copy of the data through the chain without touching the input.

$transformer->update($dataset);

Transformers that are stateless and non-elastic are applied as-is — they are always fitted(), so they contribute nothing to the fitted() status of the pipeline. An empty pipeline is always considered fitted and behaves as a no-op.

Since fitting does not transform the data, transform your dataset explicitly with the pipeline's transform() method or, more idiomatically, with Dataset::apply():

$dataset->apply($transformer);

Persistence#

A fitted pipeline can be saved to and loaded from storage by decorating it with the Persistent Transformer, which interfaces with the persistence subsystem on your behalf.

use Rubix\ML\Transformers\Pipeline;
use Rubix\ML\Transformers\OneHotEncoder;
use Rubix\ML\Transformers\PersistentTransformer;
use Rubix\ML\Persisters\Filesystem;

$transformer = new PersistentTransformer(
    new Pipeline([new OneHotEncoder()]),
    new Filesystem('example.rbx')
);

$transformer->fit($dataset);

$transformer->save();
$transformer = PersistentTransformer::load(new Filesystem('example.rbx'));