Pipeline#
Pipeline is a Transformer decorator capable of composing an arbitrarily long series of Transformer middleware into a single unit. It fits the stack to a training dataset without mutating it, transforms incoming samples by streaming them through each transformer in order, and — when updated — refines the fitting of Elastic transformers (or lazily fits any Stateful ones that have not yet been seen) while streaming a working copy of the data through the chain.
Interfaces: Transformer, Stateful, Elastic, Persistable
Data Type Compatibility: Categorical, Continuous, Image, Other
Parameters#
| # | Name | Default | Type | Description |
|---|---|---|---|---|
| 1 | transformers | array | A list of transformers to be composed in order. |
Example#
use Rubix\ML\Transformers\Pipeline;
use Rubix\ML\Transformers\HotDeckImputer;
use Rubix\ML\Transformers\OneHotEncoder;
use Rubix\ML\Transformers\ZScaleStandardizer;
$transformer = new Pipeline([
new HotDeckImputer(5),
new OneHotEncoder(),
new ZScaleStandardizer(),
]);
Fitting and Updating#
Because Pipeline is a Stateful transformer itself, it exposes fit() and fitted(). Calling fit() will refit every stateful transformer in the stack to the current dataset, streaming a working copy of the data through the chain — the input dataset is left unaltered.
$transformer = new Pipeline([
new OneHotEncoder(),
new ZScaleStandardizer(),
]);
$transformer->fit($dataset);
If any transformer in the stack is Elastic, the pipeline is also Elastic: update() will refine each elastic fitting and lazily fit any stateful transformer that has not yet been seen, again streaming a working copy of the data through the chain without touching the input.
$transformer->update($dataset);
Transformers that are stateless and non-elastic are applied as-is — they are always fitted(), so they contribute nothing to the fitted() status of the pipeline. An empty pipeline is always considered fitted and behaves as a no-op.
Since fitting does not transform the data, transform your dataset explicitly with the pipeline's transform() method or, more idiomatically, with Dataset::apply():
$dataset->apply($transformer);
Persistence#
A fitted pipeline can be saved to and loaded from storage by decorating it with the Persistent Transformer, which interfaces with the persistence subsystem on your behalf.
use Rubix\ML\Transformers\Pipeline;
use Rubix\ML\Transformers\OneHotEncoder;
use Rubix\ML\Transformers\PersistentTransformer;
use Rubix\ML\Persisters\Filesystem;
$transformer = new PersistentTransformer(
new Pipeline([new OneHotEncoder()]),
new Filesystem('example.rbx')
);
$transformer->fit($dataset);
$transformer->save();
$transformer = PersistentTransformer::load(new Filesystem('example.rbx'));