Typed workflows
Typed workflows let you declare the types of a named workflow's inputs and outputs in its take: and emit: sections. Nextflow uses the annotations to validate the values that enter and leave the workflow against their declared types.
nextflow.enable.types = true
workflow hello_bye {
take:
samples: Channel<Path>
main:
ch_hello = hello(samples)
val_bye = bye(ch_hello.collect())
emit:
result: Value<Path> = val_bye
}
In the above example, hello_bye takes a channel of files (Channel<Path>) and emits a dataflow value with a single file (Value<Path>).
Typed workflows require the nextflow.enable.types feature flag. For an overview of static typing and how to enable it, see Static typing. For the complete syntax reference, see Workflow (typed).
Supported types
Channels and values are modeled using the dataflow types:
- Use
Channel<V>for a channel of values. - Use
Value<V>for a single dataflow value.
The element type V can be any standard type, such as Path, String, or a record.
Workflow inputs can be channels, dataflow values, or regular values. Workflow outputs can be channels or dataflow values.
Restricted syntax
The following syntax patterns are no longer supported in typed workflows:
-
Using
Channelto access channel factories (usechannelinstead) -
Using implicit closure parameters (declare parameters explicitly instead)
-
Using
setortapto assign channels (use standard assignments instead) -
Composing dataflow logic with
|and&(use standard method calls instead) -
Accessing process and workflow outputs via
.out(use standard assignments instead)
For more information about preparing existing code, see Preparing for static typing.
Operators
The operator library supports static typing and records. All operators work in both typed and legacy workflows, but only a core subset of operators is recommended for static typing.
For more information about best practices when migrating existing code, see Using operators with static typing.
Pipeline composition
An entire pipeline -- the params block, entry workflow, and output block of a script -- can be included as a named workflow and called like any other workflow. The params block acts as the take: section and the output block acts as the emit: section.
Given the following pipeline:
// pipelines/rnaseq.nf
nextflow.enable.types = true
params {
input: Channel<Sample>
aligner: String = 'star_salmon'
fasta: Path
}
workflow {
main:
// ...
publish:
bams = ch_bams
multiqc = val_multiqc
}
output {
bams: Channel<Path> { path 'bams' }
multiqc: Path { path 'multiqc' }
}
It can be included and called as follows:
include { workflow as RNASEQ } from './pipelines/rnaseq.nf'
workflow {
main:
rnaseq = RNASEQ(record(
input: samples,
fasta: file('index.fasta')
))
rnaseq.bams.view() // Channel<Path>
rnaseq.multiqc.view() // Value<Path>
}
The included pipeline must be aliased to a specific name (RNASEQ), which is also used to scope its processes in the config:
process {
withName: 'RNASEQ:STAR_ALIGN' {
cpus = 12
memory = 72.GB
}
}
Because the included pipeline is part of the calling pipeline's dataflow graph, it can consume a channel produced by another pipeline, and begin working on each item as soon as it is emitted.
Note the following:
-
Both the calling script and the included pipeline must enable static typing (
nextflow.enable.types = true). -
The pipeline is called with a single record, with one field for each param. Params with a default value can be omitted.
-
The outputs published by the included pipeline are emitted to the calling workflow instead of being published. Only the calling pipeline decides what is published, by declaring its own
outputblock. -
Params and outputs are not inherited. The calling pipeline declares its own params and outputs and passes them explicitly to the included pipeline.
Importing the params block
Redeclaring every param of an included pipeline gets tedious. The params block can be imported as a record type instead:
include {
params as RnaseqParams ;
workflow as NFCORE_RNASEQ
} from './pipelines/rnaseq.nf'
params {
strandedness: String = 'auto' // unique to this pipeline
rnaseq: RnaseqParams // input, aligner, fasta
}
workflow {
main:
rnaseq = NFCORE_RNASEQ( params.rnaseq + record(input: samples) )
}
RnaseqParams is a partial record type: every field is nullable, so a user can provide any rnaseq param as --rnaseq.<name>, the calling pipeline can override specific params with params.rnaseq + record(input: samples), and the NFCORE_RNASEQ() call reports any param that is still missing.
Note the following:
-
A param that the pipeline defaults (
aligner) can be omitted from the record. The pipeline applies its own default when it is called. -
rnaseq.inputis supplied by the dataflow, which overrides any value given by the user. -
rnaseq.fastamust still be provided, but the error surfaces at theNFCORE_RNASEQ()call rather than at launch.
A pipeline receives its params when it is called, so params refers to a single execution of the pipeline. A pipeline can be included under any number of aliases, and each alias resolves its own params. As with a named workflow, a pipeline can be called only once per alias -- include it again under a different alias to call it again.
Best practices
Pipeline inclusion only captures the pipeline script and the modules it includes. It does not capture external context such as the config or the lib directory. As a result, an included pipeline should be written so that it works when included by another pipeline:
-
Pipeline parameters should be declared in the
paramsblock. The config should only declare config params, i.e. params that only affect config settings. -
Project-level assets (
projectDir,bin,lib) should not be used, since the calling pipeline has a different project root. Module-level assets can be safely used through the moduleresources/bundle andmoduleDir. -
Params should be referred to only in the entry workflow and output block. A process or workflow that reads
paramsdirectly should declare an explicit input instead. -
Process configuration (
container,conda,ext) should be specified in the process definition or avoided in favor of process inputs. -
Workflow outputs should be published using the
outputblock, notpublishDir.
None of these constraints are absolute. Each of them can be circumvented by replicating the external context in the calling pipeline. Following them simply makes it easier to include a pipeline with minimal extra work.
Validation
The language server validates each workflow input and output against its declared type. Calling a workflow with an argument whose type does not match its take: declaration, or emitting a value that does not match its emit: declaration, is reported as an error before the pipeline runs.
See also
- Static typing: Overview of static typing and how to enable it.
- Standard types: The standard types available for annotations.
- Operators: The core operators recommended for static typing.