Sculpin's Data Provider and Generator Pipeline: A Developer's Reference

Datum

This page documents the 2 interfaces behind every listing page, tag index and paginated archive in Sculpin: DataProviderInterface and GeneratorInterface. It covers what each one has to implement, how to register one so Sculpin finds it, and the exact order the pipeline runs them in.

This reference covers the Sculpin 3.x series, cross-checked against sculpin/sculpin at commit bff3efa.

The Two Interfaces

DataProviderInterface
1 method: provideData(): array. Whatever array it returns becomes available to templates under data.<name>, once a page opts in with a use: key (see Configuring Front Matter).
GeneratorInterface
1 method: generate(SourceInterface $source): array. It receives the source that named it via generator: and returns an array of new SourceInterface objects, one per page it wants to create.

Both interfaces are deliberately minimal. Everything else, filtering, sorting, taxonomy grouping, pagination math, is implementation detail inside whichever class you write.

Source in code: DataProviderInterface, GeneratorInterface.

Registering One

Implement the interface, register the class as a Symfony service, and tag it. The tag’s alias attribute is the name your front matter will refer to later.

services:
    app.recent_projects_provider:
        class: App\DataProvider\RecentProjectsProvider
        tags:
            - { name: sculpin.data_provider, alias: recent_projects }

    app.archive_generator:
        class: App\Generator\ArchiveGenerator
        tags:
            - { name: sculpin.generator, alias: archive }

Two compiler passes do the actual wiring: they collect every service tagged sculpin.data_provider or sculpin.generator and call registerDataProvider($alias, $service) or registerGenerator($alias, $service) on the corresponding manager. An unregistered alias fails at generate time with a message naming both the missing alias and the offending source file, not at compile time.

Source in code: GeneratorManagerPass, DataProviderManagerPass, both registered in SculpinBundle::build().

How a Page Becomes a Generator

This is the part that isn’t written down anywhere else. During a build, Sculpin walks every updated source exactly once and calls generatorManager->generate($source, $sourceSet) on each. Inside that call:

  1. It reads the source’s generator front matter key. Nothing set, nothing happens.
  2. If set, the source is flagged as a generator (it will be excluded from the final written output later).
  3. Each named generator runs in turn. List more than one, and they chain: the pages the first generator produces become the input the second generator runs against, not the original source again.
  4. Every page a generator produces is flagged as generated and merged back into the build’s source set, where it’s picked up by the normal formatting and writing steps like any other page.

At the very end of the build, Sculpin writes every source to disk except the ones flagged as a generator. That’s the whole mechanism: a generator source is a template that never gets written itself, only the pages it produces do.

Source in code: GeneratorManager::generate(), orchestrated from Sculpin::run().

A Minimal Worked Example

A generator that turns one template into a fixed set of pages, and a data provider that feeds it a title for each:

<?php

class GreetingDataProvider implements DataProviderInterface
{
    public function provideData(): array
    {
        return ['Guten Tag', 'Grüezi', 'Servus'];
    }
}

class GreetingGenerator implements GeneratorInterface
{
    public function __construct(private DataProviderManager $dataProviders) {}

    public function generate(SourceInterface $source): array
    {
        $pages = [];
        foreach ($this->dataProviders->dataProvider('greetings')->provideData() as $greeting) {
            $page = clone $source;
            $page->data()->set('greeting', $greeting);
            $page->data()->set('permalink', 'hello/' . strtolower($greeting));
            $pages[] = $page;
        }
        return $pages;
    }
}

With generator: [greeting] and use: [greetings] in a template’s front matter, this produces 3 pages, 1 per greeting, each with its own page.greeting and its own permalink. The exact cloning and data-mutation approach shown here is illustrative, not the only valid strategy; ProxySourceCollectionDataProvider shows how Sculpin’s own content types bundle solves the same problem at production scale.

What Sculpin Builds With This Machinery

Every dynamic listing in Sculpin, without exception, is these 2 interfaces wearing different clothes:

Feature Data provider Generator
A content type’s items (e.g. data.posts) Yes, 1 per type No generator needed just for this
A taxonomy (e.g. tags) Yes, a <type>_tags mapping Yes, a <type>_tag_index generator, 1 page per tag
Pagination Consumes an existing provider Yes, the built-in pagination generator

The first 2 rows are covered in Configuring sculpin_kernel.yml, the third in Configuring Front Matter. Nothing on either of those pages happens by any mechanism other than the one described here.

A Real Example: posts_categories

Trace the taxonomy row above all the way through, using the built-in posts type and its categories taxonomy.

The name itself comes from a plain string concatenation in SculpinContentTypesExtension::load(): $type.'_'.$taxonomyName, which for the type posts and the taxonomy categories produces posts_categories and registers it under exactly that alias.

The class behind that alias, ProxySourceTaxonomyDataProvider, is both a DataProviderInterface and a Symfony event subscriber on Sculpin::EVENT_BEFORE_RUN, so it populates itself once, before the build runs, rather than on every read:

  1. It asks the posts data provider for every post.
  2. For each post, it reads the raw categories front matter value.
  3. It trims every entry and builds a map of category name to the list of posts carrying it.
  4. It writes that trimmed array back onto each post’s own categories data, which is why page.categories always comes out clean in a template even if the front matter itself had stray whitespace.

provideData() then simply returns that map. Here’s a complete, real page consuming it, categories.html in sculpin-blog-skeleton:

---
layout: default
title: Categories
use:
  - posts_categories
---
<h2>Categories</h2>
<div>
{% for category,posts in data.posts_categories %}
<a href="{{ site.url }}/blog/categories/{{ category|url_encode(true) }}">{{ category }}</a>
{% endfor %}
</div>

The use: key loads the provider as data.posts_categories; the for category,posts in ... loop is Twig destructuring the map this article just traced, key and value at once.

The same loop in SculpinContentTypesExtension that registers posts_categories also registers its generator counterpart in the same breath: ProxySourceTaxonomyIndexGenerator, tagged with the alias posts_category_index (<type>_<taxon-singular>_index). Attach it to a template with generator: [posts_category_index] and it produces 1 generated page per distinct category, each carrying page.category (the category name) and page.category_posts, literally named $reversedName in the source, the list of posts in that category, set directly on the generated page rather than pulled in with use:.

Source in code: ProxySourceTaxonomyIndexGenerator::generate(), naming in SculpinContentTypesExtension.

Recipe: A Sitewide Category Sidebar and Archive Pages

Putting both halves to work: every page gets a sidebar listing every category site-wide, and each category gets its own page listing that category’s posts, newest first.

The sidebar, on every page

use: is read from each source’s own front matter (FormatterManager::buildBaseFormatContext()), so a shared layout only sees data.posts_categories if the page currently being rendered declared it, not because the layout wants it. Repeating use: [posts_categories] on every single source works but isn’t durable. A small event subscriber on the same event the data provider itself uses is:

class InjectCategoriesSubscriber implements EventSubscriberInterface
{
    public static function getSubscribedEvents()
    {
        return [Sculpin::EVENT_BEFORE_RUN => 'beforeRun'];
    }

    public function beforeRun(SourceSetEvent $event)
    {
        foreach ($event->sourceSet()->allSources() as $source) {
            $use = $source->data()->get('use') ?: [];
            $source->data()->set('use', array_unique([...$use, 'posts_categories']));
        }
    }
}

Register it as a service and tag it kernel.event_subscriber, same pattern as „Registering One“ earlier on this page:

services:
    app.inject_categories_subscriber:
        class: InjectCategoriesSubscriber
        tags:
            - { name: kernel.event_subscriber }

With that in place, every page, without exception, can now render:

<aside>
{% for category, posts in data.posts_categories %}
<a href="{{ site.url }}/blog/categories/{{ category|url_encode(true) }}">{{ category }}</a>
{% endfor %}
</aside>

The archive page, one per category

A single template, needing no use: at all:

---
generator: [posts_category_index]
layout: default
---
<h2>{{ page.category }}</h2>
<ul>
{% for post in page.category_posts %}
<li><a href="{{ site.url }}{{ post.url }}">{{ post.title }}</a></li>
{% endfor %}
</ul>

The newest-first order comes for free. posts, like every content type, is sorted by DefaultSorter, which compares 2 posts‘ date and title with the arguments swapped, turning an ascending comparison into a descending one. posts_categories is built by iterating posts in that already-sorted order, and PHP arrays preserve insertion order, so every per-category bucket, and therefore page.category_posts on every generated archive page, comes out newest first without any sorting code of your own.

Source in code: DefaultSorter::sort(), wired as the default in SculpinContentTypesExtension.

Where the archive page’s URL segment actually comes from

generate() never hardcodes /category/ or /categories/ anywhere. It takes the permalink, or failing that the file path, of the template carrying generator: [posts_category_index], splits it into directory and filename, discards the filename, and appends only the processed category value to that directory. The official sculpin-blog-skeleton keeps that template at source/blog/categories/category.html and its tag equivalent at source/blog/tags/tag.html, landing archives at blog/categories/$name and blog/tags/$tag respectively. A site whose category archives instead show up at blog/category/$name, singular, simply keeps that same template one folder name differently, e.g. source/blog/category/index.html. Same mechanism, whatever folder you put the template in.

The skeleton’s own category.html is also the concrete example of the chaining this article described earlier in „How a Page Becomes a Generator“: its front matter reads generator: [posts_category_index, pagination], and the pagination generator that runs second is told provider: page.category_posts, exactly the field the first generator, posts_category_index, just set on that same generated page.

---
layout: default
title: Category Archive
generator: [posts_category_index, pagination]
pagination:
    provider: page.category_posts
---

A sibling file, category.xml, carries only generator: [posts_category_index] (no pagination) and loops page.category_posts|slice(0, 10) directly to build a per-category Atom feed. Its own front matter file extension, .xml, is exactly the indexType this article’s code excerpt below branches on.

$permalink = $source->data()->get('permalink') ?: $source->relativePathname();
$basename = basename($permalink);
$permalink = dirname($permalink);
// ...
$urlTaxon = $this->permalinkStrategyCollection->process($taxon);
$permalink = $permalink.'/'.$urlTaxon.'/';

$name itself, the processed category value, only changes if the taxonomy is configured as an array with a strategies list (see Configuring sculpin_kernel.yml). Left as a plain string, the way the built-in posts type ships, PermalinkStrategyCollectionFactory::create() returns an empty strategy chain, so the category name lands in the URL exactly as written in front matter, unlowercased and unslugified.

Source in code: ProxySourceTaxonomyIndexGenerator::generate().

A Documentation Correction

Sculpin’s own Generators page states that pagination.provider defaults to page.posts. The actual default in the shipped code is data.posts.

if (!isset($config['provider'])) {
    $config['provider'] = 'data.posts';
}

Source in code: PaginationGenerator.

Where to look in the source

Topic Class / file Lines
The 2 interfaces DataProviderInterface, GeneratorInterface full files, both under 30
The posts_categories walkthrough ProxySourceTaxonomyDataProvider, naming in SculpinContentTypesExtension full file, 233-243
The posts_category_index generator and category_posts ProxySourceTaxonomyIndexGenerator full file, naming 244-256 in SculpinContentTypesExtension
Why the archive URL’s parent folder is whatever you name it ProxySourceTaxonomyIndexGenerator::generate() 53-64
Why the category value is unslugified by default PermalinkStrategyCollectionFactory full file
Why use: needs a subscriber to reach a shared layout FormatterManager::buildBaseFormatContext() 63-89
Default newest-first sorting DefaultSorter full file
Registration wiring GeneratorManagerPass, DataProviderManagerPass full files, both under 42
The chaining and skip-on-write logic GeneratorManager::generate() 65-111
Where generate() is called during a build Sculpin::run() 102-135, 191-198
A production-scale data provider ProxySourceCollectionDataProvider full file
The pagination.provider default mismatch PaginationGenerator 58-67

Why This Needed Writing Down

None of the above is officially documented in one place. Custom Types uses both interfaces constantly without naming them; Generators documents Pagination but not the mechanism itself; there is no „Data Providers“ page at all. A Sculpin maintainer confirms as much directly: issue #348 on sculpin/sculpin states that the custom types section „should explain in more detail how providers and generators work.“

Sources

  1. Generators – sculpin.io
  2. Content Types: Custom Types – sculpin.io
  3. Issue #348 – sculpin/sculpin, GitHub
  4. Creating virtual pages with Sculpin – Matthias Noback
  5. sculpin/sculpin source code (commit bff3efa) – sculpin, GitHub
  6. sculpin-blog-skeleton – sculpin, GitHub

License

This document is © Robert Wetzlmayr and licensed under Creative Commons Attribution-ShareAlike 4.0 International. Reuse, adaptation and redistribution are welcome, including commercially, as long as attribution is given and any adapted version carries the same license. Sculpin itself remains under its own MIT license; this notice covers the text and structure of this page only, not the software it describes.

Fork me!


Kategorien Sculpin