This page documents the 2 interfaces behind every listing page, tag index and paginated archive in Sculpin: DataProviderInterface and GeneratorInterface. It covers what each one has to implement, how to register one so Sculpin finds it, and the exact order the pipeline runs them in.
This reference covers the Sculpin 3.x series, cross-checked against sculpin/sculpin at commit bff3efa.
The Two Interfaces
- DataProviderInterface
- 1 method:
provideData(): array. Whatever array it returns becomes available to templates underdata.<name>, once a page opts in with ause:key (see Configuring Front Matter). - GeneratorInterface
- 1 method:
generate(SourceInterface $source): array. It receives the source that named it viagenerator:and returns an array of newSourceInterfaceobjects, one per page it wants to create.
Both interfaces are deliberately minimal. Everything else, filtering, sorting, taxonomy grouping, pagination math, is implementation detail inside whichever class you write.
Source in code: DataProviderInterface, GeneratorInterface.
Registering One
Implement the interface, register the class as a Symfony service, and tag it. The tag’s alias attribute is the name your front matter will refer to later.
services:
app.recent_projects_provider:
class: App\DataProvider\RecentProjectsProvider
tags:
- { name: sculpin.data_provider, alias: recent_projects }
app.archive_generator:
class: App\Generator\ArchiveGenerator
tags:
- { name: sculpin.generator, alias: archive }
Two compiler passes do the actual wiring: they collect every service tagged sculpin.data_provider or sculpin.generator and call registerDataProvider($alias, $service) or registerGenerator($alias, $service) on the corresponding manager. An unregistered alias fails at generate time with a message naming both the missing alias and the offending source file, not at compile time.
Source in code: GeneratorManagerPass, DataProviderManagerPass, both registered in SculpinBundle::build().
How a Page Becomes a Generator
This is the part that isn’t written down anywhere else. During a build, Sculpin walks every updated source exactly once and calls generatorManager->generate($source, $sourceSet) on each. Inside that call:
- It reads the source’s
generatorfront matter key. Nothing set, nothing happens. - If set, the source is flagged as a generator (it will be excluded from the final written output later).
- Each named generator runs in turn. List more than one, and they chain: the pages the first generator produces become the input the second generator runs against, not the original source again.
- Every page a generator produces is flagged as generated and merged back into the build’s source set, where it’s picked up by the normal formatting and writing steps like any other page.
At the very end of the build, Sculpin writes every source to disk except the ones flagged as a generator. That’s the whole mechanism: a generator source is a template that never gets written itself, only the pages it produces do.
Source in code: GeneratorManager::generate(), orchestrated from Sculpin::run().
A Minimal Worked Example
A generator that turns one template into a fixed set of pages, and a data provider that feeds it a title for each:
<?php
class GreetingDataProvider implements DataProviderInterface
{
public function provideData(): array
{
return ['Guten Tag', 'Grüezi', 'Servus'];
}
}
class GreetingGenerator implements GeneratorInterface
{
public function __construct(private DataProviderManager $dataProviders) {}
public function generate(SourceInterface $source): array
{
$pages = [];
foreach ($this->dataProviders->dataProvider('greetings')->provideData() as $greeting) {
$page = clone $source;
$page->data()->set('greeting', $greeting);
$page->data()->set('permalink', 'hello/' . strtolower($greeting));
$pages[] = $page;
}
return $pages;
}
}
With generator: [greeting] and use: [greetings] in a template’s front matter, this produces 3 pages, 1 per greeting, each with its own page.greeting and its own permalink. The exact cloning and data-mutation approach shown here is illustrative, not the only valid strategy; ProxySourceCollectionDataProvider shows how Sculpin’s own content types bundle solves the same problem at production scale.
What Sculpin Builds With This Machinery
Every dynamic listing in Sculpin, without exception, is these 2 interfaces wearing different clothes:
| Feature | Data provider | Generator |
|---|---|---|
A content type’s items (e.g. data.posts) |
Yes, 1 per type | No generator needed just for this |
| A taxonomy (e.g. tags) | Yes, a <type>_tags mapping |
Yes, a <type>_tag_index generator, 1 page per tag |
| Pagination | Consumes an existing provider | Yes, the built-in pagination generator |
The first 2 rows are covered in Configuring sculpin_kernel.yml, the third in Configuring Front Matter. Nothing on either of those pages happens by any mechanism other than the one described here.
A Real Example: posts_categories
Trace the taxonomy row above all the way through, using the built-in posts type and its categories taxonomy.
The name itself comes from a plain string concatenation in SculpinContentTypesExtension::load(): $type.'_'.$taxonomyName, which for the type posts and the taxonomy categories produces posts_categories and registers it under exactly that alias.
The class behind that alias, ProxySourceTaxonomyDataProvider, is both a DataProviderInterface and a Symfony event subscriber on Sculpin::EVENT_BEFORE_RUN, so it populates itself once, before the build runs, rather than on every read:
- It asks the
postsdata provider for every post. - For each post, it reads the raw
categoriesfront matter value. - It trims every entry and builds a map of category name to the list of posts carrying it.
- It writes that trimmed array back onto each post’s own
categoriesdata, which is whypage.categoriesalways comes out clean in a template even if the front matter itself had stray whitespace.
provideData() then simply returns that map. Here’s a complete, real page consuming it, categories.html in sculpin-blog-skeleton:
---
layout: default
title: Categories
use:
- posts_categories
---
<h2>Categories</h2>
<div>
{% for category,posts in data.posts_categories %}
<a href="{{ site.url }}/blog/categories/{{ category|url_encode(true) }}">{{ category }}</a>
{% endfor %}
</div>
The use: key loads the provider as data.posts_categories; the for category,posts in ... loop is Twig destructuring the map this article just traced, key and value at once.
The same loop in SculpinContentTypesExtension that registers posts_categories also registers its generator counterpart in the same breath: ProxySourceTaxonomyIndexGenerator, tagged with the alias posts_category_index (<type>_<taxon-singular>_index). Attach it to a template with generator: [posts_category_index] and it produces 1 generated page per distinct category, each carrying page.category (the category name) and page.category_posts, literally named $reversedName in the source, the list of posts in that category, set directly on the generated page rather than pulled in with use:.
Source in code: ProxySourceTaxonomyIndexGenerator::generate(), naming in SculpinContentTypesExtension.
Recipe: A Sitewide Category Sidebar and Archive Pages
Putting both halves to work: every page gets a sidebar listing every category site-wide, and each category gets its own page listing that category’s posts, newest first.
The sidebar, on every page
use: is read from each source’s own front matter (FormatterManager::buildBaseFormatContext()), so a shared layout only sees data.posts_categories if the page currently being rendered declared it, not because the layout wants it. Repeating use: [posts_categories] on every single source works but isn’t durable. A small event subscriber on the same event the data provider itself uses is:
class InjectCategoriesSubscriber implements EventSubscriberInterface
{
public static function getSubscribedEvents()
{
return [Sculpin::EVENT_BEFORE_RUN => 'beforeRun'];
}
public function beforeRun(SourceSetEvent $event)
{
foreach ($event->sourceSet()->allSources() as $source) {
$use = $source->data()->get('use') ?: [];
$source->data()->set('use', array_unique([...$use, 'posts_categories']));
}
}
}
Register it as a service and tag it kernel.event_subscriber, same pattern as „Registering One“ earlier on this page:
services:
app.inject_categories_subscriber:
class: InjectCategoriesSubscriber
tags:
- { name: kernel.event_subscriber }
With that in place, every page, without exception, can now render:
<aside>
{% for category, posts in data.posts_categories %}
<a href="{{ site.url }}/blog/categories/{{ category|url_encode(true) }}">{{ category }}</a>
{% endfor %}
</aside>
The archive page, one per category
A single template, needing no use: at all:
---
generator: [posts_category_index]
layout: default
---
<h2>{{ page.category }}</h2>
<ul>
{% for post in page.category_posts %}
<li><a href="{{ site.url }}{{ post.url }}">{{ post.title }}</a></li>
{% endfor %}
</ul>
The newest-first order comes for free. posts, like every content type, is sorted by DefaultSorter, which compares 2 posts‘ date and title with the arguments swapped, turning an ascending comparison into a descending one. posts_categories is built by iterating posts in that already-sorted order, and PHP arrays preserve insertion order, so every per-category bucket, and therefore page.category_posts on every generated archive page, comes out newest first without any sorting code of your own.
Source in code: DefaultSorter::sort(), wired as the default in SculpinContentTypesExtension.
Where the archive page’s URL segment actually comes from
generate() never hardcodes /category/ or /categories/ anywhere. It takes the permalink, or failing that the file path, of the template carrying generator: [posts_category_index], splits it into directory and filename, discards the filename, and appends only the processed category value to that directory. The official sculpin-blog-skeleton keeps that template at source/blog/categories/category.html and its tag equivalent at source/blog/tags/tag.html, landing archives at blog/categories/$name and blog/tags/$tag respectively. A site whose category archives instead show up at blog/category/$name, singular, simply keeps that same template one folder name differently, e.g. source/blog/category/index.html. Same mechanism, whatever folder you put the template in.
The skeleton’s own category.html is also the concrete example of the chaining this article described earlier in „How a Page Becomes a Generator“: its front matter reads generator: [posts_category_index, pagination], and the pagination generator that runs second is told provider: page.category_posts, exactly the field the first generator, posts_category_index, just set on that same generated page.
---
layout: default
title: Category Archive
generator: [posts_category_index, pagination]
pagination:
provider: page.category_posts
---
A sibling file, category.xml, carries only generator: [posts_category_index] (no pagination) and loops page.category_posts|slice(0, 10) directly to build a per-category Atom feed. Its own front matter file extension, .xml, is exactly the indexType this article’s code excerpt below branches on.
$permalink = $source->data()->get('permalink') ?: $source->relativePathname();
$basename = basename($permalink);
$permalink = dirname($permalink);
// ...
$urlTaxon = $this->permalinkStrategyCollection->process($taxon);
$permalink = $permalink.'/'.$urlTaxon.'/';
$name itself, the processed category value, only changes if the taxonomy is configured as an array with a strategies list (see Configuring sculpin_kernel.yml). Left as a plain string, the way the built-in posts type ships, PermalinkStrategyCollectionFactory::create() returns an empty strategy chain, so the category name lands in the URL exactly as written in front matter, unlowercased and unslugified.
Source in code: ProxySourceTaxonomyIndexGenerator::generate().
A Documentation Correction
Sculpin’s own Generators page states that pagination.provider defaults to page.posts. The actual default in the shipped code is data.posts.
if (!isset($config['provider'])) {
$config['provider'] = 'data.posts';
}
Source in code: PaginationGenerator.
Where to look in the source
| Topic | Class / file | Lines |
|---|---|---|
| The 2 interfaces | DataProviderInterface, GeneratorInterface | full files, both under 30 |
The posts_categories walkthrough |
ProxySourceTaxonomyDataProvider, naming in SculpinContentTypesExtension | full file, 233-243 |
The posts_category_index generator and category_posts |
ProxySourceTaxonomyIndexGenerator | full file, naming 244-256 in SculpinContentTypesExtension |
| Why the archive URL’s parent folder is whatever you name it | ProxySourceTaxonomyIndexGenerator::generate() | 53-64 |
| Why the category value is unslugified by default | PermalinkStrategyCollectionFactory | full file |
Why use: needs a subscriber to reach a shared layout |
FormatterManager::buildBaseFormatContext() | 63-89 |
| Default newest-first sorting | DefaultSorter | full file |
| Registration wiring | GeneratorManagerPass, DataProviderManagerPass | full files, both under 42 |
| The chaining and skip-on-write logic | GeneratorManager::generate() | 65-111 |
| Where generate() is called during a build | Sculpin::run() | 102-135, 191-198 |
| A production-scale data provider | ProxySourceCollectionDataProvider | full file |
| The pagination.provider default mismatch | PaginationGenerator | 58-67 |
Why This Needed Writing Down
None of the above is officially documented in one place. Custom Types uses both interfaces constantly without naming them; Generators documents Pagination but not the mechanism itself; there is no „Data Providers“ page at all. A Sculpin maintainer confirms as much directly: issue #348 on sculpin/sculpin states that the custom types section „should explain in more detail how providers and generators work.“
Sources
- Generators – sculpin.io
- Content Types: Custom Types – sculpin.io
- Issue #348 – sculpin/sculpin, GitHub
- Creating virtual pages with Sculpin – Matthias Noback
- sculpin/sculpin source code (commit
bff3efa) – sculpin, GitHub - sculpin-blog-skeleton – sculpin, GitHub
License
This document is © Robert Wetzlmayr and licensed under Creative Commons Attribution-ShareAlike 4.0 International. Reuse, adaptation and redistribution are welcome, including commercially, as long as attribution is given and any adapted version carries the same license. Sculpin itself remains under its own MIT license; this notice covers the text and structure of this page only, not the software it describes.