Technical documentation for the vojtamaur-web project
1. Project overview
vojtamaur-web is a static website built with Astro. Content is managed as files in the repository and converted into static output during the build process. The project does not use a CMS or a database. The source of truth is the repository containing the source code, content, and static assets.
The project is divided into the following main content sections:
- Volná tvorba
- Výstavy
- Cestování
- Propagační videa
- O mně
- Kontakt
The architecture is based on the following components:
- Astro as the static site generator
- MDX as the authoring format for content
- Content Collections for metadata validation and typing
- Components for reusable content blocks and layouts
- Static assets in
public/for images, PDFs, demos, and other files - A generated Gemini/Gopher text edition in
dist-gemini/as a text-oriented hypertext edition of the finished website, usable as both a Gemini capsule and a Gopher directory tree
This model makes it easier to version content, archive build outputs, and potentially migrate the project to another environment without relying on a database runtime.
2. Project structure
2.1 Directory structure
The current project structure is divided into two main parts:
src/– project source filespublic/– static assets copied unchanged during the build
Provided structure:
public/
demos/
files/
images/
keys/
src/
components/
content/
content.config.ts
env.d.ts
layouts/
lib/
pages/
styles/Generated output directories are kept outside the source tree:
dist/ # normal web or portable build output
dist-arweave/ # derived Arweave / Permaweb deployment output
dist-gemini/ # generated Gemini capsuleThese directories are build artifacts and should not be edited as source content.
2.2 Meaning of the main directories
public/
Contains static files that are simply copied into the output during the build:
public/images/– article images, thumbnails, and other visual contentpublic/files/– PDFs and other downloadable or embeddable filespublic/demos/– standalone HTML/JS demos and legacy static pagespublic/keys/– public OpenPGP material copied unchanged into the build, including the public key and its fingerprint
src/content/
Project content files. In the current configuration, the following are mainly used:
src/content/posts/– articlessrc/content/videos/– metadata for promotional videos
src/components/
Reusable components for working with content and listings:
ImageFigure.astroMediaRow.astroEmbed.astroPostTileGrid.astroOpenPgpContact.astroOpenPgpFingerprint.astroHeader.astro
src/layouts/
Layouts for individual page types:
BaseLayout.astroPostLayout.astro- potentially specialized layouts such as
TravelLayout.astroandExhibitionLayout.astro
src/pages/
Application routes. Includes the homepage, category pages, and dynamic article routing.
src/styles/
Global and optionally other style files.
src/lib/
Helper utilities and shared logic used across the project.
source-bundle/
Templates used by scripts/generate-source-bundle.mjs when creating the reconstructable source package. This directory contains the Python asset downloader and the reconstruction README that are copied into the generated source ZIP.
BUILD_HASH_HISTORY.txt
The project-root file is the canonical cumulative history of successfully signed builds. It is source data to preserve in version control. Builds receive a snapshot of the previous records; signing appends the current record only to the source file (section 9.6.1).
3. Key configuration files
astro.config.mjs
The project uses two Astro build modes. The configuration switches behavior according to the BUILD_TARGET variable. In the standard web build it uses trailingSlash: "always" and build.format: "directory". In the portable file-based build it uses trailingSlash: "never" and build.format: "file".
This produces two primary Astro output types:
- web build – suitable for standard hosting
- portable file-based build – suitable, for example, for offline use, archiving, or transfer on external media
A third deployment artifact is derived from the standard web build:
- Arweave / Permaweb build – generated into
dist-arweave/by copying the finisheddist/output and rewriting root-relative paths so that the site can work under an Arweave manifest transaction path
The Arweave build is not a separate Astro BUILD_TARGET. It is a postprocessed version of the normal web build.
A fourth publication artifact is also derived from the finished standard web build:
- Gemini/Gopher text edition – generated into
dist-gemini/byscripts/generate-gemini-capsule.mjs
The Gemini/Gopher text edition is not an Astro BUILD_TARGET and is not placed inside dist/. It reads the finished HTML in dist/ so that the English postprocess is already reflected in the exported content, while article frontmatter remains available for section membership, dates, slugs, and ordering.
package.json
Basic project workflow:
npm run dev
npm run build:web
npm run build:web:signed
npm run build:web:translate
npm run build:web:translate:signed
npm run build:translate:signed
npm run build:web:strict
npm run build:web:strict:signed
npm run build:web:prune:dry
npm run build:web:prune
npm run preview
npm run build:usb
npm run build:usb:signed
npm run build:usb:translate
npm run build:usb:translate:signed
npm run build:usb:strict
npm run build:usb:strict:signed
npm run build:usb:prune:dry
npm run build:usb:prune
npm run build:arweave
npm run build:arweave:signed
npm run generate:all-posts
npm run generate:gemini
npm run build:gemini
npm run generate:source-bundle
npm run generate:sitemap
npm run generate:integrity
npm run generate:integrity:arweave
npm run sign:build
npm run sign:build:arweave
npm run test:integrity
npm run export:epub
npm run export:epub:metaweb
npm run export:pdf
npm run export:pdf:metaweb
npm run export:sstvMeaning of the main scripts:
npm run dev– starts the Astro development servernpm run build:web– creates the standard production web build, applies the English postprocess using existing cache only, generatesdist/ALL_POSTS.txtanddist/ALL_POSTS.json, generates the reconstructable source package atdist/source/vojtamaur-web-source.zip, enriches the sitemap with images and PDF/TXT files, and writes unsigned build integrity metadata fordist/npm run build:web:signed– runs the standard web build and then creates and verifies the detached OpenPGP signature fordist/SHA256SUMS.txtnpm run build:signed– alias fornpm run build:web:signednpm run build:web:translate– creates the web build, fills missing English translation cache entries through DeepL, generatesdist/ALL_POSTS.txtanddist/ALL_POSTS.json, generates the reconstructable source package, enriches the sitemap with images and PDF/TXT files, writes unsigned build integrity metadata fordist/, and then generates the Gemini/Gopher text edition indist-gemini/npm run build:web:translate:signed– runs the translating web build, including Gemini/Gopher text-edition generation, and then signs and verifies the checksum manifest fordist/; the detached signature does not cover the separatedist-gemini/directorynpm run build:translate:signed– alias fornpm run build:web:translate:signednpm run build:web:strict– creates the web build, fails if an English translation cache entry is missing, generatesdist/ALL_POSTS.txtanddist/ALL_POSTS.json, generates the reconstructable source package, enriches the sitemap with images and PDF/TXT files, and writes unsigned build integrity metadata fordist/npm run build:web:strict:signed– runs the strict web build and then signs and verifies its checksum manifestnpm run build:web:prune:dry– runs the strict web build and prints unused English translation cache files that would be deleted, without deleting themnpm run build:web:prune– runs the strict web build and deletes unused, unprotected English translation cache filesnpm run preview– local preview of the buildnpm run build:usb– creates a portable file-based build, applies the English postprocess, generatesdist/ALL_POSTS.txtanddist/ALL_POSTS.json, generates the reconstructable source package, rewrites paths for offline use, and writes unsigned build integrity metadata fordist/npm run build:usb:signed– runs the portable build and then signs and verifies its checksum manifestnpm run build:usb:translate– creates a portable file-based build, fills missing English translation cache entries, generatesdist/ALL_POSTS.txtanddist/ALL_POSTS.json, generates the reconstructable source package, rewrites paths for offline use, and writes unsigned build integrity metadata fordist/npm run build:usb:translate:signed– runs the translating portable build, signs and verifies its manifest, and records its hash in the source history with typeusbnpm run build:usb:strict– creates a strict portable file-based build and fails if an English translation cache entry is missingnpm run build:usb:strict:signed– runs the strict portable build and then signs and verifies its checksum manifestnpm run build:usb:prune:dry– runs the strict portable build and prints unused English translation cache files that would be deleted, without deleting themnpm run build:usb:prune– runs the strict portable build and deletes unused, unprotected English translation cache filesnpm run build:arweave– creates the strict standard web build, createsdist-arweave/as a Permaweb-compatible deployment copy with rewritten paths, and then writes separate unsigned build integrity metadata fordist-arweave/npm run build:arweave:signed– runs the Arweave build and then signs and verifiesdist-arweave/SHA256SUMS.txtnpm run generate:all-posts– generatesdist/ALL_POSTS.txtand the structureddist/ALL_POSTS.jsonfrom the finished build output, and embeds the text export intodist/404.htmlnpm run generate:gemini– replacesdist-gemini/with a freshly generated bilingual Gemini/Gopher text edition based on the existing finisheddist/build and the post frontmatternpm run build:gemini– runs the strict standard web build and then generatesdist-gemini/; this is the standalone complete Gemini/Gopher text-edition build commandnpm run generate:source-bundle– generatesdist/source/vojtamaur-web-source.zip, a small reconstructable source package with a media manifest and an asset downloadernpm run generate:sitemap– enriches existing web sitemaps with page-associated images and public PDF/TXT files after content generation and before integrity generation; USB builds omit this stepnpm run generate:integrity– generatesSHA256SUMS.txt,BUILD_SHA256.txt,integrity.json, and an unsignedSIGNING_STATUS.txtfordist/; any stale detached signature is removednpm run generate:integrity:arweave– performs the same integrity pass fordist-arweave/after the Arweave copy-and-rewrite stepnpm run sign:build– signs the existingdist/SHA256SUMS.txt, verifies the detached signature immediately, and updates the informational signing status and JSON metadatanpm run sign:build:arweave– performs the same explicit local signing step fordist-arweave/npm run test:integrity– checks history formatting, archive snapshots, signing and failure handling in isolated temporary projects; the signing tests use a temporary GPG key and are skipped if GnuPG is unavailablenpm run export:epub– manually creates separate Czech and English reflowable EPUB books from the finished article pages indist/, including a generated title/index section; this command is not part of the normal build pipelinenpm run export:epub:metaweb– manually creates one bilingual Metaweb EPUB with both article languages, its image appendix, archive map, technical documentation, and build identity files; this command is not part of any build pipelinenpm run export:pdf– manually exports article pages from the finisheddist/build to archival PDF files; this command is not part of the normal build pipelinenpm run export:pdf:metaweb– manually creates one archival PDF containing the bilingual metaweb article and its directly related preservation files; this command is not part of any build pipelinenpm run export:sstv– manually renders indexed articles from the finisheddist/build to numbered PNG frames at native SSTV mode dimensions; defaults to Czech PD120 output underexports/sstv/pd120/, does not run a build or generate audio, and is not part of any build pipeline (section 9.6.8)npm run export:morse– manually converts the existingdist/ALL_POSTS.jsonthrough the compact text serializer into an ASCII Morse edition underexports/; supports--lang cs|en|both, defaults to one bilingual file, and is also included inexport-all.bat(section 9.6.11)
All signed build variants and both standalone signing commands use the same history mechanism: after successful signature verification and status/metadata writes, they record the current build in the project-root BUILD_HASH_HISTORY.txt. Unsigned build commands do not append records. The source-bundle step supplies the history snapshot required by the subsequent integrity pass; see section 9.6.1.
content.config.ts
Defines content collections and metadata schemas using Zod validation. The project uses at least the following collections:
postsvideos
4. Content Collections
4.1 posts collection
Articles are stored as .mdx files and validated against the schema in content.config.ts.
Required metadata:
titleslugsectiondatethumbnailthumbnailAlt
Optional metadata:
excerptdraft
Section-specific metadata:
For section: "vystavy"
dateFromdateTocityvenueexhibition
For section: "cestovani"
yearmedia
4.2 videos collection
Used for the Propagační videa section. Contains metadata for external YouTube videos. Typically:
titleurlthumbnailthumbnailAltdraft
5. Layout logic
5.1 PostLayout.astro
PostLayout.astro wraps article content in the main layout and creates a shared wrapper for article pages.
5.2 Dynamic routing through [slug].astro
The [slug].astro file is the central route for content from the posts collection. It handles:
- loading articles via
getCollection - generating static paths based on
slug - rendering specific content via
render(post)
5.3 Conditional rendering by section
Different sections have different meta blocks:
- Volná tvorba – title + formatted month and year
- Výstavy – title + event date + city + venue + exhibition title
- Cestování – title + supplementary metadata such as medium
5.4 Date formatting
In the Volná tvorba section, the date is displayed as month and year, for example:
duben 2026Internally, the standard date field is still used for sorting.
5.5 Sorting
Articles are sorted by date. This also applies in cases where the UI does not display the exact day, but only the month and year.
5.5.1 Legacy date migration note
This project was created as a replacement for an older website with fragile infrastructure (outdated PHP, WordPress, unmaintained plugins, and dependence on third-party systems).
That legacy site displayed only the month and year for many articles in the Volná tvorba section (for example duben 2020). During migration, the exact original day was often no longer recoverable. In such cases, the date field was normalized to the first day of the given month (for example 2020-04-01) in order to preserve sorting behavior.
This means that for part of the legacy content, the stored day may be approximate and should be understood as a technical migration value rather than an exact historical publication date.
Articles added after April 2026 use the real day in the date field whenever that information is available.
6. Adding and managing content
6.1 Adding a new article
A new article is added by creating a new .mdx file in:
src/content/posts/The file must contain valid frontmatter according to the posts schema.
6.2 Required and optional metadata
Shared required metadata
titleslugsectiondatethumbnailthumbnailAlt
Shared optional metadata
excerptdraft
Metadata for Výstavy
dateFromdateTocityvenueexhibition
Metadata for Cestování
yearmedia
6.3 Thumbnails
Each article uses:
thumbnail– path to the thumbnailthumbnailAlt– thumbnail alt text
Thumbnails are used in article listings, on the homepage, and in individual sections.
6.3.1 Public asset metadata and privacy
Files stored in public/ are copied into the build output unchanged during the Astro phase. This includes images in public/images/ and downloadable or embeddable files in public/files/. The later source-bundle step intentionally replaces only dist/images/kurt-godel-rat.jpg with a JPEG/ZIP carrier, while retaining the complete original public JPEG as its byte-for-byte prefix. Its embedded metadata therefore remains unchanged. Embedded metadata in all public files must be treated as public data once the files are committed and published.
This applies especially to:
- EXIF camera metadata
- GPS coordinates
- XMP and IPTC metadata
- software and editing history
- PDF author, creator, producer, and title fields
- internal document IDs and document lineage metadata
- embedded comments, captions, keywords, or other hidden text fields
The project uses a separate metadata audit script for checking public assets:
python scripts/audit-public-metadata.py --exiftool "D:\Program Files\exiftool\exiftool.exe"ExifTool path may need to be adjusted depending on the local installation.
The script checks:
public/images/public/files/
It reports privacy-relevant metadata such as camera model, GPS data, author fields, software history, PDF metadata, document IDs, and embedded comments.
Default rule:
- files committed to the public repository should not contain unintended metadata
- original files with full metadata should be kept outside the public repository, if they need to be preserved
- metadata should only be kept when it is intentionally part of the work itself
- intentional metadata exceptions should be handled through the script allowlist
For normal publishing, the recommended workflow is:
- keep the original file in a private archive, if needed
- export or copy a public version of the file
- run the metadata audit
- strip unintended metadata from the public version
- run the metadata audit again
- commit only the cleaned public version into
public/
The audit and stripping process is not part of the Astro build. It is a separate maintenance step. This is intentional: metadata should be removed before commit, not only from the generated dist/ output, because the source repository, mirrors, releases, and archival snapshots may preserve the original files.
6.4 draft
The draft: true field excludes the article from the public listing and generated paths. It is used for content that is in progress or temporarily hidden.
6.5 Components used in content
Besides standard Markdown, article content can also use the following components:
ImageFigureMediaRowEmbed
These components must be explicitly imported in the MDX file.
6.5.1 EntryPoints.astro
The EntryPoints.astro component is used in the meta article to render the main archive entry points.
It reads the public/ARCHIVE.txt file, extracts the section between the markers === ENTRY POINTS START === and === ENTRY POINTS END ===, and displays it inside a <pre> block.
This ensures that the list of primary locations and snapshots is maintained in a single source of truth and does not need to be manually duplicated in the article content.
6.6 Example frontmatter
Volná tvorba
---
title: "Název článku"
slug: "nazev-clanku"
section: "volna-tvorba"
date: 2026-04-19
thumbnail: "/vojtamaur-web/images/nazev-clanku-nahled.jpg"
thumbnailAlt: "Náhled článku"
excerpt: ""
draft: false
---Výstavy
---
title: "Recamánova struktura"
slug: "vystavy-recamanova-struktura"
section: "vystavy"
date: 2024-01-01
thumbnail: "/vojtamaur-web/images/vystavy-recamanova-struktura-nahled.jpg"
thumbnailAlt: "Recamánova struktura"
dateFrom: "1. 1. 2024"
dateTo: "31. 1. 2024"
city: "Jindřichův Hradec"
venue: "Muzeum fotografie a moderních obrazových médií"
exhibition: "Obrazy nad čísly"
draft: false
---Cestování
---
title: "Itálie - Benátky 2019"
slug: "cestovani-italie-benatky-2019"
section: "cestovani"
date: 2019-01-01
thumbnail: "/vojtamaur-web/images/cestovani-italie-benatky-2019-nahled.jpg"
thumbnailAlt: "Itálie - Benátky 2019"
year: "2019"
media: "Fotografie"
draft: false
---6.7 Representative example of content
The following file was provided as a representative example of an MDX article:
recamanova-posloupnost-zelvi-grafice.mdxThis file is suitable as a reference example of content structure, frontmatter, and component usage.
7. Media components
7.1 ImageFigure.astro
A component for standalone images with support for the following parameters:
srcaltcaptionwidthalignborderedlink
Supported width variants:
fullhalfone-third
Supported alignment options:
leftcenterright
link defaults to true: clicking the image opens its source file in the current tab. Setting link={false} renders the image without a wrapping link.
7.2 MediaRow.astro
A component for arranging multiple items in a row. Supports the following types:
imagepdftextempty
Use cases:
- galleries with multiple images
- combinations of images and PDFs
- combinations of visual and text blocks in a single grid
- intentional empty cells used to keep a row layout without creating fake media
The component also supports a bordered variant.
The component-level link parameter defaults to true and opens image items in the current tab. Setting link={false} removes the image links for the whole row. Items of type pdf remain embedded in an iframe; this setting only controls image links.
Image items support alt text. The value is rendered directly into the image alt attribute. If alt is omitted, the component renders an empty alt attribute.
Example:
<MediaRow
bordered
items={[
{
type: "image",
src: "/vojtamaur-web/images/example-1.jpg",
alt: "Popis prvního obrázku"
},
{
type: "image",
src: "/vojtamaur-web/images/example-2.jpg",
alt: "Popis druhého obrázku"
},
{
type: "pdf",
src: "/vojtamaur-web/files/example.pdf"
},
{
type: "text",
content: "Textový blok v řádku médií."
}
]}
/>The alt value applies only to type: "image" items. It is not automatically translated by the English postprocess.
Use type: "empty" when a row should keep an intentionally blank cell. Do not create placeholder image items with an empty src, because that produces broken media in the rendered HTML and can also leak into preservation exports such as ALL_POSTS.txt.
Example with one image and two empty cells:
<MediaRow
items={[
{
type: "image",
src: "/vojtamaur-web/images/nahodna-cisla-zvolena-clovekem-obr-1.jpg",
alt: "Graf náhodných čísel zvolených počítačem"
},
{ type: "empty" },
{ type: "empty" }
]}
/>Empty items render only an empty .media-row__item cell marked with aria-hidden="true". They do not render an image, link, iframe, or text block.
7.3 Embed.astro
A generic wrapper for embedded iframe content. It is used, for example, for:
- YouTube
- Google Maps
- Sketchfab
- local HTML demos
Supported parameters:
srcratiokindwidthalign
For YouTube embeds, prefer the privacy-enhanced youtube-nocookie.com domain instead of the standard youtube.com embed URL:
<Embed
src="https://www.youtube-nocookie.com/embed/6wDN62Xq3pA"
kind="youtube"
/>This reduces unnecessary YouTube cookie use, although the iframe still loads content from a third-party service.
7.4 Edge cases
PDF in MediaRow
When the grid collapses, a PDF iframe may require special adjustment of height or aspect ratio. This was handled using CSS for .media-row__pdf.
Responsive grid collapse
Listings and media layouts have multiple states depending on screen width. Some elements, such as the “show all” button or PDF embeds, required separate behavior adjustments for 3, 2, and 1 column layouts.
7.5 Link behavior
Site links open in the current tab by default, including linked images in ImageFigure and MediaRow, text links to images and files, and archive entry points. A linked image and a text link to the same file therefore behave consistently. File links do not force a download: the browser displays supported formats or handles the file according to its settings and the server response.
The intentional exceptions are the YouTube video tiles on the homepage and the Propagační videa playlist links (including the section heading and the show-all tile). They continue to open in a new tab in both language versions.
This policy uses normal HTML links and requires no client-side JavaScript. Visitors can still choose to open web and file links in a new tab. It also applies to the portable/offline and Arweave builds; controls inside third-party embeds are managed by their providers. Links using protocols such as mailto:, tel:, gemini:, or gopher: are handled by the browser or a registered application rather than being guaranteed to open in a browser tab.
8. Homepage architecture
The homepage combines content from multiple parts of the website.
8.1 Dynamic content loading
Both Czech and English homepages display the latest items by section:
- Volná tvorba: 9 items (three rows on desktop)
- Výstavy: 9 items (three rows on desktop)
- Cestování: 3 items (one row on desktop)
The grids retain their responsive two-column and single-column layouts on smaller screens.
8.2 “Show all” button
If a given section contains more items than the number displayed on the homepage, a “show all” tile is added.
8.3 Propagační videa
The Propagační videa section uses a visual model similar to article listings, but the items link to external YouTube URLs. The listing is based on the videos collection. The video tiles, show-all playlist tile, and playlist section heading intentionally open in a new tab in both language versions (section 7.5).
8.4 Clickable section headings
Section headings on the homepage are clickable and serve as quick navigation to the relevant categories or an external playlist.
8.5 O mně and Kontakt
The homepage also contains specialized content blocks outside the standard article system:
- an O mně block with text and a Sketchfab iframe
- a Kontakt block with phone number, email, and the public OpenPGP identity
8.6 OpenPGP contact block
Both homepage variants use the shared component:
<OpenPgpContact lang={lang} />The component is rendered from both src/pages/index.astro and src/pages/en/index.astro. It provides short Czech or English labels while keeping the cryptographic material shared between both language versions.
The component reads these files during the Astro build:
public/keys/vojta-maur-openpgp.asc
public/keys/vojta-maur-openpgp-fingerprint.txtOpenPgpContact.astro therefore does not duplicate the armored public key or fingerprint inside either homepage file. The public key remains directly visible in the rendered HTML, while the fingerprint is normalized to a single space-separated line. Both values are marked with translate="no" so that the translation postprocess cannot modify them.
8.7 Gemini homepage parity
The generated Gemini homepage intentionally follows the content logic of the HTML homepage rather than reducing the capsule to a single archive dump.
For Volná tvorba / Personal Work, Výstavy / Exhibitions, and Cestování / Travel, the capsule shows the same entries in the same order as the built homepage: up to nine for Personal Work and Exhibitions, and three for Travel. If the HTML homepage contains its final “show all” tile, the Gemini homepage appends an uppercase ZOBRAZIT VŠE or SHOW ALL link after those entries.
The remaining sections are handled differently:
- Propagační videa / Promotional Videos lists the individual external video links directly on the homepage and ends with
ZOBRAZIT VŠEorSHOW ALLpointing directly to the YouTube playlist. - O mně / About Me is rendered directly on the Gemini homepage, including the external link to the Sketchfab model.
- Kontakt / Contact is rendered directly on the Gemini homepage, including active
tel:,mailto:, and OpenPGP material links or blocks.
No separate Gemini section pages are created for Promotional Videos, About Me, or Contact. English is the capsule root language at /index.gmi; Czech is available under /cs/index.gmi.
8.8 Homepage archive metadata
Both the Czech (/) and English (/en/) homepages include three custom archive metadata tags in their HTML <head>:
<meta name="archive-documentation" content="https://vojtamaur.cz/metawebovy-clanek/" />
<meta name="archive-full-text" content="https://vojtamaur.cz/ALL_POSTS.txt" />
<meta name="archive-index" content="https://vojtamaur.cz/ARCHIVE.txt" />The tags are rendered by src/layouts/BaseLayout.astro only when pathForCs === "/vojtamaur-web/", which identifies both homepage variants. Both languages use the same absolute public URLs: the metaweb article explaining the archive, the full-text export, and the archive index. These metadata tags do not embed the contents of those resources.
The tags provide machine-readable archive entry points without adding visible page content or hidden anchor links. Their names are project-specific conventions; automatic discovery or indexing by crawlers is not guaranteed. They complement the existing sitemap and ordinary archive links.
9. Development and normal workflow
9.1 Local development
npm install
npm run devAstro normally starts the local server at http://localhost:4321/ and reacts continuously to changes in the project.
9.2 Practical note about the dev server
In this specific project, it sometimes happens that after adding a new .mdx file, the new article does not appear correctly in dev mode, or temporarily replaces another article in the listing. Restarting the development server usually fixes it immediately.
Recommended troubleshooting step:
Ctrl + C
npm run dev9.2.1 Public asset metadata audit
Before publishing new or changed public assets, run the metadata audit:
python scripts/audit-public-metadata.py --exiftool "D:\Program Files\exiftool\exiftool.exe"The script scans public/images/ and public/files/ and reports files containing privacy-relevant embedded metadata.
To preview metadata stripping without modifying files:
python scripts/audit-public-metadata.py --exiftool "D:\Program Files\exiftool\exiftool.exe" --strip --dry-runTo strip unintended metadata from supported image files:
python scripts/audit-public-metadata.py --exiftool "D:\Program Files\exiftool\exiftool.exe" --stripPDF files are treated more conservatively and are not stripped by default. If PDF metadata needs to be removed, create a checked public copy and verify the output afterwards.
After stripping metadata, run the audit again and check the changed files before committing.
Files with intentional embedded metadata can be kept through the script allowlist. These exceptions should remain explicit, because otherwise hidden metadata becomes indistinguishable from accidental leakage.
9.3 Web build
npm run build:webCreates the standard build intended for normal deployment to web hosting. This command uses existing English translation cache entries, but it does not create new DeepL translations.
The web build also creates the reconstructable source package at dist/source/vojtamaur-web-source.zip. The link to this package is root-relative in the source content (/source/vojtamaur-web-source.zip), which is correct for normal web hosting.
To create missing English translations, use:
npm run build:web:translateTo verify that no English translation cache entry is missing, use:
npm run build:web:strictTo inspect unused English translation cache files without deleting them, use:
npm run build:web:prune:dryTo delete unused, unprotected English translation cache files after the dry run looks correct, use:
npm run build:web:prune9.4 Build preview
npm run previewUsed for local verification of the production build.
9.5 Portable file-based build
npm run build:usbThis build is suitable, for example, for offline use, archiving, or transfer as a set of files.
The USB build also includes the reconstructable source package. In USB mode, scripts/usb-rewrite.mjs must run after the source package has been generated so that /source/vojtamaur-web-source.zip is rewritten to a relative file://-safe link.
To create missing English translations in the portable build, use:
npm run build:usb:translateFor a strict USB check, use:
npm run build:usb:strictTo inspect unused English translation cache files while producing the portable build, use:
npm run build:usb:prune:dryTo delete unused, unprotected English translation cache files after the dry run looks correct, use:
npm run build:usb:prune9.6 Plain-text export of all posts
During the build process, the project also generates a plain-text export of all article content:
/ALL_POSTS.txtThe export is produced by:
scripts/generate-all-posts.mjsThe script reads the finished static HTML from dist/, extracts the main article content, converts it into plain text, writes the result to dist/ALL_POSTS.txt, and embeds the same generated text into the finished dist/404.html recovery page. It then calls scripts/export-site-json.mjs to create dist/ALL_POSTS.json beside the text export. Both files are generated before the build integrity pass. Section 9.6.9 describes the JSON-LD format and its full, untruncated article content.
This is intentionally generated from the built output rather than directly from the source .mdx files. The English version is created as a post-build static artifact by scripts/en-postprocess.mjs, so reading from dist/ allows the export to include both Czech and English article versions.
The file is intended as a minimal preservation layer for:
- long-term archiving
- indexing
- full-text search
- offline reading
- recovery of textual content even if the original HTML structure, CSS, JavaScript, or build system becomes unusable
The export includes metadata for each article, such as title, slug, canonical URL, language, section, date, source file, and built HTML path.
Non-textual and embedded content is represented by explicit placeholders instead of being silently removed. For example:
[MEDIA: image]
[VIDEO EMBED]
[INTERACTIVE EMBED]
[PDF EMBED]
[SVG CONTENT OMITTED]HTML tables are converted into readable plain-text table blocks marked with [TABLE] and [/TABLE]. Table rows and cells are preserved in a Markdown-like form so that tabular data remains legible in the linear text export.
Code blocks and generated output blocks are preserved as [CODE BLOCK] sections. Very large code or generated output blocks are truncated when they exceed the configured size limits. The export keeps the beginning of the block up to the configured line or character limit, then adds an explicit truncation note with the original size and the amount of omitted content.
This prevents one unusually large generated block from making the entire text export difficult to read or process while still preserving a readable sample of the original block. The full version remains available in ALL_POSTS.json, the rendered website, source repository, or static snapshots.
The output file is written as UTF-8 with BOM to improve encoding detection in text editors and archival systems.
Recovery 404 page
The source template src/pages/404.astro turns the normal error page into a human-readable recovery interface. The explanatory text and recovery links appear first, followed by the complete visible content of ARCHIVE.txt and then the complete visible content of ALL_POSTS.txt. Nothing is hidden behind an expansion control, encoded as Base64, or fetched by client-side JavaScript.
During the Astro phase, the page reads public/ARCHIVE.txt and renders its URL lines as real HTML links. After the article pages and English postprocess are complete, scripts/generate-all-posts.mjs generates dist/ALL_POSTS.txt and replaces the marked <pre data-all-posts-embed> content in dist/404.html with an HTML-escaped copy of the same text. This order ensures that the standalone file and the copy carried by the 404 page always come from the same build.
The published page must still be served with the real HTTP status 404; 404.html is a custom response body, not a normal indexable success page. The separate /ARCHIVE.txt, /ALL_POSTS.txt, and /metawebovy-clanek/ resources remain normal recovery entry points. The large text blocks are intentionally rendered in full, with ordinary HTTP compression left to the hosting server or CDN.
Manual filtered and compact exports
The generated dist/ALL_POSTS.json can be filtered into a plain-text reading or archival copy manually with:
scripts/filter-all-posts.pyThis Python script is a separate post-processing tool. It is not called by npm run build:web, npm run build:usb, npm run build:arweave, npm run generate:all-posts, or any other normal build command. It reads an already generated dist/ALL_POSTS.json and does not modify the source export. Article metadata comes from vm:sourceMetadata, ordering from vm:position, and full text from articleBody. Structured media references are adapted to the existing text and compact-media records. Code is not shortened to the limits of ALL_POSTS.txt.
The script has no external Python dependencies. When it is stored in scripts/, the default input is resolved automatically as dist/ALL_POSTS.json. Available languages, sections, metadata fields, and current article counts can be inspected with:
python scripts/filter-all-posts.py --list-valuesArticle selection can be restricted by language, section, or an inclusive date range. Multiple languages or sections can be supplied by repeating an option or by separating values with commas. For example:
python scripts/filter-all-posts.py --language cs --section volna-tvorba,vystavy --format structured --metadata full
python scripts/filter-all-posts.py --language en --section volna-tvorba --format compact
python scripts/filter-all-posts.py --from-date 2020-01-01 --to-date 2026-12-31 --format structuredTwo output formats are available:
--format structuredis the default. It retains the delimited article-block structure and can preserve all metadata or use thefull,archive,minimal, ornonemetadata profiles. Exact metadata fields can also be selected with--metadata-fields, for example--metadata-fields TITLE,URL,DATE.--format compactcreates a compact reading and transfer copy. Long separator lines and repeated article metadata blocks are removed. Each article starts with [rendered title|YYYY-MM]; image references are reduced to records such as [img:filename|ALT: …|CAPTION: …]; and video, PDF, and interactive embeds use equivalent one-line records. Redundant blank lines outside code blocks and generic embed notes are omitted. Code indentation and blank lines inside code blocks are preserved.
A different JSON or JSON-LD snapshot can be supplied as the positional input argument. Legacy structured TXT exports remain supported when selected explicitly, for example:
python scripts/filter-all-posts.py exports/ALL_POSTS-2026-09-23.json --language cs --format compact
python scripts/filter-all-posts.py dist/ALL_POSTS.txt --language cs --metadata archiveThe legacy TXT input retains any truncation already present in that source. Compact output and structured output with --metadata none are not suitable as inputs for another filtering pass because the metadata required for reliable filtering has been removed. Use the original JSON again, or structured TXT output with sufficient metadata, when another filtering step is needed.
The default output directory is:
exports/Output filenames describe the selected language, section, date range, format, and metadata profile, followed by the export timestamp. A different path can be supplied with --output. The exports/ directory is generated and excluded through .gitignore, while scripts/filter-all-posts.py remains a versioned source file.
Filtered exports are written outside dist/ by default so that they do not silently change the contents described by the existing build integrity manifests. If an explicit --output path is placed inside dist/, the script prints a warning because the existing SHA256SUMS.txt, BUILD_SHA256.txt, and integrity.json do not cover the newly created file.
A selection can be fully parsed and validated without writing a file:
python scripts/filter-all-posts.py --language cs --section volna-tvorba --format structured --metadata full --dry-runThe script refuses to overwrite its input, rejects unknown language or section values, and rejects an empty selection unless --allow-empty is supplied explicitly. It reports the selected article count, output path, and SHA-256 hash after processing. Like the source export, written output uses UTF-8 with BOM.
9.6.1 Build integrity and OpenPGP signatures
During the final post-build phase, the project generates integrity metadata for the completed static output:
/SHA256SUMS.txt
/BUILD_SHA256.txt
/integrity.json
/SIGNING_STATUS.txtThese files are produced by:
scripts/generate-integrity.mjsThe script walks through the selected output directory, calculates a SHA-256 checksum for each included file, writes a sorted per-file manifest to SHA256SUMS.txt, and then calculates a global build hash from that manifest.
The output also contains BUILD_HASH_HISTORY.txt, supplied earlier by scripts/generate-source-bundle.mjs. The integrity generator requires and validates that snapshot; it does not refresh it from source, because the already generated source ZIP must contain the same bytes. Unlike the metadata files excluded below, BUILD_HASH_HISTORY.txt is included in SHA256SUMS.txt.
The global build hash is stored in:
BUILD_SHA256.txtMachine-readable informational metadata is stored in:
integrity.jsonEvery ordinary integrity pass starts in an unsigned state. It removes any stale SHA256SUMS.txt.asc, writes SIGNING_STATUS.txt as unsigned, and records openPgp.present: false in integrity.json. This prevents a detached signature from an older build from surviving after the manifest has changed.
The following files are intentionally excluded from the checksum manifest:
SHA256SUMS.txt
SHA256SUMS
BUILD_SHA256.txt
integrity.json
SHA256SUMS.txt.asc
SIGNING_STATUS.txt
.DS_Store
Thumbs.dbThe checksum and signing metadata files cannot describe themselves without creating a recursive dependency. In particular, integrity.json and SIGNING_STATUS.txt are informational and are not cryptographic proof that a build is signed.
The manifest uses the .txt extension intentionally. A bare file named SHA256SUMS can be interpreted badly by local preview or static hosting setups that use directory-style routes and trailingSlash: "always". SHA256SUMS.txt behaves as a normal downloadable text file.
The integrity files are generated from the final build output, not from the source files. They are intended to verify the published static artifact after Astro build, English postprocessing, generated text exports, path rewriting for portable builds, and any other post-build changes that happen before scripts/generate-integrity.mjs runs.
For the normal web build and USB build, the integrity files describe the final contents of dist/.
For the Arweave / Permaweb build, integrity is generated again after dist-arweave/ is created, so the integrity files inside dist-arweave/ describe the final Arweave deployment artifact rather than the original dist/ directory.
The public OpenPGP identity used for signed build artifacts is published separately in the static build:
/keys/vojta-maur-openpgp.asc
/keys/vojta-maur-openpgp-fingerprint.txtSHA-256 verifies whether the files match a particular manifest. The detached OpenPGP signature authenticates that exact manifest as one signed by the holder of the corresponding private key. Publishing the public key alone does not authenticate a deployment.
The explicit signing step is implemented by:
scripts/sign-build.mjsIt reads the expected primary-key fingerprint from public/keys/vojta-maur-openpgp-fingerprint.txt, locates the private key through GNUPGHOME, creates the armored detached signature:
/SHA256SUMS.txt.ascand immediately verifies the new signature against SHA256SUMS.txt. Only after successful verification does it write a signed SIGNING_STATUS.txt and update the informational openPgp object in integrity.json.
The final step appends the current build hash to the canonical source history. Before signing and again after verification, the signer checks that BUILD_SHA256.txt matches the manifest, that the manifest covers the exact history snapshot, and that integrity.json contains a valid build type and matching build hash/history reference. Failure before the final history commit leaves the canonical file unchanged.
If signing fails, the script removes SHA256SUMS.txt.asc, writes a failed unsigned status where possible, records openPgp.present: false, and exits with a non-zero status. A build is therefore signed only when the detached signature is present and verifies successfully.
When GNUPGHOME is already set, the script validates that directory and confirms that it contains the requested secret key. When it is not set and an interactive terminal is available, the script asks for the GnuPG home directory. In a non-interactive environment, GNUPGHOME must be set explicitly.
A Windows CMD example is:
set "GNUPGHOME=C:\path\to\gnupg"
npm run build:web:signedThe private signing key is never stored in the repository, copied into public/, included in the build, or provided to third-party CI systems.
The available signed build commands are:
npm run build:web:signed
npm run build:web:translate:signed
npm run build:translate:signed
npm run build:web:strict:signed
npm run build:usb:signed
npm run build:usb:translate:signed
npm run build:usb:strict:signed
npm run build:arweave:signedThe signing step can also be applied to an already generated output:
npm run sign:build
npm run sign:build:arweaveCumulative signed build history
BUILD_HASH_HISTORY.txt at the project root is the canonical source file shared by the web, USB, and Arweave signing workflows. It contains one record per successfully signed build, in append order, without a header:
ISO-8601 UTC timestamp | build type | SHA256 | hashThe timestamp is the local record-creation time in UTC, written as YYYY-MM-DDTHH:mm:ss.sssZ. The stable build type is web, usb, or arweave; translation and strict modes do not create additional types. The hash is the lowercase 64-character SHA-256 digest of that build’s SHA256SUMS.txt, also declared in its BUILD_SHA256.txt. Each record ends with a newline; an empty file is valid before the first recorded signed build.
For build N, the lifecycle is:
- Source-bundle generation embeds the existing history in both source archives and copies those exact bytes into the build as
BUILD_HASH_HISTORY.txt. - After all build-specific rewriting, integrity generation includes that snapshot in
SHA256SUMS.txtand writes the current manifest hash toBUILD_SHA256.txt. - The explicit signing step creates and verifies
SHA256SUMS.txt.asc, then writes the signed status and informational metadata. - Only then is the current hash appended to the project-root history. The history snapshot, manifest, and source archives inside build N remain unchanged. The next build inherits the new record.
This separates previous signed builds (BUILD_HASH_HISTORY.txt in the output) from the current build (BUILD_SHA256.txt). It avoids putting the current hash into a file that contributes to that same hash. The existing signature on SHA256SUMS.txt also authenticates the history snapshot through its checksum; no separate BUILD_HASH_HISTORY.txt.asc is created.
integrity.json records buildType and buildHistoryFile alongside buildHash. Both web and USB builds use dist/, and the USB runner’s environment is not automatically inherited by the later standalone signing process. The integrity step therefore stores usb in the metadata while BUILD_TARGET=usb is available, and the signer reads the stored type. The Arweave integrity command passes arweave explicitly. These metadata fields remain informational; verification uses the detached signature and the manifest checksums.
Unsigned builds and failed signing attempts do not append records. Re-signing the same already recorded build does not append a duplicate. If another build has advanced the canonical history since the snapshot was taken, a new unrecorded build using that old snapshot is rejected; rebuild from the current source history before signing it.
History updates use an exclusive lock and a temporary file under .source-bundle-staging/, which is excluded from Git and the source archives, followed by atomic replacement of the canonical source file. A lock conflict stops the update. After an interrupted signing process, remove a leftover .source-bundle-staging/build-hash-history.lock only after confirming that no signer is still running.
Preserve the updated project-root history in version control after a successful signed build. Do not replace the history inside an already signed output with that newer source file: the output intentionally contains only the previous records, and changing it would break its manifest checksum.
Verifying a signed build
Download the public key, checksum manifest, and detached signature into the same directory:
vojta-maur-openpgp.asc
SHA256SUMS.txt
SHA256SUMS.txt.ascImport the public key:
gpg --import vojta-maur-openpgp.ascDisplay its fingerprint:
gpg --fingerprint C5F5B3905220BE59The primary-key fingerprint must match exactly:
57E9 D455 FB10 A228 F66E 18AE C5F5 B390 5220 BE59The fingerprint should also be compared with a copy obtained from an independent trusted source or archive. Downloading the public key, manifest, and signature from the same compromised mirror would not by itself establish the author’s identity.
Verify the detached signature:
gpg --verify SHA256SUMS.txt.asc SHA256SUMS.txtA valid result includes a message similar to:
Good signature from "Vojta Maur vojtamaur.cz"A trust label such as [unknown] concerns the local GnuPG trust assigned to the identity. It does not mean that the mathematical signature verification failed. The important results are a good signature and an independently confirmed primary-key fingerprint.
Verifying the OpenPGP signature authenticates SHA256SUMS.txt. To check the complete downloaded build against that authenticated manifest, run this command from the build root:
sha256sum -c SHA256SUMS.txtThe last command requires sha256sum. It is normally available on Linux, through GNU coreutils on macOS, and on Windows through environments such as Git Bash or WSL.
The checksum check includes BUILD_HASH_HISTORY.txt. To also compare the current manifest digest with the value declared in BUILD_SHA256.txt, run sha256sum -c BUILD_SHA256.txt from the build root. A history file by itself is not a substitute for verifying the signature and checksums of the build that carries it.
9.6.2 Reconstructable source package
During the web and USB builds, the project generates a reconstructable source package:
/source/vojtamaur-web-source.zipIn the local build output this file is written to:
dist/source/vojtamaur-web-source.zipThe package is produced by:
scripts/generate-source-bundle.mjsThis source package is not a full copy of the generated website and it is not intended to duplicate all media files inside the ZIP. It is a compact reconstruction layer: it contains the project source, content, build scripts, package.json, package-lock.json when available, Astro configuration, selected small files from public/, and the files required to reconstruct omitted media.
Large or externally useful media assets from public/images/ and public/files/ are omitted from the ZIP and described in MEDIA_MANIFEST.json and MEDIA_SHA256SUMS.txt. The generated download-assets.py script can then restore them from the configured mirrors and verify them using SHA-256. The clean base file public/images/kurt-godel-rat.jpg is the intentional exception: it is bundled directly and listed in the manifest’s bundledPublicFiles field.
The public/demos/ directory is bundled directly into the source ZIP rather than restored through the asset downloader. The demo files are small enough to include, and their HTML can differ between source, web output, USB output, and static mirrors after path rewriting or deployment-specific postprocessing. Treating them as downloadable checksum assets would create false hash mismatches.
The generated package contains the downloader and reconstruction notes both at the ZIP root and inside source-bundle/. This allows a reconstructed copy of the project to build the website and then generate a new source package again. In other words, the source package is recursively reconstructable: a source ZIP can produce a rebuilt website, and that rebuilt website can produce another source ZIP.
The default reconstruction flow is:
python download-assets.py
npm install
npm run build:web:strictFor a portable file-based reconstruction:
python download-assets.py
npm install
npm run build:usbThe downloader first checks local candidate paths when possible, then uses the configured public mirrors. Files are accepted only when their SHA-256 hash matches the manifest. Hash mismatches are reported and rejected by default rather than silently written into public/.
After writing the standalone source package, the same generator also creates a JPEG/ZIP archival carrier at:
dist/images/kurt-godel-rat.jpgThe generator always reads the clean carrier base from public/images/kurt-godel-rat.jpg. Its complete JPEG byte stream, including all six redundant copies of the archival text in JPEG Comment, EXIF, Windows XP and XMP metadata fields, is written unchanged as the output prefix. The source-package ZIP is then written after the JPEG EOI marker with offsets that remain valid for strict ZIP readers. The image is never decoded, recompressed or passed through a metadata editor.
The clean carrier base is included directly in the source package and excluded from the external asset list. This breaks the otherwise unavoidable hash cycle in which a source ZIP would need to contain the final hash of the JPEG that contains that same ZIP. It also makes generation idempotent: every run starts from the clean public/ JPEG instead of appending another archive to the previous dist/ result.
ZIP-aware tools such as 7-Zip can open the published .jpg directly. Tools that require a .zip suffix can operate on a copied or renamed kurt-godel-rat.zip. Ordinary image software continues to read the same JPEG image and ignores the archive data after EOI.
The source package and JPEG/ZIP carrier are generated before the final integrity pass. Therefore SHA256SUMS.txt includes both artifacts as part of the final published or portable build.
Both archives include the project-root BUILD_HASH_HISTORY.txt. The generator copies the exact archived history bytes into the build output, so all three snapshots agree. They contain the records preceding the current build, even after that build is successfully signed. Signing updates only the canonical file in the working source tree; it does not regenerate either archive. A reconstruction from an older source ZIP therefore starts with that ZIP’s historical snapshot, not any records subsequently added to the maintained source repository.
9.6.3 Manual PDF export
The project includes a separate manual PDF export tool:
scripts/export-site-pdf.pyThe exporter is intentionally not run by any normal web, usb, or Arweave build. PDF generation is comparatively slow and can create large output files, so it is started only when an archival or reading copy is needed.
The script reads the finished static HTML from dist/, not the source .mdx files. This is necessary because English article pages are finalized by scripts/en-postprocess.mjs after the Astro build. The PDF export therefore reflects the same rendered Czech and English pages that exist in the finished build.
Before running the PDF exporter, create a standard web build (npm run build:web or another web-build variant), not a portable USB build. USB-rewritten paths in dist/ may produce invalid hyperlinks in the exported PDF.
Before the first use, install the Python dependencies and the Playwright Chromium browser:
python -m pip install -r requirements-pdf-export.txt
python -m playwright install chromiumThe standard command is:
npm run export:pdfEquivalently, the script can be called directly:
python scripts/export-site-pdf.pyThe default export:
- includes all non-draft articles from all article sections
- includes both Czech and English routes
- creates one combined A4 PDF
- adds a generated cover and article index
- adds PDF outline entries for the included pages
- writes
vojtamaur-web-export-pdf.manifest.jsonwith article metadata, output file sizes, and SHA-256 hashes
The generated files are written to:
exports/This is the same generated output directory used by scripts/filter-all-posts.py, so manually created text and PDF exports are kept together outside dist/. The exports/ directory is excluded through .gitignore. The exporter script and requirements-pdf-export.txt remain versioned source files.
Useful selection options include:
python scripts/export-site-pdf.py --lang cs
python scripts/export-site-pdf.py --lang en
python scripts/export-site-pdf.py --lang both
python scripts/export-site-pdf.py --section cestovani
python scripts/export-site-pdf.py --section volna-tvorba,vystavy
python scripts/export-site-pdf.py --separate--section can be repeated or supplied as a comma-separated list. --separate creates one PDF for each article/language page instead of one combined file.
The script normally expects an existing dist/ directory. A build can be requested explicitly for a particular run:
python scripts/export-site-pdf.py --build-command "npm run build:web:strict"This does not change the normal project build scripts; it only runs the supplied command before that individual PDF export.
By default, the exporter keeps Chromium’s original PDF output unchanged:
--pdf-quality archiveThis mode does not require Ghostscript. Optional PDF compression and image downsampling use Ghostscript as a second post-processing stage. A practical medium-quality export is:
python scripts/export-site-pdf.py --pdf-quality ebook --image-dpi 150 --jpeg-quality 75The available quality presets are:
archive
printer
ebook
screenThe archive preset preserves the original output. The other presets require Ghostscript and may be further adjusted with --image-dpi and --jpeg-quality. If Ghostscript is not detected automatically, its executable can be supplied explicitly:
python scripts/export-site-pdf.py --pdf-quality ebook --image-dpi 150 --jpeg-quality 75 --ghostscript "C:\Program Files\gs\gs10.07.1\bin\gswin64c.exe"--keep-uncompressed preserves the original uncompressed PDF alongside the compressed result.
The exporter renders pages through a temporary local server, but hyperlinks written into the PDF are converted to their public equivalents under https://vojtamaur.cz/. This prevents relative article, image, PDF, download, and other media links from pointing to a temporary address such as http://127.0.0.1:54321/. Loopback preview aliases such as localhost and 127.0.0.1 on the export server port are treated as the same temporary server. Existing external links remain external. Article images and other directly rendered media that do not already have a link are linked to their public source file, using the original HTML source attributes where possible rather than browser-resolved local preview URLs. The generated cover index also links each built route to the corresponding public article page.
If the same exporter is used for another deployed domain, the public root can be overridden:
python scripts/export-site-pdf.py --site-url https://example.comIn the Ghostscript-compressed modes (printer, ebook, and screen), iframes are replaced before printing by PNG snapshots. Each snapshot links to the original public embed target. YouTube embed URLs are converted to normal YouTube watch URLs; local PDF, scan, map, and other iframe sources link to their corresponding public file or page. The archive mode keeps Chromium’s original iframe rendering unchanged while still correcting normal page and media hyperlinks.
For English pages, the index title is read from the rendered HTML rather than copied from Czech MDX frontmatter. If the English postprocess marked a page as incomplete, the cover index and manifest label it as Incomplete / Czech fallback. An /en/ route can therefore exist even when some or all of its article content remains Czech.
The PDF is a paginated visual snapshot, not a lossless replacement for the website or source repository. Wide tables, long source-code lines, and other horizontally overflowing content can be clipped, wrapped differently, or extend beyond the printable area. The generated cover includes this limitation explicitly. The archived HTML, source files, repository, ALL_POSTS.txt, and other preservation layers remain the authoritative complete versions.
9.6.4 Ultra-compact PDF export
The project also includes a second PDF exporter designed for high-density archival copies:
scripts/export-site-pdf-ultra.pyUnlike scripts/export-site-pdf.py, this exporter does not reproduce the visual design of each web page. It reads the existing dist/ALL_POSTS.json, reuses the compact serialization implemented by scripts/filter-all-posts.py, and lays the selected content out as a dense A4 document. The tool is intended for cases where page count and storage size matter, while images and links should still remain present and usable.
The ultra-compact export:
- uses four text columns and 4 pt body text by default
- starts every article title on a new line
- includes aggressively resized image thumbnails while preserving their natural aspect ratios
- keeps consecutive image groups in regular rows of up to three images
- uses incomplete image rows efficiently by allowing a complete following paragraph to occupy the remaining width, rather than leaving an isolated sentence fragment beside the images
- writes video, PDF, and interactive-media targets as compact clickable text URLs instead of large embed cards
- keeps links clickable without visually underlining them
- omits unrecoverable image placeholders from the PDF and reports their count on the command line
- omits duplicate media by default when the same media occurs in both language versions
- writes an adjacent JSON manifest containing the selection, layout settings, media statistics, output size, page count, link count, and SHA-256 hash
The exporter is manual and is not called by the normal web, USB, Arweave, Gemini, or integrity build workflows. It expects a current standard web build containing:
dist/ALL_POSTS.json
dist/images/
dist/files/
dist/demos/The --input option can select another JSON-LD snapshot or a legacy structured TXT export. JSON supplies full article text and code; thumbnails still come from the local build assets. Ultra shortens only the recognized example output log in the dullgpt article, using the first 120 lines capped at 12000 characters, matching the MoM-like exporter. DullGPT’s executable Python program and all code/output blocks in other articles remain complete, regardless of length or whether the MDX fence specifies a language. The log is identified by its Input: ..., Output: records, not by an unlabelled code fence or a fixed block position.
Each shortened log contains a visible [TRUNCATED: ...] note with original and omitted counts. The console and both combined/separate manifests record the shortened logs per language and block index under content.dullgpt_log. The source JSON stays unchanged. Existing explicit truncation notes in legacy TXT are preserved; full content already omitted from such a TXT cannot be recovered by the exporter.
To retain the full DullGPT log from JSON, including with PDF/A, compression or separate exports:
python scripts/export-site-pdf-ultra.py --full-dullgpt-log--full-code-blocks remains a compatibility alias for --full-dullgpt-log. No flag is needed to preserve complete Python programs or other articles’ code blocks.
Create or refresh that input before exporting, for example with:
npm run build:web:strictThe Python packages used by the exporter are Playwright, pypdf, and Pillow. Playwright also requires its Chromium browser:
python -m pip install playwright pypdf Pillow
python -m playwright install chromiumThe recommended high-resolution Czech export is:
python scripts/export-site-pdf-ultra.py --lang cs --image-dpi 400The default output is:
exports/vojtamaur-web-export-ultra.pdfThe adjacent manifest is written as:
exports/vojtamaur-web-export-ultra.manifest.jsonLanguage and section selection work in the same style as the normal PDF exporter:
python scripts/export-site-pdf-ultra.py --lang cs
python scripts/export-site-pdf-ultra.py --lang en
python scripts/export-site-pdf-ultra.py --lang both
python scripts/export-site-pdf-ultra.py --section cestovani
python scripts/export-site-pdf-ultra.py --section volna-tvorba,vystavy
python scripts/export-site-pdf-ultra.py --lang cs --section cestovani
python scripts/export-site-pdf-ultra.py --separate--section can be repeated or supplied as a comma-separated list. Date selection is available through --from-date YYYY-MM-DD and --to-date YYYY-MM-DD. --separate creates one compact PDF for each selected article and language under exports/ultra-media-separate/ instead of creating one combined document.
Image quality can be adjusted independently from the physical thumbnail size:
python scripts/export-site-pdf-ultra.py --lang cs --image-dpi 300 --jpeg-quality 80The default thumbnail raster resolution is 240 DPI and the default JPEG quality is 70. Higher values generally produce sharper images and a larger PDF. --thumbnail-width-mm and --thumbnail-height-mm control the maximum physical thumbnail box; images are fitted inside that box without changing their original proportions.
The page format, column count, text size, spacing, margins, and page numbers can also be adjusted. All available options and their current defaults can be displayed with:
python scripts/export-site-pdf-ultra.py --helpThe script detects the project root from its own location, so it can be launched from the repository root or directly from the scripts/ directory. --project-root remains available when the script is stored or invoked from another location.
This output is a compact preservation and reading layer, not a lossless replacement for the website, source repository, original media, or ALL_POSTS.txt. Text is deliberately very small, images are downsampled, interactive media is represented by links, and missing built assets cannot be embedded. The source files and normal build artifacts remain the authoritative complete versions.
9.6.5 Metaweb archival PDF export
The project includes a dedicated manual exporter for the metaweb article and the preservation material linked from it:
scripts/export-metaweb-pdf.pyThe exporter reads the existing finished standard web build in dist/. It does not run Astro, regenerate integrity metadata, sign the build, or run automatically from any web, USB, Arweave, Gemini, or integrity build command. This preserves the exact relationship between the rendered English page and the identity files already present in the selected build.
Before the first use, install the same Playwright and pypdf dependencies used by the normal PDF exporter:
python -m pip install -r requirements-pdf-export.txt
python -m playwright install chromiumRun the export manually with:
npm run export:pdf:metawebThere is no generate:metaweb npm script. The command above is the supported npm entry point for this manual export.
All standard web and USB builds write to the same dist/ directory, so the most recently run build determines its routing and link format. Do not run the metaweb exporter against dist/ left by build:usb, build:usb:translate:signed, or another USB build variant. USB mode uses flat files such as metawebovy-clanek.html and rewrites links for file://, while the exporter expects the standard directory route /metawebovy-clanek/. A typical symptom of using USB output is Page.evaluate: Error: Article content container not found.
Restore a standard translated and signed web build before exporting:
npm run build:web:translate:signed
npm run export:pdf:metawebIf the English translation cache is already complete and a new Gemini export is not required, the strict standard build can be used instead:
npm run build:web:strict:signed
npm run export:pdf:metawebThe metaweb exporter intentionally does not rebuild or change dist/ by itself.
Equivalently:
python scripts/export-metaweb-pdf.pyThe default output is one PDF plus an adjacent JSON manifest:
exports/vojtamaur-web-export-metaweb.pdf
exports/vojtamaur-web-export-metaweb.manifest.jsonThe PDF contains:
- a simple white bilingual title page with the article title, author, and public website, followed by a separate bilingual contents page
- the Czech page
/metawebovy-clanek/ - the postprocessed English page
/en/metawebovy-clanek/ - a deduplicated appendix of every image linked from the article, with the physical archive ID and item name plus Czech and English descriptions derived from the corresponding article versions, the source filename, and public link
- a bilingual explanatory page followed by
exports/vojtamaur-web-export-ultra.pdf, the ultra-compact Czech export of every article with its images ARCHIVE.txt- the complete rendered technical documentation at
/documentation/ - every file linked from the article section
INTEGRITA A IDENTITA BUILDU, including the checksum manifest, detached signature when present, build checksum, JSON metadata, signing status, and public OpenPGP key
The payloads of ALL_POSTS.txt and source/vojtamaur-web-source.zip are deliberately not appended or embedded. Their descriptions and public links remain visible inside both rendered article versions because they are part of the article itself.
The physical archive registry is wider than an A4 page. During this export only, each table row is converted in the browser DOM into a labeled registry card. The card heading contains the ID and item name; the remaining table headers become field labels. This preserves every original cell and hyperlink while preventing the holder, location, access, and notes fields from being clipped. The website HTML and source MDX are not modified.
The English postprocessed HTML can split the build-integrity link paragraph at invalid positions and expose stray Markdown backticks. For the PDF only, the exporter normalizes those same six links and their descriptions into a readable list. It does not edit the finished build or change any linked payload.
The first page is a deliberately simple white bilingual title page. The generated timestamp, source build, section counts, notices, and outline are placed on a separate bilingual contents page. The image appendix is discovered from article links rather than a hardcoded filename list. The Czech and English pages point to the same images, so every image target is included only once while its card shows the shared archive item title and separate Czech and English article-derived descriptions. The exporter fails instead of silently creating a monolingual card if an image is missing from either article version. The article thumbnail is not included unless it is also linked from the article body. By default, three linked images are placed on each gallery page.
The combined PDF expects the ultra export at:
exports/vojtamaur-web-export-ultra.pdfIf that file does not exist, the metaweb exporter creates it automatically before rendering the combined document by running the equivalent of:
python scripts/export-site-pdf-ultra.py --lang cs --image-dpi 400An existing ultra PDF is reused unchanged. Immediately after the bilingual image appendix, the metaweb exporter inserts one generated bilingual introduction page. It identifies the embedded filename and explains that the following pages are an ultra-compact Czech export of every website article with its images. It also states the number of following ultra pages and that their intentionally dense multi-column layout serves as an archival overview. The complete pages of the standalone ultra PDF follow this introduction without being edited or overprinted.
Useful options include:
python scripts/export-metaweb-pdf.py --images-per-page 2
python scripts/export-metaweb-pdf.py --images-per-page 4
python scripts/export-metaweb-pdf.py --output exports/metaweb-custom.pdf
python scripts/export-metaweb-pdf.py --site-url https://example.com
python scripts/export-metaweb-pdf.py --no-manifestARCHIVE.txt, checksums, JSON, signatures, and public-key material are rendered completely as wrapping monospace text. Long code blocks and URLs in the technical documentation are also wrapped for A4. The final PDF contains outline bookmarks and validates that no file payload was embedded.
An unsigned build legitimately has no SHA256SUMS.txt.asc. If that optional link is present in the article but the file is absent from the selected build, the PDF includes an explicit absence notice rather than inventing signature content or failing the whole export. Other missing identity files and missing linked images are treated as build errors.
The exporter writes a manifest describing page ranges, including the generated ultra introduction and the unchanged standalone ultra PDF as separate sections, source routes, bilingual linked-image descriptions and hashes, the included ultra PDF and its hash, identity-file hashes, the registry-card validation result, the final PDF hash, and the two deliberately excluded payloads. Use --no-manifest only when the adjacent machine-readable record is not wanted; that option also removes a stale adjacent manifest left by an earlier run.
Optional PDF/A-2b and shared compression controls
All three manual PDF exporters (export-site-pdf.py, export-site-pdf-ultra.py,
and export-metaweb-pdf.py) accept the same optional final-processing controls.
Without these new options their existing ordinary-PDF defaults remain unchanged.
PDF/A is never generated by export-all.bat automatically.
python scripts/export-site-pdf.py --pdfa 2b
python scripts/export-site-pdf-ultra.py --pdfa 2b --compress medium
python scripts/export-metaweb-pdf.py --pdfa 2b --compress light
python scripts/export-site-pdf-ultra.py --pdfa 2b --compress medium --image-dpi 120 --jpeg-quality 75
python scripts/export-site-pdf-ultra.py --pdfa 2b --target-size 2M--pdfa 2b first retains the ordinary export at its normal output path, then
creates a separate *-pdfa-2b.pdf. For example, exports/vojtamaur-web-export-ultra.pdf
and exports/vojtamaur-web-export-ultra-pdfa-2b.pdf coexist. New compression
presets and target-size iterations apply to the PDF/A derivative in this mode.
Without --pdfa, these options process the ordinary PDF at its output path.
The site export retains its legacy --pdf-quality archive|printer|ebook|screen
interface; choose either a non-archive legacy preset or --compress, not both.
Its --keep-uncompressed still keeps a *-uncompressed.pdf for plain compression.
--compress | Maximum raster DPI | JPEG quality |
|---|---|---|
none | no downsampling | no forced lossy recompression |
light | 300 | 90 |
medium | 150 | 80 |
high | 96 | 60 |
--image-dpi N (at least 36) and --jpeg-quality N (1–100) independently
override the preset, including none. DPI caps apply to color, grayscale and
monochrome raster images; text and vectors remain resolution-independent.
JPEG quality controls Ghostscript’s DCT quantization dictionaries and disables
JPEG/JPX pass-through when recompression is requested. Actual JPEG encoders can
produce different sizes at the same nominal quality.
Ultra retains its original thumbnail defaults of 240 DPI and JPEG quality 70.
Its explicit image overrides still affect thumbnail creation as before, and
also override final processing when a preset, target or PDF/A is requested.
Final processing cannot restore detail already discarded in thumbnail creation.
Use --image-dpi 400 when a sharper original thumbnail edition is required.
--target-size 2M means strictly below 2,000,000 bytes; 2MiB means below
2,097,152 bytes. It starts from the requested preset/overrides, or medium if
none was selected, tries up to seven descending DPI/quality combinations, and
refines the first fitting bracket with up to three further trials. Every trial
reads the same original, avoiding accumulated recompression. Overrides are
starting ceilings in this mode; automatic trials may reduce them, with a
36 DPI floor and a quality floor of 30 (or a lower explicitly supplied quality).
The chosen file has the highest quality along this bounded search that was
measured to fit; a global quality optimum or exact byte size is not guaranteed.
The limit applies per PDF in --separate mode. If no trial fits, the smallest
candidate is retained, the manifest records target_met: false, a warning is
printed, and the exporter exits with status 2.
Ghostscript is required for final compression or PDF/A. Use --ghostscript PATH
if automatic detection fails. PDF/A uses -dPDFA=2, an embedded RGB ICC output
intent, font embedding and strict PDFACompatibilityPolicy=2. Compression and
PDF/A conversion occur together in the final Ghostscript invocation; the result
is never merged, annotated or compressed afterward. The RGB profile normally
comes from Ghostscript’s iccprofiles/srgb.icc; --icc-profile PATH supplies
another valid RGB ICC profile.
Page counts, dimensions, public URI links, outline presence and searchable-text
presence are checked against the original. PDF/A identification, embedded fonts
(including fonts in form resources), and an embedded ICC output intent are checked too. These structural checks do not establish
full PDF/A compliance. If veraPDF is available on PATH or in a standard Windows
installation, its CLI validates the selected candidate explicitly against 2b;
--verapdf PATH can select its executable or verapdf.bat. A failed or unusable
validator report aborts publication of the derivative. When veraPDF is unavailable,
the exporter explicitly reports that external validation was not performed and
records external_validation.status: not_performed.
On 2026-10-02, the maintainer manually checked the Site (CS), Ultra and Metaweb PDF/A-2b outputs with veraPDF; all three passed validation. Validate each new archival export separately.
Unless --no-manifest is used, each PDF/A derivative receives its own
*-pdfa-2b.manifest.json containing the source hash, final settings, trial sizes,
target result, structural checks, output hash and external-validation status.
The ordinary export manifest also records optional processing results.
Ghostscript or validation failures preserve the ordinary PDF and do not replace
an existing derivative; always check the command’s exit status before archival submission.
Implementation references: Ghostscript PDF/A and pdfwrite controls and veraPDF CLI validation. Run the regression checks from the project root with:
python -B -m unittest discover -s scripts/tests -p test_pdf_export.py -vThese tests do not require Ghostscript or veraPDF. Real conversions additionally require the local tools described above.
Run all manual exports on Windows
The manual TXT, Morse, PDF, Ultra PDF, MoM-like PDF and EPUB exporters append the local export start time to automatically generated filenames, before the extension: -YYYY-MM-DD_HH-MM-SS-ffffff. For example, vojtamaur-web-export-ultra-2026-10-05_14-30-00-123456.pdf and vojtamaur-web-export-ultra-2026-10-05_14-30-00-123456.manifest.json belong to the same run. Microseconds distinguish repeated runs, including multiple exports on the same day. Separate article exports also receive the timestamp. An explicit --output is used exactly as supplied, including its adjacent TXT and manifest names. --no-timestamp restores legacy automatic names; --timestamp YYYY-MM-DD_HH-MM-SS-ffffff selects a specific edition timestamp. PDF/A and uncompressed derivative labels precede the timestamp. The generated build files in dist/ retain their stable names.
export-all.bat generates one shared timestamp in VOJTAMAUR_EXPORT_TIMESTAMP for all eleven steps, including the SSTV output directory. A new batch run creates a new timestamp and preserves older editions. Metaweb PDF includes the Ultra PDF with that same timestamp, creating it at that exact path if needed; a standalone Metaweb PDF run creates its own matching Ultra edition. MoM-like automatic English and bilingual names add -en and -cs-en, respectively, before the timestamp. Metaweb EPUB uses an .epub.manifest.json sidecar so it can coexist with the PDF’s .manifest.json without replacing it.
The default output listings in the sections below and above show the base names; the timestamp is appended automatically unless an explicit --output or --no-timestamp is used.
The repository root contains export-all.bat. It switches the Windows console and Python standard streams to UTF-8, changes to the repository directory, runs the eleven export steps sequentially, and stops immediately if any command fails. The ultra export runs before the metaweb PDF export, so the combined PDF receives the freshly generated ultra PDF rather than a stale copy. The workflow requires both dist/ALL_POSTS.txt and dist/ALL_POSTS.json from the existing build and begins with separate English and Czech compact text exports. It uses the built JSON without regenerating it, preserving the bytes covered by any existing build checksum manifest or signature. The PDF steps include separate Czech and English MoM-like editions of volna-tvorba at the default 4 pt font size. After the PDF steps it creates the practical compact Metaweb EPUB and separate compact Czech and English site EPUBs, preserving animated GIFs. The SSTV PNG export is the final step:
python scripts/filter-all-posts.py --language en --section volna-tvorba --format compact
python scripts/filter-all-posts.py --language cs --section volna-tvorba --format compact
python scripts/export-site-morse.py --lang both
python scripts/export-site-pdf.py --pdf-quality ebook --image-dpi 150 --jpeg-quality 75 --ghostscript "C:\Program Files\gs\gs10.07.1\bin\gswin64c.exe"
python scripts/export-site-pdf-ultra.py --lang cs --image-dpi 400
python scripts/export-site-mom-like.py --lang cs
python scripts/export-site-mom-like.py --lang en
python scripts/export-metaweb-pdf.py
python scripts/export-metaweb-epub.py --image-quality compact --gif-mode preserve
python scripts/export-site-epub.py --lang both --image-quality compact --gif-mode preserve
python scripts/export-site-sstv.pyRun the complete workflow with:
export-all.batThe batch file also reuses the current dist/ and does not create a web build. If the last build was a USB variant, restore a standard web build before running export-all.bat.
The SSTV PNG export runs last in export-all.bat and can also be run separately with npm run export:sstv (section 9.6.8). The batch does not generate SSTV audio.
9.6.6 Manual EPUB exports
The project includes two manual reflowable EPUB exporters:
scripts/export-site-epub.py
scripts/export-metaweb-epub.pyBoth exporters read the finished HTML and assets from dist/. They do not convert the source MDX directly. This preserves the postprocessed English pages and makes the EPUB content reflect the selected finished build. The exporters support both normal directory routes such as dist/article/index.html and portable flat routes such as dist/article.html.
Install their Python dependencies once. Pillow 11.3 or newer is required so that any AVIF image discovered in a finished build can be decoded and converted to an EPUB core image format:
python -m pip install -r requirements-epub-export.txtThe general site exporter also reads article selection metadata from the generated dist/ALL_POSTS.txt. Its default command is:
npm run export:epubThe default --lang both selection intentionally creates two independent publications:
exports/vojtamaur-web-export-all-cs.epub
exports/vojtamaur-web-export-all-en.epub
exports/vojtamaur-web-export-epub.manifest.jsonIt does not combine the Czech and English website editions into one bilingual book. Language and section selection follow the normal PDF exporter:
python scripts/export-site-epub.py --lang cs
python scripts/export-site-epub.py --lang en
python scripts/export-site-epub.py --lang both
python scripts/export-site-epub.py --section cestovani
python scripts/export-site-epub.py --section volna-tvorba,vystavy
python scripts/export-site-epub.py --lang cs --section cestovani
python scripts/export-site-epub.py --separateEach combined publication begins with a visible generated title/frontmatter and reflowable article index. It contains the same export metadata, limitations, language status, article dates, sections, titles, and built routes as the generated index in the normal site PDF. The routes are internal links to the corresponding EPUB chapters. This is an export index, not a copy of the website homepage. Pass --no-cover only when this frontmatter is not wanted.
--section can be repeated or comma-separated. --separate creates one EPUB per article under exports/<language>/<section>/; like the PDF exporter, these single-article files do not receive the combined export index. A custom --output can be used only for a combined single-language export, because --lang both always has two output files.
As with the PDF exporter, a build command is run only when explicitly requested:
python scripts/export-site-epub.py --build-command "npm run build:web:strict"Each book keeps the existing article order and content structure. The technical EPUB transformation:
- adds the generated title/frontmatter and linked article index as the first spine document in combined mode
- places every article in its own XHTML spine document with the correct
langandxml:lang - generates EPUB 3 package metadata, manifest, spine, navigation document, and stable reflowable CSS
- embeds and deduplicates local JPEG, PNG, GIF, SVG, and WebP resources used by article images or direct image links; local AVIF is converted to JPEG or PNG because AVIF is not an EPUB core image format
- rewrites links between included articles to internal EPUB links
- keeps downloads, PDFs, demos, and other non-image payloads as public
https://vojtamaur.cz/links - replaces iframes and other interactive embeds with labelled public links because a non-scripted reflowable EPUB cannot preserve their runtime
- removes empty built image placeholders that have no source
The exporter can reduce image size without changing article structure. The default preserves the exact source bytes of EPUB-core image formats; AVIF is the intentional exception and is always converted to JPEG or PNG:
--image-quality archiveOther presets are:
printer maximum side 2400 px, JPEG quality 90
ebook maximum side 1600 px, JPEG quality 82
screen maximum side 1200 px, JPEG quality 72
compact maximum side 800 px, JPEG quality 60, static GIF poster framesFor example:
python scripts/export-site-epub.py --image-quality ebook
python scripts/export-site-epub.py --image-quality compact
python scripts/export-site-epub.py --lang both --image-quality compact --gif-mode preserve
python scripts/export-site-epub.py --lang cs --image-max-px 1400 --jpeg-quality 78The third command is the recommended practical complete-site export. It creates separate Czech and English EPUB files with compact still images while preserving animated GIFs. Use plain --image-quality compact only when the smallest practical files are more important than GIF animation.
--image-max-px and --jpeg-quality override the selected preset. JPEG files are downscaled and recompressed; PNG files are downscaled losslessly when necessary. The archive, printer, ebook, and screen presets preserve animated GIF and SVG resources unchanged. The deliberately lossy compact preset converts opaque PNG files to JPEG and, unless overridden, animated GIF files to static first-frame PNG posters. These two decisions can be controlled independently with --png-mode preserve|jpeg and --gif-mode preserve|poster. Preserved animated GIF files retain their original bytes and therefore remain the main size limit. The JSON manifest records the selected and actually used JPEG quality, original and packaged image byte counts, the number of optimized images, interactive fallbacks, and removed empty placeholders. The Metaweb manifest also records separate source and packaged byte sizes and SHA-256 hashes for every asset, so a converted or optimized file is never described by the hash of a different representation.
The dedicated Metaweb EPUB remains one bilingual publication:
npm run export:epub:metawebIts default outputs are:
exports/vojtamaur-web-export-metaweb.epub
exports/vojtamaur-web-export-metaweb.epub.manifest.jsonThe Metaweb EPUB follows the semantic section order of export-metaweb-pdf.py:
- bilingual title page
- bilingual contents and export notes
- Czech Metaweb article
- English Metaweb article
- a visible bilingual image appendix with every locally linked physical-archive image, its Czech and English descriptions, filename, and source URL
- the complete
ARCHIVE.txtpayload - the complete rendered technical documentation from
/documentation/ - every file discovered in the article’s build identity and integrity section, each as a text appendix with source URL, byte size, SHA-256, and full payload
The publication metadata declares both cs and en. Both Metaweb article chapters keep their own language and heading IDs. Their wide physical archive tables are converted to reflowable record cards while preserving every cell and link, all <details> blocks are opened, and the malformed English integrity paragraph is normalized to the same six-item list used by the PDF exporter.
The PDF’s embedded vojtamaur-web-export-ultra.pdf is intentionally not inserted into the EPUB and is not replaced by expanded article chapters. The dense standalone PDF remains a separate publication artifact; the Metaweb EPUB contains no substitute for that section.
As in the PDF exporter, the ALL_POSTS.txt and source/vojtamaur-web-source.zip payloads are deliberately not appended; their mentions and public links remain in the article. A missing optional SHA256SUMS.txt.asc in an unsigned build receives an explicit explanatory appendix instead of invented content.
The Metaweb exporter supports the same image controls. The first command below is the recommended practical Metaweb export:
python scripts/export-metaweb-epub.py --image-quality compact --gif-mode preserve
python scripts/export-metaweb-epub.py --image-quality compact
python scripts/export-metaweb-epub.py --image-quality ebook
python scripts/export-metaweb-epub.py --output exports/metaweb-custom.epub
python scripts/export-metaweb-epub.py --no-manifestThe recommended command keeps animated GIFs intact while applying the compact preset to still images. Use plain --image-quality compact only when static first-frame posters are acceptable in exchange for a smaller EPUB.
The Metaweb EPUB therefore remains focused on the Metaweb material and its direct archival appendices instead of duplicating the complete website corpus.
Every generated book is validated before it replaces the destination file. When the site exporter creates multiple books, all uniquely named candidates are completed first; a conversion failure changes none of the requested outputs, and a handled failure during installation restores the previous set. JSON manifests are likewise written through unique temporary files and atomically replaced. The built-in validator checks the ZIP/mimetype rules, container and package documents, required metadata, correctly namespaced MathML and inline SVG properties, manifest, spine, navigation item, XML/XHTML well-formedness, duplicate IDs, and all packaged local references. The Metaweb exporter additionally verifies its section counts, visible image count, and deliberate payload exclusions before replacing an existing output. This does not replace a final EPUBCheck run when the external Java validator is available, but malformed or structurally incomplete output is rejected locally.
9.6.7 Gemini capsule and Gopher map generation
The project generates a separate bilingual Gemini capsule and a Gopher-compatible map layer with:
npm run generate:geminiThe generator is:
scripts/generate-gemini-capsule.mjsThe output directory is a sibling of the normal web build:
dist/
dist-gemini/generate:gemini requires an existing finished standard web build in dist/. It does not run Astro or the English translation postprocess by itself. Do not use it against dist/ left by a USB build. Although the generator can locate both directory-style article files and flat slug.html files, USB postprocessing has already changed internal links and asset paths for file://. The Gemini command can therefore finish without an error while deriving incorrect web paths from that USB-specific HTML.
For a complete build from source, use:
npm run build:geminiThis command runs the strict standard web build first and then generates the text edition.
The translating production workflow also generates the text edition automatically. Therefore:
npm run build:web:translate:signedbuilds and translates the normal website, generates dist-gemini/, and then signs the checksum manifest belonging to dist/.
The Gemini/Gopher text edition is deliberately generated from the finished HTML rather than directly from the MDX body. This preserves the actual postprocessed English output and the rendered behavior of shared Astro components. The generator also reads article frontmatter from src/content/posts/ for stable slugs, section membership, dates, draft filtering, and ordering.
The default output structure is:
dist-gemini/
index.gmi
gophermap
favicon.txt
keys/
personal-work/
index.gmi
gophermap
article-slug.gmi
exhibitions/
index.gmi
gophermap
article-slug.gmi
travel/
index.gmi
gophermap
article-slug.gmi
cs/
index.gmi
gophermap
volna-tvorba/
index.gmi
gophermap
article-slug.gmi
vystavy/
index.gmi
gophermap
article-slug.gmi
cestovani/
index.gmi
gophermap
article-slug.gmiEnglish is intentionally the default language at the capsule root. Czech pages are stored below /cs/. Each article page includes navigation back to the homepage and section, a link to the counterpart language when available, and a link to the full HTTPS web version.
For envs.net, the Gemini capsule is published below the user path rather than directly at the host root. Internal generated Gemini links therefore use the configured prefix /~vojtamaur so that links resolve under gemini://envs.net/~vojtamaur/... instead of incorrectly resolving under gemini://envs.net/.... If the capsule is moved to a host where it is served directly from the Gemini host root, this prefix must be changed or removed in the generator.
The same dist-gemini/ output can also be uploaded to public_gopher/. Gemini clients ignore gophermap files as ordinary unrelated files, while the Gopher server uses them as directory menus. The intended envs.net deployment mapping is:
dist-gemini/ -> public_gemini/ -> gemini://envs.net/~vojtamaur/
dist-gemini/ -> public_gopher/ -> gopher://envs.net/1/~vojtamaurGopher item types matter in the generated maps. Directory/menu links use type 1; plain text or Gemtext files use type 0; informational separator or label lines use type i. Some Gopher clients or web gateways may visibly show the leading i on informational lines. For this reason, gophermap should remain a thin navigation layer, while the real page content stays in .gmi files. Clients such as Lagrange can still display those .gmi files as Gemtext after opening them through Gopher.
The Gemini homepage contains:
- the site title and motto
- a language switch
- a short description of the capsule
- a provider-neutral capsule description without host-specific acknowledgements
- up to nine latest items for Personal Work and Exhibitions, and three for Travel
SHOW ALL/ZOBRAZIT VŠEafter the previews where the HTML homepage exposes that action- promotional video links followed by a direct link to the YouTube playlist
- the full About Me and Contact content directly on the homepage
The generator converts the built HTML into Gemtext using the following rules:
- headings become Gemtext headings, limited to the three heading levels supported by Gemtext
- paragraphs remain paragraphs
- explicit HTML
<br>elements remain line breaks instead of being collapsed into spaces - unordered and ordered lists become Gemtext list items
- blockquotes become
>lines <pre>and code blocks become Gemtext preformatted blocks- HTML tables become tab-separated preformatted blocks
- horizontal rules become plain separator lines
- images become active links to the canonical HTTPS asset URL
- image ALT text or captions are retained when useful, and the visible link label also includes the target URL so that the destination remains explicit
- PDF, video, map, Sketchfab, and interactive-demo embeds become external links
- local demo paths such as
/demos/example/example.htmlbecome complete links underhttps://vojtamaur.cz/ - internal article and section links are rewritten to local Gemini routes when a corresponding generated page exists
- unknown internal routes remain HTTPS links to the main website instead of becoming broken Gemini paths
mailto:andtel:references remain active Gemtext links without duplicated plain-text copies- generic embeds without a useful title use their URL as the visible label rather than the meaningless text “Embedded content”
The output is written as UTF-8. One client-specific workaround is applied during final writing: a literal Unicode replacement character U+FFFD (�) is serialized as the visible text \uFFFD. The original MDX and HTML remain unchanged. This exists because the tested Lagrange renderer stops drawing the remainder of a line immediately after a literal U+FFFD, even though the character is valid UTF-8.
The generator copies favicon.txt and the public keys/ directory from dist/, or falls back to public/ when necessary. Binary article media are not copied into dist-gemini/; they remain available through stable HTTPS links to the main website.
Before generation, the existing output directory is removed and recreated. The script contains safety checks that refuse to delete the project root, the normal dist/ directory, the filesystem root, or an output directory whose name does not contain gemini.
Available command-line overrides are:
--dist <dir>
--output <dir>
--content <dir>
--site-url <url>Because dist-gemini/ is outside dist/, the existing web integrity and signing commands do not cover it. Even when the capsule is produced during build:web:translate:signed, the resulting detached signature authenticates only dist/SHA256SUMS.txt. A separate checksum or signing workflow for dist-gemini/ has not been implemented. This applies to both the Gemini files and the generated Gopher map layer.
9.6.8 Manual SSTV PNG export
The project includes a separate manual exporter for an SSTV (Slow-Scan Television) edition of its articles:
scripts/export-site-sstv.py
requirements-sstv-export.txtIt creates numbered RGB PNG pages at the native resolution of the selected standard SSTV mode. Each PNG represents one complete frame. The exporter is started explicitly with npm run export:sstv or as the final step of export-all.bat; it is not called by any normal web, USB, Arweave, or Gemini build.
Setup and build-before-export workflow
Install the Python dependencies once. Chromium is used to rasterize local SVG images and can reuse the installation already needed by the PDF exporters:
python -m pip install -r requirements-sstv-export.txt
python -m playwright install chromiumRun the export commands from the repository root. The exporter reads article selection metadata from dist/ALL_POSTS.txt, then takes the full article text and media from the corresponding finished HTML and local assets in dist/. It does not use the shortened text representation in ALL_POSTS.txt as the article body and does not render the source .mdx directly. This also preserves the English content produced by the translation postprocess.
After editing an article, a media component, or a source image, rebuild the website before exporting again. Running only npm run export:sstv reuses the existing dist/, including any old content or broken image references still present there. The exporter never starts a build or refreshes translations itself.
For a Czech export, or when the existing English translation cache is sufficient:
npm run build:web
npm run export:sstvThe usual translating and signing workflow can also be used:
npm run build:web:translate:signed
npm run export:sstvThe second workflow fills missing translations and signs the web build. Signing is not required to generate SSTV PNGs, and the web signature does not cover the separate SSTV export directory. If dist/ already contains the desired finished build, only the export command is needed. For the shared PDF, EPUB, and SSTV workflow, use a standard web build; the SSTV exporter itself can also resolve directory routes and flat USB article routes.
Mode, language, and preview selection
The default command exports all indexed Czech articles from all article sections, using PD120:
npm run export:sstvThe following mode choices are implemented:
--mode | Native PNG dimensions | Aspect ratio |
|---|---|---|
pd120 (default) | 640 × 496 px | 40:31 |
pd180 | 640 × 496 px | 40:31 |
pd240 | 640 × 496 px | 40:31 |
pd290 | 800 × 616 px | 100:77 |
--mode pd290 changes the PNG raster to 800 × 616 pixels and scales the default typography accordingly. PD120, PD180, and PD240 use the same raster and layout; their mode labels and export metadata differ. The exporter keeps each mode’s native frame dimensions without introducing a square format. Mode references are recorded in manifest.json and in the script’s MODE_SOURCES list.
Useful commands:
npm run export:sstv -- --max-pages 10
npm run export:sstv -- --slug koncepty --max-pages 10
npm run export:sstv -- --mode pd290
npm run export:sstv -- --lang en
npm run export:sstv -- --mode pd290 --lang both
npm run export:sstv -- --section volna-tvorba,vystavy
npm run export:sstv -- --helpThe -- after the npm script name passes the remaining options to the exporter. --section and --slug accept repeated options or comma-separated values. --lang both puts the selected Czech and English article pages into one numbered sequence, with language labels on the frames and in the manifest. Incomplete English translations remain visibly marked as EN incomplete / Czech fallback.
--max-pages 10 is a preview limit of ten complete PNG frames for the whole run, not ten articles or ten frames per article. If the limit stops the export before all selected content has been rendered, the run is marked as partial. Omit this option for a complete export.
Body text defaults to 24 native pixels for PD120/PD180/PD240 and 30 pixels for PD290. --font-size accepts 16–48 pixels; --font and --bold-font accept local font files. Additional path options are --project-root, --dist, --output-dir, and --site-url; use --help for their current meanings.
Layout and output files
The export includes articles listed in the finished build index. Homepage sections and the contents of linked PDF, download, or source-package files are not appended automatically.
The SSTV layout uses:
- black text on white, a single reading column, page numbers, and article identification
- text wrapping and pagination, including long code blocks; wide tables become labelled cell records with their links retained
- separate image frames fitted proportionally without cropping or stretching, with EXIF orientation applied and transparency composited onto white
- the first frame of animated images; local SVG image files are rasterized through Chromium
- visible references for interactive media and unavailable images; remote media are not downloaded
- available fallback fonts for additional Unicode characters, with explicit code-point text if a glyph cannot be rendered
Every successful run creates a new timestamped directory rather than mixing its frames with an earlier run. For example:
exports/sstv/pd120/cs-<timestamp>/
000001.png
000002.png
...
manifest.json
README.txtA preview stopped by the page limit has -preview before its final timestamp. --output-dir overrides the parent directory; the exporter still creates a new timestamped child inside it. The default exports/ tree is already ignored by Git and remains outside the signed dist/ build.
Use the numbered PNGs in filename order. README.txt identifies the selected mode and dimensions. manifest.json records the selection, article page ranges, translation status, rendered text and image references, native dimensions, fonts, source and PNG hashes, warnings, and whether the export is partial.
partial: False means all content selected for that run was processed. It does not mean that the warnings list is empty or that SSTV transmission quality has been tested. An unfiltered Czech export covers the indexed Czech articles; a filtered export is complete only for its selected subset.
Warning diagnosis and stale build output
Read the warnings array in the run’s manifest.json for the reasons behind the console warning count.
Missing, remote or unsupported image sourcemeans an image could not be resolved to an available local build asset. An emptysourcevalue can come from aMediaRowitem written as{ type: "image", src: "" }. Use{ type: "empty" }for an intentionally blank cell, as described in section 7.2. The exporter renders unavailable images as[Image unavailable]; correctly rendered empty cells are skipped. After correcting the MDX, rebuild the web and then export again. Otherwise the old<img>elements remain indist/and the same warnings return.Cannot render image: ...includes the reason why an existing image could not be decoded or rasterized. Check the named source and any required image-rendering dependencies.Missing font glyphs rendered as explicit Unicode code pointslists characters that the available fonts could not render. The PNG uses visible text such as[U+FFFD]; the manifest retains the original text. This warning groups affected characters rather than reporting every occurrence separately.U+FFFDis the Unicode replacement character (�), which can already occur in experimental article content such as Anti-jazyk. Its presence does not by itself establish that the SSTV exporter damaged the text.
Audio tools and working directories
The PNG-to-audio and recovery procedure below was verified on 2026-09-10 using the scripts, archival README, and WAV files described in this section. It documents the working sequence, including MMSSTV for continuous reception.
The project includes the following audio tools and archival README:
scripts/sstv-audio/
encode_sstv.bat
join_wavs.py
README.txtThe complete operational instructions for these tools are maintained here in documentation.mdx. README.txt is the archival companion intended to travel with the finished WAV volumes, for example in an Internet Archive deposit. It describes the specific six-volume edition tested on 2026-09-10.
Both scripts resolve their input and output directories relative to their own location. For each new edition, copy scripts/sstv-audio/ to a new working directory under exports/sstv-audio/, then put the numbered PNGs from one finished PD120 export into its input_images/ directory. Keep the PNG export’s manifest as png-manifest.json beside the scripts. Keep the archival README.txt from scripts/sstv-audio/; the shorter README.txt generated with the PNGs is a different document.
For example, in Windows Explorer copy the tool directory to exports/sstv-audio/2026-09-10-cs/ and use this layout:
exports/sstv-audio/2026-09-10-cs/
encode_sstv.bat
join_wavs.py
README.txt
png-manifest.json
input_images/
000001.png
000002.png
...
input_wav/
output_wav/Run the audio commands below in the working copy, not in scripts/sstv-audio/. Use a new directory name for a new edition; reuse a working directory only to resume the same unchanged PNG set. Working files under exports/ are excluded from Git and the reconstructable source bundle. Keep scripts/sstv-audio/ for the small source files: the source-bundle generator traverses scripts/, so putting generated PNG/WAV files or an Open-SSTV installation there would also put them in the source ZIP.
All audio stages are manual. npm run export:sstv continues to generate PNGs only; neither the audio scripts nor the decoding software are invoked by the website build or export-all.bat.
Open-SSTV encoder setup
The tested encoder is open-sstv-encode from Open-SSTV. The working installation used a source checkout at %USERPROFILE%\Downloads\Open-SSTV with a Python 3.11.9 virtual environment. The inspected checkout declares version 0.6.10, requires Python 3.11 or newer, and was at commit 0cf68e0d2ee0ca85d1f9947403319a2a92043be3. These details record the tested installation rather than a promise about future releases.
The Windows GUI ZIP used in the experiment did not provide the required CLI executable. Installing with pip install open-sstv also failed in that test. The successful setup used the source repository and a Python 3.11 virtual environment. For a fresh installation, after installing Python 3.11 and Git, the commands are:
cd /d "%USERPROFILE%\Downloads"
git clone https://github.com/bucknova/Open-SSTV.git
cd Open-SSTV
py -3.11 -m venv .venv
.venv\Scripts\activate
python -m pip install -e .
open-sstv-encode --helpReuse an existing working installation instead of recreating it. The test initially failed with Python 3.10; explicitly selecting Python 3.11 when creating the environment resolved that dependency requirement. Open-SSTV is separate from the Python environment used by the website’s PNG exporter.
encode_sstv.bat finds the encoder in this order:
- the executable path in
OPEN_SSTV_ENCODER, if set and present open-sstv-encode.exeonPATH.venv\Scripts\open-sstv-encode.exebeside the batchOpen-SSTV\.venv\Scripts\open-sstv-encode.exebeside the batch%USERPROFILE%\Downloads\Open-SSTV\.venv\Scripts\open-sstv-encode.exe
The final fallback matches the tested installation without hard-coding the account name. If the encoder is installed elsewhere, set OPEN_SSTV_ENCODER to its full executable path before running the batch. The installed encoder can then be used without manually activating its virtual environment for every export.
PNG to WAV encoding
Use PNGs from the website exporter’s --mode pd120 output: 640 × 496 pixels. The copied batch is fixed to PD120 and invokes Open-SSTV with the differently spelled mode identifier pd_120. It does not read the PNG manifest or accept a mode argument. Choosing --mode pd290 in the PNG exporter does not change this batch; using another audio mode requires a corresponding change in a working copy and a new test.
From the working directory containing the scripts and input_images/, run:
encode_sstv.batFor each input PNG, the batch runs the equivalent of:
open-sstv-encode "input_images\000001.png" --mode pd_120 -o "input_wav\000001.wav"The result is one complete SSTV transmission per image, with its own VIS mode-identification signal. The actual test files are uncompressed mono PCM WAVs at 48,000 Hz and 16 bits per sample. The batch uses the encoder’s default sample rate; it does not pass --sample-rate explicitly.
The batch creates input_wav/ if needed, reports missing input PNGs, stops on an encoder error, and pauses for keyboard input when finished. Existing WAV filenames are skipped. This allows an interrupted run to resume, but the skip checks only file existence: it does not check PNG changes or validate an existing WAV. Use a fresh working directory when the input edition changes. Inspect an output left by an interrupted encoding before reusing it.
Joining WAVs into archive volumes
After encoding has finished successfully, run this in the same working directory:
py join_wavs.pyjoin_wavs.py uses only Python’s standard library. It reads all .wav files from input_wav/, sorts them naturally by filename, and writes to output_wav/. There is no manually maintained files.txt list and no FFmpeg step in this final workflow.
The script checks that the input WAVs are uncompressed PCM with matching channel count, sample width, sample rate, and compression type. It writes their PCM sample data in sequence without resampling or re-encoding, inserting one second of digital silence between transmissions. A frame remains whole when a volume boundary is reached; the inter-frame silence is placed at the beginning of the next volume. No extra silence is appended after the final frame.
The configured data-size limit is MAX_WAV_DATA_BYTES = 3_900_000_000, below the capacity of a classic RIFF/WAV file. Splitting is determined by byte size, not a fixed number of images. If the complete sequence fits, the output is combined-with-gaps.wav. Otherwise the names are:
combined-with-gaps_001.wav
combined-with-gaps_002.wav
combined-with-gaps_003.wav
...The joiner overwrites matching output names and does not remove leftover volumes from an older run. Use the separate working directory for the current edition to keep the inputs and output set consistent. Some messages in the preserved script still mention input or output; its actual directory constants are input_wav and output_wav.
WAV to images: tested MMSSTV reception
The successful continuous decoder was MMSSTV for Windows. Open-SSTV successfully encoded frames and decoded individual image WAVs during the experiment, but the tested continuous multi-image playback did not produce the desired sequence there. MMSSTV with automatic restart was then verified to reconstruct consecutive pages. This is a record of the tested setup, not a general claim about all Open-SSTV versions.
The working audio path is:
WAV volume played at normal speed
-> Windows playback output
-> Stereo Mix / Směšovač stereo (Realtek)
-> MMSSTV recording input
-> automatic VIS detection and PD120 reception
-> progressively drawn pages and received-image historyThe initial MMSSTV failure was caused by Stereo Mix being disabled while a microphone was the default recording device. The working Windows setup was:
- Open Control Panel → Sound → Recording (
Ovládací panely → Zvuk → Záznam). - Enable Stereo Mix / Směšovač stereo (Realtek). If hidden, show disabled devices first.
- Set it as the default recording device.
- Restart MMSSTV so it picks up the recording-device change.
Use these settings, as recorded in the successful test and the archival README:
| Location | Setting | Value |
|---|---|---|
| Main RX window | RX Mode | Auto |
| Main RX window | Auto history | ON |
| Option → Setup MMSSTV → RX | Auto start | VIS only |
| Option → Setup MMSSTV → RX | Auto restart | ON |
| Option → Setup MMSSTV → RX | Auto resync | ON |
| Option → Setup MMSSTV → RX | Auto stop | OFF |
| Option → Setup MMSSTV → Misc | Sound Card → In | Default, with Stereo Mix selected as the Windows default recording input |
If MMSSTV lists the Stereo Mix input directly, it can be selected there. Play the WAV in a normal Windows audio player at its original speed and let MMSSTV listen to that input. It does not need to open the long WAV as an offline image file. Each new VIS signal identifies another PD120 frame, and automatic restart allows the next page to begin after the previous one.
The tested installation keeps received images in C:\Ham\MMSSTV\History\ as Hist*.bmp; the shortcut in the experimental workspace points to that directory. The observed history is BMP, not a recreation of the original PNG filenames. Preserve wanted received images separately and use the page number printed in each image to relate them to the source PNG sequence. Other installations may use a different program or history location.
The user also verified recovery after seeking to an arbitrary position inside a long WAV volume. A partial frame at that point may be unusable; keep playback running until the next complete VIS signal, when reception can start from the next page. Decoding through playback happens in real time. Keep the original playback speed and pitch and use the original WAVs to retain the tested signal timing.
This recovers a readable visual edition of the web content. It does not recreate the original HTML/MDX or guarantee pixel-identical copies of the source PNGs. The observed decoded pages contain SSTV artifacts. Successful individual, consecutive, and seek-recovery tests have been reported; a page-by-page audit of every frame across the entire archive has not been recorded.
Archival README and the verified audio edition
The audio edition verified on 2026-09-10 contained 1,724 input PNGs, 1,724 individual WAVs, and these six final volumes. Their sizes and PCM format were checked directly:
| Volume | Bytes |
|---|---|
combined-with-gaps_001.wav | 3,895,596,294 |
combined-with-gaps_002.wav | 3,895,692,294 |
combined-with-gaps_003.wav | 3,895,692,294 |
combined-with-gaps_004.wav | 3,895,692,294 |
combined-with-gaps_005.wav | 3,895,692,294 |
combined-with-gaps_006.wav | 1,708,205,794 |
| Total | 21,186,571,264 |
All six volumes are 48 kHz, 16-bit mono PCM WAVs. These counts and sizes describe this edition; a later build or different PNG selection may produce a different set.
For an archive deposit, place the archival README.txt alongside the original numbered WAV volumes. It explains what the audio contains, the PD120 mode, volume order, the tested MMSSTV settings, audio routing, recovery after seeking, and the archival purpose; it includes a short Czech explanation as well as the English instructions.
The copied scripts/sstv-audio/README.txt already names the six volumes and their total size above. For a later edition, update those edition-specific details in the distribution copy to match its actual outputs. Preserve the source copy in scripts/sstv-audio/README.txt as the record of the tested edition. The scripts do not generate or update that archival README automatically. Retaining the corresponding png-manifest.json with the archive also preserves the article-to-page mapping. Archive publication, including an Internet Archive upload, is a separate manual step.
Calibration images and a systematic comparison of source and received pages remain future work. Basic PNG-to-WAV encoding, joined-volume playback, and visual recovery are now implemented and experimentally verified as described above.
Spotify for Creators publication workaround
When publishing this SSTV export through Spotify for Creators, the original large WAV volumes uploaded successfully, appeared as published, and were playable in Creators, but an ingest problem prevented the episodes/show from appearing publicly. Spotify support confirmed that the issue was related to the audio file type. Converting the same volumes to MP3 and uploading them again resolved public publication. For future Spotify uploads of this export, use MP3 copies of the volumes; this records the workaround for this specific ingest issue, not a general claim that Spotify does not support WAV.
9.6.9 JSON-LD article export and vocabulary
The structured article export is generated automatically alongside ALL_POSTS.txt by npm run generate:all-posts, which is part of the web and USB build variants. The Arweave build carries both files into dist-arweave/. The standalone command remains available for an existing build:
node scripts/export-site-json.mjsBy default, it writes dist/ALL_POSTS.json as UTF-8 JSON without BOM, also valid as JSON-LD 1.1. It uses the existing Node.js and Cheerio project dependencies. The exporter reads an existing build; it does not start Astro, fill translation caches, or access the network. export-all.bat consumes the built JSON and no longer regenerates it.
The JSON is initially generated after the English postprocess and TXT generation. After USB or Arweave path rewrites, the respective rewrite script refreshes the JSON when it is present, so vm:articleHtml and vm:builtHtmlSha256 describe the final HTML files. The following integrity pass includes ALL_POSTS.json in that build’s checksum manifest; signed variants cover it through the signed manifest.
The finished dist/ALL_POSTS.txt supplies article selection, order and original metadata. Full content comes from the corresponding built HTML files, including postprocessed English titles and bodies. Both directory routes and flat USB article files are supported. The export includes all indexed article language versions; it does not automatically include homepage-only content, standalone pages such as this documentation or /ns/, or the contents of linked downloads.
The root is a Schema.org Collection with article records in hasPart. Each record is a BlogPosting with its canonical URL, rendered headline, language, section, publication date, author, full articleBody, structured media references, and translation links when the counterpart is present. Code blocks in the JSON article body are not truncated. vm:articleHtml separately retains the parsed article HTML fragment, including its heading and visible metadata. Images and other media remain references; binary files and external embed payloads are not downloaded or bundled.
The inline @context combines Schema.org with the custom namespace https://vojtamaur.cz/ns/. The vm: vocabulary documents every custom property and the vm:Link class, including JSON types, scopes, examples and interpretation limits. For example, vm:slug expands to https://vojtamaur.cz/ns/slug. The namespace ends with /, not #; anchors on the vocabulary page are navigation aids, not replacement identifiers. Each known term also has a static description page, for example /ns/slug/.
Vocabulary descriptions are maintained in src/lib/export-vocabulary.ts. The static route src/pages/ns/[...term].astro generates /ns/ and the individual term pages from this shared data. These ordinary website pages become public when the built website is deployed; no separate service or runtime is required. Keep this vocabulary synchronized whenever custom export fields change. The export’s inline context remains self-contained and does not fetch the vocabulary page during JSON-LD processing.
Important interpretation rules:
vm:sourceMetadatais a JSON literal (@type: "@json"); its original index keys are not interpreted as Schema.org terms. The index’sTITLEmay remain Czech even for an English article, whileheadlinecomes from the rendered language version.vm:positionrecords article order starting at 1, since ordinary JSON-LD arrays do not imply RDF ordering. Each language version counts separately invm:articleCount.vm:builtHtmlSha256hashes the whole selected HTML file;vm:sourceIndexSha256hashes the exact input index bytes. The JSON inside the build omitsvm:declaredBuildSha256: it cannot contain the final hash of a build that includes itself. An explicit snapshot written underexports/may copy the declared hash from an existingBUILD_SHA256.txt; this does not verify the manifest or OpenPGP signature.vm:generatedAtis the JSON export time, not an article modification date.datePublishedretains the index date, including legacy first-of-month migration values described in section 5.5.1.vm:languageStatusrecords the existing English fallback convention based on thenoindexmarker. It is not an independent translation-quality assessment, and English records may contain Czech ALT text or untranslated passages.vm:sourcePathrecords the source path from the finished index. The exporter does not mix the current MDX contents into an older build.
Useful options:
node scripts/export-site-json.mjs --dry-run
node scripts/export-site-json.mjs --dist dist-arweave --dry-run
node scripts/export-site-json.mjs --dist dist --output exports/ALL_POSTS.json
node scripts/export-site-json.mjs --project-root "H:\vojtamaur-web"Relative paths are resolved from the project root, which defaults to the parent of scripts/. Without --output, the destination is ALL_POSTS.json at the root of the directory selected by --dist (default: dist/). An explicit output may be that same build file, or a .json / .jsonld snapshot under the project’s exports/ directory and outside the selected build. Such separate snapshots are not included in the build signature.
Dry run validates the inputs and constructs the export without writing it. Missing article files, malformed metadata, duplicate records and unsafe output paths cause a nonzero exit; successful writes replace the output through a temporary file. The exporter reports article count, byte size and the SHA-256 of the resulting JSON. Rerunning the standalone exporter into an already checksummed or signed build changes the export, including its generation time; regenerate integrity metadata and sign again when applicable, or use --dry-run / an explicit snapshot under exports/ to preserve the existing build.
9.6.10 Rosetta multilingual text export
scripts/export-site-rosetta.mjs creates a machine-translated reading and archival edition of the articles using the Google Gemini Batch API. It runs manually with Node.js 20 or newer and the project’s existing Cheerio dependency. It reads the JSON-LD article export described above and writes its own files under exports/; it does not build the site, change published pages, or update the DeepL English translation cache. The Gemini API used here is separate from the Gemini/Gopher publishing format described in section 9.6.7.
Source and translation boundaries
By default, the script reads dist/ALL_POSTS.json from the current build. It does not search exports/ or automatically select a dated snapshot. If the default file is missing, create a build or select an existing JSON / JSON-LD snapshot with --input. Only records with inLanguage == "cs" are translated, in vm:position order. English is also translated from Czech for this edition; the existing English article records are not reused. The 2026-09-23 source contained 82 Czech articles and 82 English counterparts; later exports can have different counts.
articleBody supplies the linear text, including headings, captions and image ALT descriptions. vm:articleHtml supplies the matching protection map for notranslate, translate="no", and code/output elements such as pre, code, kbd and samp. Ordinary prose and the article headline are translated unless protected. Structured media remain references in the text; media files are not downloaded or bundled.
Before sending text to Gemini, the exporter replaces protected spans with placeholders and retains their originals locally. These include complete [CODE BLOCK]...[/CODE BLOCK] sections, explicitly protected passages and program output, inline code, URLs, slugs, paths, filenames, hashes, IDs, numeric data and other recognized technical tokens. Source line/paragraph separators and metadata are assembled locally. Protected content is restored unchanged after translation. Technical or experimental text that is not identifiable by these rules must be marked notranslate in the source before generating the JSON-LD snapshot.
Preparation verifies HTML-to-body mapping, paired code markers and an exact protection/restoration round trip. Response validation rejects missing, duplicated or invented placeholders, missing text items and invalid structural boundaries. These checks preserve the archive’s structure and protected text; they do not establish the linguistic accuracy of a machine translation. Each TXT identifies the model and states that its translation has not been independently reviewed.
Preparation and translation commands
Run commands from the project root. Paths supplied to the exporter resolve relative to that root. Offline preparation needs no API key and makes no API calls, but does create or update the local rosetta.json state file:
node scripts/export-site-rosetta.mjs --list-languages
node scripts/export-site-rosetta.mjs prepare
node scripts/export-site-rosetta.mjs run --languages de,ja --limit 1 --dry-runThe default selection is all Czech articles and a fixed list of 98 target languages. --languages accepts comma-separated codes; --limit restricts the selection to the first N Czech articles. The language list is a recorded selection derived from Gemini Live’s 99-language list with Czech excluded, not a translation-quality guarantee for the selected model. --list-languages shows the exact supported CLI codes and full names.
For API operations, set GEMINI_API_KEY (or the fallback GOOGLE_API_KEY) in the launching shell. The exporter reads the environment and does not save the key in its output. The API project must have access and quota for the selected model and Batch API. Translation submissions and explicit retries may incur API charges; offline preparation does not check account billing or quota.
Start with a small translation, or run the complete selection:
node scripts/export-site-rosetta.mjs run --languages de,ja --limit 1
node scripts/export-site-rosetta.mjs runThe default model is gemini-3.1-flash-lite; --model overrides it. run prepares the selection, submits missing work and waits for the batch, checking its state every 60 seconds by default (--poll-seconds overrides this). Running without --languages or --limit expands a completed pilot to all articles and target languages while reusing saved results. Collect an outstanding batch before changing its selection.
To submit and collect in separate sessions:
node scripts/export-site-rosetta.mjs submit
node scripts/export-site-rosetta.mjs collectsubmit uses the saved preparation. collect checks the saved batch and downloads its results when ready; if it is still running, call collect again later or use run with the same selection to wait. After results have been collected, another collect rebuilds and validates the TXT files locally without another API call. Console messages identify language codes and names, queued work, available batch progress, completed exports and article/part failures. Detailed per-language completion becomes available when Gemini returns the results.
Output directory and resumable state
All languages share one directory, with one TXT per complete language and one supporting JSON:
exports/
rosetta/
ALL_POSTS__lang-de.txt
ALL_POSTS__lang-ja.txt
ALL_POSTS__lang-uk.txt
...
rosetta.jsonTXT files use UTF-8 with BOM. Filenames always follow ALL_POSTS__lang-CODE.txt, with no __section-all suffix. Existing outputs with the old suffix are migrated after their saved checksums are verified. A language TXT is written only when every required part for the saved article set passes validation. A pilot TXT therefore contains only its limited selection; its header records that article count. When expanding to the full archive, a stale partial pilot file is removed until the complete language can be exported.
rosetta.json is both the language manifest and the resumable working state. Its manifest.languages entries map code, full language name and filename, with completion status, article counts and output checksums. The same file retains the source identity, protected originals, translation plan, API results, recovered valid text items, batch history and validation errors. Keep it alongside the TXT files: deleting it discards the local cache needed to avoid repeating paid translations. No per-language subdirectories or separate request, preview or error files are required.
--output exports/rosetta-NAME selects another shared directory under exports/; pass that option consistently when preparing, submitting and collecting that edition. A different source snapshot or model requires a different output directory. After rebuilding the site, use a new output directory for the new JSON, or keep selecting the exact archived input with --input when continuing an earlier edition. submit and collect reuse the saved configuration and do not need --input. For example:
node scripts/export-site-rosetta.mjs prepare --input exports/ALL_POSTS-2026-09-23.json --output exports/rosetta-2026-09-23
node scripts/export-site-rosetta.mjs run --input exports/ALL_POSTS-2026-09-23.json --output exports/rosetta-2026-09-23Failed items and recovery
A Gemini batch marked SUCCEEDED can still contain truncated, blocked or structurally invalid translations. The exporter retains valid work, writes complete languages and exits with an error if selected parts remain invalid. Inspect errors in rosetta.json for the language, article, part and pending text-item IDs. There is no automatic paid retry.
To request the remaining failed text items explicitly:
node scripts/export-site-rosetta.mjs run --retry-failedThe repair path preserves valid items even from incomplete responses and submits only invalid or missing items from failed parts as individual requests. Already valid translations are reused. If another item still fails, the next explicit retry targets only what remains; it does not restart complete languages. Validation is not relaxed to accept missing source information.
An interrupted run keeps its batch ID and results in rosetta.json. If a submission’s outcome is uncertain and no batch ID was saved, locate the actual job in Google AI Studio and recover it with collect --batch-name batches/ID (replace ID with that job’s identifier), adding the same --output when applicable. Do not submit a duplicate job blindly. Completed exports from an older timestamped Rosetta directory can be imported into the shared directory without API calls using node scripts/export-site-rosetta.mjs import --from exports/rosetta-OLD.
Published Rosetta reading pages
Finished translations selected for publication live in the stable, versioned public/rosetta/ directory, using ALL_POSTS__lang-CODE.txt filenames. Do not put the Rosetta English TXT here: /en/ remains the existing full English homepage. The build never selects files from exports/, translates content, calls an API, or modifies the selected TXT files.
/rosetta/ is a directory of available languages with their native names. Each language has a small page, for example /de/, containing a localized notice, the text-translation edition date, and links to the TXT, the Czech and English homepages, and the language directory. The notice explains that the language URL leads to an unreviewed machine translation in plain text that may be out of date. No article text is embedded in these HTML pages. They are ordinary static pages, not 404 recovery responses; unsupported languages retain the normal 404 behavior. Right-to-left notices use the appropriate document direction.
The localized notices and link labels are maintained in src/lib/rosetta-copy.ts. src/lib/rosetta.ts discovers actual TXT files and reads their headers. Only exports with the complete Czech article set are accepted. The visible date describes the text-translation edition, never the source text or the current website build time. The directory has Czech and English introductions stacked above the shared native-language list.
A repository-root rosetta.json can retain the export’s original working state. When present, the build uses manifest.updatedAt (recorded after TXT export) for the visible translation-edition date and checks each selected TXT’s completion status, source identity, article count and SHA-256 against its manifest. An incomplete overall manifest is allowed when the selected TXT languages are complete. Missing TXT files are not advertised. Source dates in TXT headers remain validation metadata. When the JSON or its TXT export timestamp is absent, the translation-date line is omitted rather than mislabeling a source date or a copied file timestamp. Neither development nor builds modify this JSON.
rosetta.json must remain outside public/. It is excluded from the public source ZIP as well as from the website itself. The TXT files are copied once into dist/rosetta/ and included in normal sitemap and integrity generation. To avoid another copy inside the public source ZIP (and its JPEG carrier), public/rosetta/ is treated as an external asset directory in the source-bundle manifest. The existing asset downloader restores and verifies those TXT files when reconstructing the source package. The root JSON is unnecessary for reconstruction.
To update the published edition, manually replace the selected TXT files in public/rosetta/ and, if kept, their matching root rosetta.json. Adding a new supported TXT automatically adds its landing page and directory entry; removing it removes both. Add localized UI copy before publishing an entirely new language code. Do not mix files from different source editions with a retained manifest.
Run npm run test:rosetta after a web or USB build to verify the directory, language pages, dates, links, untouched TXT bytes, integrity inclusion and absence of public JSON. USB builds use flat HTML pages with relative links; the Arweave copy also carries the Rosetta files through its existing path rewrite.
9.6.11 Morse-code text export
scripts/export-site-morse.py creates a manual archival edition from the existing dist/ALL_POSTS.json, using only the Python standard library. It reuses the JSON parser, language selection, article order, full articleBody and compact media serializer from scripts/filter-all-posts.py. It does not encode raw JSON syntax or JSON-LD bookkeeping. Each article retains its rendered title and month, language, section and public URL, followed by compact text and media references (image filenames, alt text, captions and embed sources). Explicit ARTICLE / END ARTICLE lines mark article boundaries, and the closing code marker becomes END CODE BLOCK before normalization.
npm run export:morse -- --lang cs
npm run export:morse -- --lang en
npm run export:morse -- --lang bothThe default is both, which writes one combined file containing Czech and English entries in source order. All sections are included. Filenames follow the existing compact export convention:
exports/ALL_POSTS__lang-cs__section-all__format-morse.txt
exports/ALL_POSTS__lang-en__section-all__format-morse.txt
exports/ALL_POSTS__lang-cs+en__section-all__format-morse.txtThe pipeline is JSON → compact text with article metadata → normalized ASCII → International Morse. Normalization uses Unicode NFKD decomposition and uppercase A–Z / 0–9. The versioned, fixed scripts/morse-glossary.json supplies letter transliterations such as ł → L and named symbols such as π → PI, φ → PHI, × → TIMES and ≈ → APPROXIMATELY. Named symbols become separate words. Remaining punctuation becomes word separators; combining marks and invisible Unicode formatting controls are removed. Unsupported characters become UNICODE followed by their uppercase hexadecimal code point, for example 天 → UNICODE 5929, rather than disappearing silently. This is a lossy reading representation: punctuation, case, diacritics and indentation cannot be reconstructed exactly.
Horizontal whitespace becomes single spaces; every LF and blank line in the compact representation is preserved. As with compact text export, redundant prose blank lines are already removed by the shared serializer, while blank lines inside code remain. CRLF, CR and Unicode line separators normalize to LF. The final file contains only ., -, spaces, / and LF: Morse letters are separated by a space and words by /. It is ASCII without BOM, with LF even on Windows; any header is Morse-encoded too. The same input, glossary and options produce the same bytes, and the command reports article counts and SHA-256.
Paths resolve from the project root (the parent of scripts/, or --project-root). Explicit inputs must be JSON-LD article archives; legacy TXT is not accepted. An explicit output must remain a .txt file under exports/ and cannot overwrite the input. Missing requested languages are errors. Examples for archived input and validation without writing:
python scripts/export-site-morse.py --input exports/ALL_POSTS-2026-09-23.json --lang cs --output exports/morse-cs-2026-09-23.txt
python scripts/export-site-morse.py --lang both --dry-run
npm run test:morseexport-all.bat runs the bilingual Morse export after its compact text exports. No normal web, USB, Arweave or Gemini build invokes it; the build continues to generate the source JSON, and the manual export leaves dist/ and its integrity manifests unchanged. The generated exports/ files remain excluded from version control.
9.6.12 MoM-like text-only PDF export
scripts/export-site-mom-like.py creates an additional text-only edition inspired by the supplied Memory of Mankind PDF. It shares the JSON parser, article selection and compact serializer in scripts/filter-all-posts.py. It does not import or change export-site-pdf-ultra.py.
Install the two Python dependencies and run it manually:
python -m pip install -r requirements-mom-like-export.txt
python scripts/export-site-mom-like.pyThe npm equivalent is npm run export:mom-like. The default input is the existing dist/ALL_POSTS.json, and the default selection is Czech articles in volna-tvorba. No browser, local media assets, network access or new site build is required. export-all.bat runs this exporter twice, creating separate Czech and English PDFs with the default volna-tvorba selection and 4 pt font size. The Czech PDF uses the default path below; the English PDF uses exports/vojtamaur-web-export-mom-like-en.pdf, with its own TXT and manifest. The batch setup includes requirements-mom-like-export.txt. Normal builds, signing and integrity generation do not invoke this exporter.
Running the script without selection options is equivalent to:
python scripts/export-site-mom-like.py --lang cs --section volna-tvorbaThe section default also applies to --lang en and --lang both. An explicit --section replaces the default rather than adding to it. For example, --section vystavy exports only exhibitions; --section all exports every section. The visible Selection line and the manifest record the actual selection. Article and page counts depend on the input snapshot and selected content.
The default outputs are:
exports/vojtamaur-web-export-mom-like.pdf
exports/vojtamaur-web-export-mom-like.txt
exports/vojtamaur-web-export-mom-like.manifest.jsonThe PDF uses three columns, embedded Arial at 4 pt, character spacing of -0.2 pt, a 3.5 pt line advance, justified text and square 210 x 210 mm pages. Equal 7 mm margins create a square 196 x 196 mm text frame, with 1 mm gaps between columns. Any remaining fraction of a line at the bottom of the grid is divided equally above and below it. The negative character spacing matches the reference Word document (w:spacing=-4 twentieths of a point), and is applied both when measuring line breaks and when drawing glyphs. Use --char-spacing to override it. --margin-mm sets all margins unless --top-margin-mm is supplied explicitly. On systems without Arial, an installed Liberation Sans or DejaVu Sans is used; --font supplies an explicit TrueType font for reproducible typography. The manifest records font paths and hashes.
Set --font-size to change the text size in points; the default remains 4 pt. Decimal values such as --font-size 4.5 are supported. For example, this creates a separate edition with 5 pt text:
python scripts/export-site-mom-like.py --font-size 5 --output exports/mom-like-5pt.pdfLine advance scales with the font size: it is font-size * line-height, with --line-height 0.875 by default (4.375 pt at a 5 pt font size). Page dimensions and margins stay unchanged; larger text generally requires more pages. Character spacing remains -0.2 pt unless overridden with --char-spacing.
Original logical line breaks become gaps of seven spaces before justification. Article titles and dates are bold, while image records and prose remain ordinary. Titles still flow continuously without forced article breaks, thumbnail grids, background bands or page numbers. Emphasis uses fill-and-stroke text with the same embedded glyphs and advances, so it does not increase page count. Only actual article headings are emphasized, including across line, column and page boundaries; title/date-like strings inside code are unaffected. The visible opening label and PDF title use ALL_POSTS; MoM-like remains an internal exporter name. Images retain their filename, ALT and CAPTION as text; video, PDF and interactive records retain their source references. Media is not embedded or deduplicated. URLs and filenames can wrap at existing separators such as slashes, hyphens and underscores, making use of the remaining line width. Oversized unbroken words can wrap without added hyphens. Links are visible text, not clickable annotations.
By default, the DullGPT example output log is shortened to the same preview limits used by ALL_POSTS.txt: the first 120 lines, capped at 12000 characters. A visible [TRUNCATED: ...] note reports the omitted part, and the manifest records counts per language. This applies only to the recognized output log in the dullgpt article; its executable Python program and all other articles remain complete. The companion TXT has the same explicitly shortened log. --full-dullgpt-log restores the full log from JSON. The source JSON and original MoM Word/PDF are never modified.
This flowing treatment also applies to source-code blocks in the PDF. The companion UTF-8 TXT retains the compact serializer’s logical lines, blank lines within code and code indentation. Use that TXT or the JSON/source repository when recovering executable code. It contains the compact content, not raw JSON or the original complete website markup. JSON provides full code blocks before the explicit DullGPT log preview; an explicitly selected legacy TXT snapshot can already contain truncated code that this exporter cannot recover.
Unicode fallback fonts are used when installed. Characters unsupported by the available fonts, and supplementary-plane characters whose PDF Unicode mapping is not reliably supported by the renderer, are printed explicitly as [U+XXXX] (or a longer code point) in the PDF. The TXT preserves the original characters. Counts are reported on the console and in the manifest. Before replacing output files, the exporter verifies the PDF’s page dimensions, absence of images and exact non-whitespace text against the prepared PDF content. The manifest includes this check, article selection, layout and SHA-256 hashes of the input, PDF and TXT.
Each default run creates a new timestamped PDF, TXT and manifest. An explicit --output replaces the files at that exact path; use distinct custom paths to retain experiments with different fonts, languages or sections.
Examples:
python scripts/export-site-mom-like.py --lang cs --output exports/mom-like-cs.pdf
python scripts/export-site-mom-like.py --lang en --output exports/mom-like-en.pdf
python scripts/export-site-mom-like.py --section vystavy
python scripts/export-site-mom-like.py --section all
python scripts/export-site-mom-like.py --lang both --output exports/mom-like-both.pdf
python scripts/export-site-mom-like.py --section volna-tvorba,vystavy --from-date 2020-01-01 --to-date 2026-12-31
python scripts/export-site-mom-like.py --page-width-mm 210 --page-height-mm 210
python scripts/export-site-mom-like.py --columns 3 --font-size 4 --char-spacing -0.2 --line-height 0.875 --line-gap-spaces 7
python scripts/export-site-mom-like.py --full-dullgpt-log --output exports/mom-like-full-log.pdf
python scripts/export-site-mom-like.py --input exports/older-ALL_POSTS.json --page-numbers
python scripts/export-site-mom-like.py --help--section defaults to volna-tvorba. An explicit selection replaces that default and can be repeated or comma-separated. Use --section all to include every section. The source article order is preserved. Bilingual selections label each article’s language. Relative input and output paths are resolved against the repository root, inferred from the script location or supplied with --project-root. An explicit --input dist/ALL_POSTS.txt remains supported for legacy structured snapshots.
Run the content-preservation and pagination regression checks with:
python -B -m unittest discover -s tests -p test_mom_like_export.py9.7 Arweave / Permaweb build
The Arweave / Permaweb build is created with:
npm run build:arweaveThis command first runs the strict standard web build, then runs the Arweave postprocess, and finally regenerates integrity files for the finished Arweave output:
node scripts/make-arweave-build.mjs
npm run generate:integrity:arweaveThe script creates a separate output directory:
dist-arweave/This directory is intended for upload to ArDrive or another Arweave-compatible upload tool.
The Arweave build exists because a normal web build uses root-relative paths such as:
/_astro/...
/images/...
/files/...
/volna-tvorba/These paths work on a normal domain root such as https://vojtamaur.cz/, but they do not work automatically when the site is served below an Arweave manifest transaction URL.
The Arweave postprocess copies the finished dist/ output into dist-arweave/ and rewrites root-relative references into relative paths appropriate for the location of each HTML or CSS file.
After this copy-and-rewrite step, the build runs the integrity generator again for dist-arweave/. This is necessary because the Arweave postprocess can modify files after the normal web build has already created integrity files for dist/.
The integrity files inside dist-arweave/ belong to the Arweave deployment output itself:
dist-arweave/SHA256SUMS.txt
dist-arweave/BUILD_SHA256.txt
dist-arweave/BUILD_HASH_HISTORY.txt
dist-arweave/integrity.jsonThe source of truth remains the normal project source. dist-arweave/ is a generated deployment artifact and should not be edited manually.
10. Exact special build logic
The build:usb script is defined as a USB-targeted Astro build followed by postprocessing. In the current translation-aware workflow, the effective order is:
set "BUILD_TARGET=usb" && astro build && node scripts/en-postprocess.mjs && npm run generate:all-posts && npm run generate:source-bundle && node scripts/usb-rewrite.mjs && npm run generate:integrityThis implies seven steps:
BUILD_TARGET=usbis set- the Astro build is run in the mode defined in
astro.config.mjs scripts/en-postprocess.mjsapplies the English translation layer using the available cachescripts/generate-all-posts.mjscreatesdist/ALL_POSTS.txt, embeds the text indist/404.html, and callsscripts/export-site-json.mjsto createdist/ALL_POSTS.jsonscripts/generate-source-bundle.mjscreates the source archives and copies their previous-build history snapshot intodist/BUILD_HASH_HISTORY.txtscripts/usb-rewrite.mjsrewrites root-based URLs into relative file paths, including links to the source ZIP and history, then refreshes the existingdist/ALL_POSTS.jsonfrom the rewritten HTMLscripts/generate-integrity.mjscreatesdist/SHA256SUMS.txt,dist/BUILD_SHA256.txt,dist/integrity.json, and an unsigneddist/SIGNING_STATUS.txtfor the final static output, recordingbuildType: "usb"
npm run build:usb:translate:signed uses the same successful-build order through scripts/build-usb-translate.mjs, with missing translations enabled. The runner attempts the USB rewrite even if an earlier step fails, but exits unsuccessfully and does not reach signing in that case. Only a successful underlying build reaches npm run sign:build; after signature verification and metadata/status updates, the signer appends a usb record to the canonical source history.
10.1 What astro.config.mjs does in USB mode
When BUILD_TARGET=usb, the following are used:
trailingSlash: "never"build.format: "file"
This means that internal routes are generated as files of the form:
slug.htmlinstead of the directory form:
slug/index.html10.2 What scripts/usb-rewrite.mjs does
The script:
- goes through all
.htmlfiles in thedistdirectory - reads their content
- rewrites selected root-based paths to relative paths
- saves the files again
Specific rewrites include:
- generated Astro assets in
/_astro/ - public assets in
/images/,/files/,/demos/, and/source/ - homepage links to
/ - internal links in the form
/slug/ - nested English paths such as
/en/slug/
The script calculates the relative path from each HTML file back to the dist/ root. This matters because a root-level file such as dist/about.html needs ./_astro/..., while a nested file such as dist/en/about.html needs ../_astro/....
The purpose of this step is to adjust the HTML so that the output works even outside a standard web server with root-relative URLs.
10.3 Exact build:arweave logic
The build:arweave script is defined as a strict standard web build followed by Arweave-specific postprocessing and a second integrity pass for the final Arweave output. The effective order is:
npm run build:web:strict && node scripts/make-arweave-build.mjs && npm run generate:integrity:arweaveThis implies five main steps:
- the standard Astro web build is created in
dist/ scripts/en-postprocess.mjsverifies or applies the English translation layer according to the strict build workflowscripts/generate-all-posts.mjscreatesdist/ALL_POSTS.txtanddist/ALL_POSTS.json,scripts/generate-source-bundle.mjscreates the source archives and history snapshot, andscripts/generate-integrity.mjscreates integrity files fordist/scripts/make-arweave-build.mjscopiesdist/todist-arweave/, rewrites root-relative paths, and refreshes the copieddist-arweave/ALL_POSTS.jsonfrom the rewritten HTMLscripts/generate-integrity.mjs dist-arweave arweavecreates fresh integrity files for the final Arweave output and recordsbuildType: "arweave"
The second integrity pass matters. The checksum files created during build:web:strict describe dist/. After scripts/make-arweave-build.mjs rewrites files in dist-arweave/, those original checksums would no longer be sufficient for the Arweave output. Therefore dist-arweave/ gets its own SHA256SUMS.txt, BUILD_SHA256.txt, and integrity.json.
The copied BUILD_HASH_HISTORY.txt snapshot remains unchanged and is included in the new Arweave manifest. npm run build:arweave:signed records only the verified final Arweave build hash with type arweave; its intermediate unsigned web build does not add a history record.
The Arweave output keeps the directory-style route model of the normal web build, for example:
fotogrammetrie/index.html
en/fotogrammetrie/index.htmlThis is different from the USB build, which uses a file-based output model.
The purpose of dist-arweave/ is to create a static folder that can be uploaded to ArDrive and then exposed through an Arweave manifest. The manifest should be created at the level where index.html, _astro/, images/, files/, ALL_POSTS.txt, ALL_POSTS.json, and ARCHIVE.txt are directly present.
10.4 Exact Gemini build logic
The standalone Gemini build is effectively:
npm run build:web:strict && npm run generate:geminiThis creates or verifies the strict bilingual HTML build in dist/ and then replaces dist-gemini/ with the generated capsule.
For rapid iteration on only the converter, the existing HTML build can be reused:
npm run generate:geminiThis avoids rebuilding Astro or rerunning translation checks after every parser change.
The translating signed production command reaches the Gemini generator through the underlying translating web build. Its relevant order is:
Astro build
-> English postprocess with missing translations enabled
-> ALL_POSTS.txt, ALL_POSTS.json and source-bundle generation
-> integrity generation for dist/
-> Gemini capsule generation in dist-gemini/
-> detached OpenPGP signing and verification of dist/SHA256SUMS.txt
-> signed status/metadata updates
-> append the web build hash to the canonical source BUILD_HASH_HISTORY.txtThe final signing step remains scoped to the normal web build. dist-gemini/ is generated in the same workflow but is a separate unsigned artifact.
The capsule copies BUILD_HASH_HISTORY.txt from the completed web output so that the article’s history link works in the text edition. During this workflow the copied file contains the same previous-build snapshot as dist/. Capsule generation itself adds no history record, and the later source-history append does not refresh the capsule copy.
11. English translation workflow
The project has an English version generated at build time. The Czech MDX files remain the source of truth. The English version is a derived static artifact produced from the rendered HTML output.
There are two separate translation layers:
- Manual UI translation – stable interface text is kept in
src/lib/i18n.ts. This covers the header, section names, homepage text, and other short visible strings. The small Czech/English labels insideOpenPgpContact.astroare a deliberate local exception because the component shares one cryptographic payload between both homepage variants. - Automatic content translation – article body regions are translated after
astro buildbyscripts/en-postprocess.mjs. The translated fragments are cached intranslations/en/.
This separation is intentional. UI text is small and highly visible, so it is translated manually. Article content is larger and less convenient to maintain twice, so it is translated automatically.
11.1 Translation configuration
The translation configuration is stored in:
scripts/i18n-config.mjsImportant values:
- source language:
CS - target language:
EN-US - cache directory:
translations/en - translated regions:
[data-i18n="translate"] - maximum translated fragment size:
80_000bytes
The translation postprocess protects code, embeds, scripts, styles, SVG, canvas, iframes, and anything marked as notranslate or translate="no". Image alt text and thumbnailAlt metadata are not automatically translated. Captions are translated when they are normal HTML text inside a translated region.
11.1.1 Translation glossary and terminology preservation
Project-specific terminology used by the automatic translation layer is stored in:
translations/glossary-cs-en.tsvThe TSV file is the canonical source of the glossary. Each non-empty line contains exactly one Czech source term, one tab character, and one English target term:
Volná tvorba Personal WorkThe tab must be a real tab character, not a sequence of spaces. Duplicate Czech source terms, empty values, extra columns, and leading or trailing whitespace cause the translation postprocess to fail instead of silently accepting an ambiguous glossary.
The glossary serves two related purposes:
- translation consistency – project-specific names and concepts are translated according to the intended terminology rather than according to a generic literal interpretation
- semantic and archival documentation – the repository preserves how the author intended a Czech term to be understood in English
For example, Volná tvorba is mapped to Personal Work, not the superficially literal Free Creation. The TSV file therefore documents part of the conceptual structure of the project, not merely a temporary setting for an external translation service. Because it is stored and versioned with the source code, this intended terminology remains available in archived repository copies even if DeepL or the current build environment is no longer available.
The glossary integration is configured in scripts/i18n-config.mjs:
glossary: {
enabled: true,
name: "vojtamaur.cz CS-EN",
sourceFile: "translations/glossary-cs-en.tsv",
stateFile: "translations/glossary-cs-en.state.json"
}During a build:*:translate run, scripts/en-postprocess.mjs:
- loads and validates the local TSV file
- reuses the remote DeepL glossary ID recorded in
translations/glossary-cs-en.state.jsonwhen possible - otherwise searches the DeepL account for a glossary with the configured name
- creates the remote glossary if no matching glossary exists
- updates the remote glossary when its entries differ from the local TSV source
- stores the resulting remote ID and local revision in the state file
The remote glossary ID does not need to be set manually as an environment variable. Only DEEPL_AUTH_KEY is required. The state JSON file is operational metadata and a reusable pointer to the remote DeepL object; it is not the authoritative glossary content. If the state file is missing, it can be recreated from the TSV and the DeepL account. If the TSV is missing, the intended terminology is no longer reliably reconstructable from the project source.
Glossary-aware cache invalidation is selective. For each translated fragment, only glossary rows whose Czech source phrase occurs in that fragment contribute to its glossary revision. Changing one glossary entry therefore invalidates translations that use that term without forcing unrelated pages to be translated again.
Translation requests are paced and retried with exponential backoff to reduce DeepL rate-limit failures. If DeepL still returns HTTP 429, translations completed before the failure remain cached and the same translate command can be run again.
11.2 DeepL API key
The DeepL key is not stored in the repository. It is provided through the DEEPL_AUTH_KEY environment variable.
In Windows CMD:
cd F:\vojtamaur-web
set "DEEPL_AUTH_KEY=YOUR_DEEPL_KEY"
npm run build:web:translateOne-line CMD version:
set "DEEPL_AUTH_KEY=YOUR_DEEPL_KEY" && npm run build:web:translateIn PowerShell:
cd F:\vojtamaur-web
$env:DEEPL_AUTH_KEY = "YOUR_DEEPL_KEY"
npm run build:web:translateThe key must never be committed to the repository, embedded in client-side JavaScript, or uploaded as a public file.
11.3 Web translation commands
Recommended workflow after changing content:
cd F:\vojtamaur-web
set "DEEPL_AUTH_KEY=YOUR_DEEPL_KEY"
npm run build:web:translate
npm run build:web:strict
npm run previewMeaning:
npm run build:web:translatefills missing translation cache entries and writes translated EN HTML intodist/.npm run build:web:strictchecks that no translated EN fragment is missing.npm run previewshows the finished build output.
For normal rebuilds with an already complete cache, npm run build:web can be enough. Before publishing, npm run build:web:strict is the safer check.
To force a full retranslation, use:
npm run build:web:refreshThis should be used carefully because it can change existing English output even if the Czech source text did not change. It is not a cache cleanup command.
To inspect unused English translation cache files without deleting them, use:
npm run build:web:prune:dryTo delete unused, unprotected English translation cache files, use:
npm run build:web:pruneThe prune commands run the strict workflow, so they should only be used when the translation cache is already complete.
11.4 USB translation commands
For the portable build with missing translation generation:
cd F:\vojtamaur-web
set "DEEPL_AUTH_KEY=YOUR_DEEPL_KEY"
npm run build:usb:translateFor a strict USB check, use:
npm run build:usb:strictTo inspect unused English translation cache files during the USB workflow, use:
npm run build:usb:prune:dryTo delete unused, unprotected English translation cache files during the USB workflow, use:
npm run build:usb:pruneThe USB wrapper scripts run usb-rewrite.mjs even if the translation postprocess fails. This prevents dist/ from being left with broken relative CSS and image paths. A rewritten dist/ is not proof of a successful translation build; the log must still be checked for [i18n] Postprocess failed:.
11.5 Translation cache
Translation cache files are stored in:
translations/en/Each cache entry contains the original source fragment, the translated fragment, and metadata about the translation configuration. The cache key is derived from the route, purpose, source fragment, language direction, DeepL options, selector policy revision, and any glossary rows that apply to the fragment.
Practical consequences:
- unchanged content is not sent to DeepL again
- changed Czech content gets a new cache key
- changes to link attributes in unprotected translated HTML also change its cache key, even when the visible Czech text stays the same; the corresponding source and translated fragments must both match the updated link behavior
- changing a glossary entry invalidates fragments that contain its Czech source term
- changing an unrelated glossary entry does not invalidate every translated page
- changing other translation policy can invalidate old cache entries
- old cache entries can remain in
translations/en/after content changes unless the cache is pruned - deleting
translations/en/forces translation generation from scratch and should not be used as routine cleanup
The cache is part of the project source, not a public runtime dependency. The published site uses the finished HTML in dist/.
11.6 Translation cache pruning
Translation cache pruning is handled by scripts/en-postprocess.mjs when the EN_PRUNE_CACHE environment variable is enabled. The script records every cache hash that is actually used while processing the current English build output. After that, it can remove cache files that are no longer used by the current site.
Recommended web workflow:
npm run build:web:prune:dry
npm run build:web:pruneRecommended USB workflow:
npm run build:usb:prune:dry
npm run build:usb:pruneThe dry-run command must be checked first. It prints files that would be deleted and reports a summary such as cache-kept, cache-would-delete, cache-locked-kept, and cache-invalid-kept. The non-dry command deletes only unused cache files that are not protected.
A cache file is kept during pruning when at least one of these is true:
- its hash is used by the current build
- the JSON file cannot be read safely
- the cache entry has
manual: true - the cache entry has
locked: true - the cache entry has
edited: true
New automatically generated cache entries are marked as not manually edited:
"manual": false,
"locked": false,
"edited": false,
"editedAt": null,
"editedBy": nullWhen a translation is manually corrected, the cache entry can be marked like this:
"edited": true,
"editedAt": "2026-07-06",
"editedBy": "Vojta"This protects the entry from future pruning even if the current source text changes and the hash is no longer used. manual: true and locked: true are also respected as stronger preservation markers.
Pruning is meant to keep translations/en/ searchable and maintainable. It should not change the published HTML except through the normal strict build process. If the goal is only to remove unused cache files, do not use build:web:refresh; refresh can retranslate existing content and change manually checked output.
11.7 Development server versus final build
npm run dev is useful for editing layout and content. The final publishing artifact is still the build output in dist/.
In some cases, npm run dev may correctly display manually translated UI elements and metadata, while article bodies remain untranslated in English routes. This usually means the DeepL postprocess has not been applied to the current output yet.
To verify the real final translated output, use:
npm run build:web:translate npm run preview
11.8 MDX hard line breaks inside translated paragraphs
Avoid using Markdown hard line breaks between a bold pseudo-heading and the text that follows it in translated article content.
Problematic pattern:
**2026-03-02: TŘÍDĚNÍ PIXELŮ FOTEK**
Pokusil jsem se pomocí algoritmu roztřídit pixely fotografií.The two trailing spaces after the bold text are meaningful Markdown. They produce a hard HTML line break:
<p><strong>2026-03-02: TŘÍDĚNÍ PIXELŮ FOTEK</strong><br>
Pokusil jsem se pomocí algoritmu roztřídit pixely fotografií.</p>This structure is fragile when the fragment is sent through the English translation postprocess. DeepL receives one translated HTML paragraph containing both the bold title and the following sentence. During translation, it may preserve the <br> but move part of the translated title across it, producing output such as:
<strong>2026-03-02: SORTING PHOTO</strong><br><strong>PIXELS</strong>
I tried to sort the pixels...The visual result looks like a broken heading even though the browser is only rendering the hard break that already exists in the HTML.
Safer pattern when the bold line is only a compact label:
**2026-03-02: TŘÍDĚNÍ PIXELŮ FOTEK**
Pokusil jsem se pomocí algoritmu roztřídit pixely fotografií.This creates separate paragraphs and does not place a <br> between the label and the body text.
Best pattern when the line is structurally a subsection heading:
### 2026-03-02: TŘÍDĚNÍ PIXELŮ FOTEK
Pokusil jsem se pomocí algoritmu roztřídit pixely fotografií.Use real headings for repeated article sections such as dated concept entries. If the default heading style is too visually large, fix that in CSS instead of simulating headings with bold text and hard breaks.
After changing this kind of structure, verify the generated English output:
npm run build:web:translate
npm run build:web:strict
npm run previewIf the broken English output persists, check whether an old translation cache entry in translations/en/ is still being reused. The cache stores translated HTML fragments, so a structural mistake can remain visible until the changed source fragment produces a new cache key or the relevant stale cache entry is removed.
11.9 Marking content as not translatable
Use NoTranslate.astro for MDX content that must remain unchanged.
Import it from an MDX file in src/content/posts/ like this:
import NoTranslate from "../../components/NoTranslate.astro";Inline use:
The term <NoTranslate>anti-language</NoTranslate> should stay unchanged.Block use:
<NoTranslate as="div">
This text should not be sent to DeepL.
It will remain exactly as written.
</NoTranslate>Use as="div" for block content. The default element is span, which is better suited for inline text.
Typical uses:
- already-English quotations
- code-like output that is not fenced as code
- generated text that should remain in its original form
- names, labels, or conceptual phrases that should not be normalized by DeepL
- very large fragments that would otherwise exceed the translation size limit
11.10 notranslate inside MediaRow
MediaRow.astro is a special case because type: "text" items are rendered through set:html. That means the content value is an HTML string, not an Astro component.
This does not work:
<MediaRow
items={[
{
type: "text",
content: "<NoTranslate>This will not run as an Astro component.</NoTranslate>"
}
]}
/>Use a normal HTML marker instead:
<MediaRow
bordered
items={[
{
type: "text",
content: "28. prosince 2014 jsem v Mombase vyfotil starou popsanou zeď."
},
{
type: "text",
content: `
<div class="notranslate" translate="no">
It doesn't matter which religion you claim you are.
It doesn't matter which country you're coming from.
It doesn't matter whether you're poor or rich.
It doesn't matter if you're black or white.
We are all the same in the eyes of God.
</div>
`.trim()
},
{
type: "text",
content: "Nezáleží, jakého jsi vyznání. Nezáleží, z jaké země pocházíš."
}
]}
/>The important part is:
<div class="notranslate" translate="no">
...
</div>The postprocess recognizes this marker, removes the protected block before sending the fragment to DeepL, and restores it afterward.
11.11 Large fragments
The current maximum translated fragment size is 80_000 bytes. If a translated region is larger, the build fails with an error similar to:
EN fragment for /en/example/ is 166405 bytes, above configured 80000.Preferred solutions:
- mark non-translatable output with
NoTranslate - use
<div class="notranslate" translate="no">...</div>insideMediaRowtext items - split unusually large pages into smaller translated regions if needed
The size guardrail exists to avoid sending oversized and fragile HTML blobs to DeepL.
11.12 Troubleshooting
DEEPL_AUTH_KEY is not set
The build tried to create a new translation but no DeepL key was available. Set the key and run the translate command again:
set "DEEPL_AUTH_KEY=YOUR_DEEPL_KEY"
npm run build:web:translateMissing EN translation cache
Strict mode found an EN page that needs a translation cache entry that does not exist yet. Run:
npm run build:web:translateor, for USB:
npm run build:usb:translateThen run the strict build again.
Remote DeepL glossary created from local TSV
This is an informational message, not an error. The TSV file in the repository is the local source of glossary entries, while DeepL requires a separate remote glossary object in the account. The message means that no reusable remote object was found through the saved state ID or configured glossary name, so the script created one and stored its ID in:
translations/glossary-cs-en.state.jsonIf this message appears on every translate build, check whether the state file is being deleted, whether a different DeepL account or API key is being used, or whether the remote glossary was removed.
DeepL HTTP 429
DeepL is rate-limiting translation requests. The script spaces requests apart, retries temporary failures with exponential backoff, and respects the Retry-After response header when provided. If all retries still fail, run the same translate command again:
npm run build:web:translateTranslations completed before the failure are already stored in translations/en/. Do not delete the translation cache or glossary state file merely because a 429 occurred; that would discard useful progress and create more requests.
Cache pruning would delete a manually edited translation
First run only the dry-run command:
npm run build:web:prune:dryIf an unused cache entry still needs to be preserved, open the corresponding JSON file and add edited: true, manual: true, or locked: true. Then run the dry-run command again and check that the file is counted as cache-locked-kept instead of cache-would-delete.
Header and UI are English, but the article body is Czech
The manual UI dictionary is working, but the automatic content translation did not run or did not have a cache entry. Check the build log and run build:web:translate.
Local dist/ is Czech after build:web:translate
Check the end of the build log. If it contains DEEPL_AUTH_KEY is not set or another Postprocess failed message, the Astro build completed but the translation postprocess failed.
Production is Czech, but local dist/ is English
The problem is upload or caching, not translation. Upload the entire dist/ directory again and force overwrite existing files. Avoid “skip if same size” and similar FTP shortcuts. Then test with a cache-busting URL parameter.
Metaweb PDF export reports Article content container not found
Check which build command last wrote to dist/. A preceding USB build, including npm run build:usb:translate:signed, leaves flat-file routes and file://-oriented links in that shared directory. The metaweb PDF exporter expects the standard web layout and cannot find the article at its directory route.
Recreate the standard web build and then rerun the export:
npm run build:web:translate:signed
npm run export:pdf:metawebIf the translation cache is complete and strict mode is sufficient, npm run build:web:strict:signed can replace the first command. Do not use npm run generate:metaweb; that package script does not exist.
Gemini was generated after a USB build
Treat that dist-gemini/ output as invalid even if npm run generate:gemini completed successfully. A USB build rewrites relative links for opening HTML directly from disk, but the Gemini converter resolves them as web links. Rebuild the capsule from a fresh standard web build:
npm run build:geminiFor a release that may also need to fill missing English translations and sign the standard web build, use npm run build:web:translate:signed instead; that workflow already generates dist-gemini/. Both commands replace the existing Gemini output, so it does not need to be deleted manually.
11.13 Publishing checklist
Before uploading the web build to FTP:
- Run
npm run build:web:translate. - Run
npm run build:web:strict. - If translation cache cleanup is needed, run
npm run build:web:prune:dryand check the summary. - If the dry-run looks correct, run
npm run build:web:prune. - Run
npm run preview. - Check that
dist/SHA256SUMS.txt,dist/BUILD_SHA256.txt,dist/BUILD_HASH_HISTORY.txt, anddist/integrity.jsonexist. - Open at least one EN article locally.
- Check that the article body is actually English, not only the header and metadata.
- Upload the complete
dist/directory, including the integrity files. - Overwrite existing files on the server.
For USB/offline output:
- Run
npm run build:usb:translate. - Run
npm run build:usb:strict. - If translation cache cleanup is needed, run
npm run build:usb:prune:dryand check the summary. - If the dry-run looks correct, run
npm run build:usb:prune. - Check that
SHA256SUMS.txt,BUILD_SHA256.txt,BUILD_HASH_HISTORY.txt, andintegrity.jsonare present in the generated output. - Open the generated HTML from disk.
- Check CSS, images, internal links, and EN article content.
For Arweave / Permaweb output:
- Run the metadata audit:
python scripts/audit-public-metadata.py --exiftool "D:\Program Files\exiftool\exiftool.exe" - Resolve unintended metadata findings before publishing.
- Run
npm run build:arweave. - Verify that the log ends with
Arweave build prepared in dist-arweave. - Verify that
dist-arweave/SHA256SUMS.txt,dist-arweave/BUILD_SHA256.txt,dist-arweave/BUILD_HASH_HISTORY.txt, anddist-arweave/integrity.jsonexist. - Test the build locally under a fake manifest-like subdirectory.
- Upload only
dist-arweave/to ArDrive. - Create the manifest at the level where
index.htmlis directly present. - Test the manifest gateway URL.
- Check several deep routes, English routes,
ALL_POSTS.txt,ALL_POSTS.json,ARCHIVE.txt, and the integrity files.
12. Deploy
12.1 dist structure
The dist/ directory contains the finished build intended for publishing.
Important generated preservation and verification files include:
dist/404.html
dist/ALL_POSTS.txt
dist/ALL_POSTS.json
dist/SHA256SUMS.txt
dist/BUILD_SHA256.txt
dist/BUILD_HASH_HISTORY.txt
dist/integrity.json
dist/SIGNING_STATUS.txt
dist/SHA256SUMS.txt.asc # only after an explicit signed build
dist/keys/vojta-maur-openpgp.asc
dist/keys/vojta-maur-openpgp-fingerprint.txt
dist/source/vojtamaur-web-source.zipThe source/ directory in dist/ is a generated output directory. It should not be treated as source input or committed as a hand-maintained project directory.
The Gemini capsule is not part of this tree. It is generated separately as the sibling directory dist-gemini/, and it is not included in the checksum manifest or detached signature stored in dist/.
12.2 Deployment of the standard web build
For the production website, the content corresponding to the standard web build is uploaded to the hosting server.
12.2.1 Neocities mirror limitation
The current Neocities mirror is subject to Neocities file-type restrictions, which reject .zip uploads. The following generated file is therefore not present on that mirror:
dist/source/vojtamaur-web-source.zipBecause the source-package link in the published content is root-relative, the link resolves on the Neocities mirror to /source/vojtamaur-web-source.zip. That URL is a known dead link on the Neocities deployment.
The source package remains part of the complete local dist/ output and other deployments that accept ZIP files. The Neocities deployment must therefore be treated as a partial hosting mirror rather than a byte-for-byte copy of dist/. Integrity files generated from the complete build may still list the source ZIP even though Neocities rejected it; this is a known deployment-specific omission, not evidence that the local build is corrupted.
12.2.2 Codeberg Pages deployment
The Codeberg repository is named pages and is available at:
https://codeberg.org/vojta_maur/pagesIt contains two different branches with separate purposes:
main– the normal project source, pushed together with the GitHub and GitLab copiespages– the generated contents ofdist/, used only by Codeberg Pages
The published website is available at:
https://vojta_maur.codeberg.page/The repository name is intentionally pages. The normal web build contains root-relative URLs such as /_astro/..., /images/..., and /slug/. If the same build were served from a repository subpath such as /vojtamaur-web/, CSS, images, the web manifest, and internal links would resolve against the domain root and return 404. Using the special pages repository exposes the build at the user-domain root and preserves the same URL model as the production website and the other root-hosted mirrors.
Codeberg Pages is triggered by a repository webhook with the following relevant settings:
Target URL: https://vojta_maur.codeberg.page/
Event: push
Branch filter: pages
Content type: application/jsonThe local setup uses two Git worktrees so that the source branch and generated deployment branch can remain checked out at the same time:
G:\vojtamaur-web -> main
G:\vojtamaur-pages -> pagesThe deployment helper is:
deploy-codeberg-pages.batIt is intended to be stored in the project root. The script:
- verifies the expected directories, branches, and the
codebergremote - refuses to deploy while the
mainworktree contains uncommitted changes - runs
npm run buildinG:\vojtamaur-web - resets and cleans only the
pagesworktree - copies the finished
dist/contents intoG:\vojtamaur-pages - creates a deployment commit only when the generated output changed
- pushes only the
pagesbranch to Codeberg
The script does not commit, modify, or push main. The normal source publishing workflow therefore remains separate:
git status
git add .
git commit -m "Commit message"
git push origin main
git push gitlab main
git push codeberg main
deploy-codeberg-pages.batAn optional deployment commit message can be passed as an argument:
deploy-codeberg-pages.bat "Deploy updated website"On a fresh workstation, the pages worktree can be recreated after configuring the codeberg remote and fetching the remote deployment branch:
git remote add codeberg https://codeberg.org/vojta_maur/pages.git
git fetch codeberg pages
git worktree add -b pages G:\vojtamaur-pages codeberg/pagesIf a local pages branch already exists, omit -b pages and attach the worktree to that existing branch instead. The deployment branch is generated output and should not be merged into main.
12.2.3 OpenPGP signing policy
Only build artifacts explicitly produced and signed in the author’s controlled local environment are OpenPGP-signed. The private signing key is kept on the author’s local storage and is never committed to the repository, copied into the build, or uploaded as a CI secret.
Automated or provider-side deployments that do not run the explicit local signing step are intentionally unsigned. They may still publish the public key and fingerprint because those are public identity material, but their presence does not authenticate that particular build.
A valid signature belongs only to the exact SHA256SUMS.txt that was signed. A provider-side rebuild is a different artifact even when it was produced from the same source revision. Generated timestamps and deployment-specific rewriting can also make otherwise equivalent builds byte-wise different.
A mirror that receives an already completed locally signed build without rebuilding or modifying it can preserve the same signature. The absence of SHA256SUMS.txt.asc on an automated deployment therefore does not indicate failed checksum generation; it indicates that the explicit private-key operation was not performed.
Deploy the history snapshot already contained in the signed output. The newer project-root BUILD_HASH_HISTORY.txt belongs to future builds and must not be copied over the signed snapshot. Provider-side unsigned rebuilds may inherit existing source-history records, but they do not append a record or authenticate their own output merely by publishing that file.
12.3 .htaccess
For a static Astro website, a minimalist configuration is appropriate.
12.4 Portable build
The portable build can be used as a file-based snapshot or offline copy. However, it is not identical to normal web hosting, and some external services may behave differently.
In the portable build, root-relative paths must be rewritten so that the site works from file:// URLs. This includes the source package link. A correct portable build must link to the source package relatively, for example:
./source/vojtamaur-web-source.ziprather than:
/source/vojtamaur-web-source.zipThe latter would point to the root of the local drive, such as C:\source\vojtamaur-web-source.zip on Windows.
12.5 Arweave / Permaweb deployment
For Arweave deployment, do not upload the normal dist/ directory directly. Use the generated Arweave-specific output:
dist-arweave/Recommended workflow:
- run the public asset metadata audit
- run
npm run build:arweave - verify that
dist-arweave/SHA256SUMS.txt,dist-arweave/BUILD_SHA256.txt,dist-arweave/BUILD_HASH_HISTORY.txt, anddist-arweave/integrity.jsonexist - optionally test
dist-arweave/locally under a fake manifest-like subdirectory - upload only
dist-arweave/to a public ArDrive drive - create an Arweave manifest at the level where
index.htmlis directly present - copy the manifest Data TX ID
- test the site through an Arweave or ArDrive gateway URL
The manifest must be created inside the uploaded build root, not above it. The correct manifest target level contains:
index.html
_astro/
images/
files/
ALL_POSTS.txt
ALL_POSTS.json
ARCHIVE.txt
SHA256SUMS.txt
BUILD_SHA256.txt
BUILD_HASH_HISTORY.txt
integrity.jsonThe uploaded Arweave version is immutable. If a mistake is uploaded, the correction must be published as a new upload and a new manifest transaction. The old version remains available.
12.6 Gemini capsule deployment
For Gemini deployment, use the generated directory:
dist-gemini/The Gemini server document root should correspond to the contents of this directory so that index.gmi is the English capsule homepage. The Czech homepage remains available at:
/cs/index.gmiA typical update consists of:
- generating a fresh capsule with
npm run build:gemini, or generating it through the translating production workflow - previewing representative
.gmifiles locally in a Gemini client - uploading the contents of
dist-gemini/through the hosting account’s supported transfer method, such as SFTP or FTPS - verifying the homepage, language switch, section indexes, several article pages, the YouTube playlist link, direct media links, demo links, and the OpenPGP block
The capsule itself contains text and links. Images, PDFs, videos, maps, 3D models, and interactive HTML demonstrations remain on the normal HTTPS website and are opened through absolute URLs. The Gemini deployment therefore depends on the continued availability of https://vojtamaur.cz/ for non-text media, while article text and capsule navigation remain native Gemtext.
The capsule can be updated independently by rerunning npm run generate:gemini against an existing current dist/. However, publishing from stale HTML would also publish stale translated content, so the complete build command should be preferred for normal releases.
13. Known issues and solutions
13.1 YouTube embed in local or file-based mode
A YouTube iframe may fail in local or file-based mode with error 153. In that case, it is recommended to account for a fallback opening of the video via an external link.
13.2 Sketchfab warnings in the console
The Sketchfab iframe may generate console warnings such as:
- permissions policy violation
- deviceorientation blocked
- accelerometer is not allowed
If the viewer works, this is not a project error, but a limitation or behavior of a third party.
13.3 Broken CSS or assets with the wrong base model
If the build uses root-relative paths in an environment where no standard server root is available, styles, images, and internal links may break.
This affects more than one special output model:
- USB / file-based output
- Arweave / Permaweb output served below a manifest transaction path
For this reason, the portable file-based build is supplemented with scripts/usb-rewrite.mjs, and the Arweave build is supplemented with scripts/make-arweave-build.mjs.
If CSS, JavaScript, images, source package links, or internal links fail in the Arweave deployment, check whether any root-relative path such as /_astro/..., /images/..., /files/..., /source/..., or /slug/ remained in the generated dist-arweave/ output.
13.4 Dev server and new articles
If the listing or routes do not match after adding a new .mdx file, the recommended first step is to restart the development server.
13.5 Lagrange and the Unicode replacement character
A tested Lagrange client stops rendering the remainder of a normal text line immediately after a literal Unicode replacement character U+FFFD (�). The source .gmi remains valid UTF-8, and the same character is allowed by Unicode; this is treated as a client rendering defect rather than malformed Gemtext.
The Gemini generator works around the issue only in the generated capsule by converting each literal U+FFFD to the visible ASCII sequence:
\uFFFDThis preserves the information that the replacement character occurred while preventing the client from truncating the rest of the line. The original MDX, rendered HTML, and other Unicode characters are not modified.
13.6 Lagrange indentation after intentional line breaks
When several normal Gemtext lines originate from intentional <br> line breaks inside one HTML paragraph (the equivalent of Shift+Enter), Lagrange may render the later lines with paragraph-like indentation. Kristall renders the same .gmi content without this indentation. This is a client-specific rendering difference, not an error in the Gemini generator.
Where the source text genuinely contains separate paragraphs, replacing <br> with separate MDX paragraphs improves both HTML and Gemini output. Intentional hard line breaks should remain unchanged if converting them to paragraphs would alter the intended HTML structure.
14. Project backlog and experiments
14.1 Future extensions
Possible future directions for the project:
- translation of image alt texts
- adding alt texts to photographs in the Travel section
- automatic translation of PDF exports, with automatic compression of the translated PDFs to reduce file size
- restructuring of the homepage biographical content by replacing the current formal “About me” text with a shorter, more informal introduction, while providing a separate detailed curriculum vitae as a linked PDF document
- evaluation of custom perforated boards manufactured by JLCPCB as a passive, punch-card-like physical storage medium, encoding binary data as a regular pattern of through-holes, with development of a self-describing format and optical decoding method using backlighting and conventional cameras or scanners
- evaluation of emerging long-term archival media (e.g. NanoFiche) as costs and availability improve
- evaluation of 5D memory crystal storage, once the technology becomes more affordable and openly accessible
- evaluation of DNA-of-Things storage embedded in physical artworks or materials, once the currently prohibitive cost of custom synthesis and production decreases
- evaluation of future opportunities for inclusion in space-based cultural archives (e.g. the Arch Mission Foundation), if suitable public submission pathways become available
- self-hosting a Tor Onion Service mirror (e.g. on a Raspberry Pi), exposing the same content from a single .onion address over multiple protocols, including HTTP/HTTPS, Gemini (gemini://…onion), Spartan (spartan://…onion), and Gopher (gopher://…onion)
- extension of the manual Internet Archive Wayback Machine sitemap exporter (section 15.3.1) to public mirror URLs and scheduled capture runs
- evaluation of embedding ARCHIVE.txt and/or ALL_POSTS.txt data invisibly within each generated HTML page, allowing individual archived page snapshots to carry a redundant copy of the website’s textual archive without affecting the visible page content
- evaluation of embedding ARCHIVE.txt data within robots.txt, taking advantage of the file’s frequent retrieval and preservation by web crawlers to provide an additional redundant copy of the website’s textual archive
- evaluation of FOREVER as an additional long-term digital preservation service, complementing the existing use of Permanent.org and providing further redundancy through an independent archival provider
- evaluation of using LLM agents to perform the archivist’s role, scaling web discovery and preservation by running many parallel instances
- evaluation of Bitcoin Ordinals inscriptions as an immutable, blockchain-based archival medium, embedding selected small digital artifacts directly into the Bitcoin blockchain to provide highly durable, publicly retrievable copies independent of conventional hosting infrastructure
- implementation of standardized RSS and JSON Feed endpoints (e.g. /rss.xml and /feed.json) for machine-readable syndication and independent subscription to newly published content
- implementation of an automated link and file integrity checker for detecting broken internal links, missing assets and invalid fragment references across the website and generated archival outputs
- production of additional ceramic tablets for deposit in Memory of Mankind containing the website’s image archive, with original filenames preserved so that individual images can be matched to their references in ALL_POSTS.txt
- evaluation of producing a self-supporting laser-engraved granite archival copy of ALL_POSTS.txt, using a substantial stone slab (approximately 35 × 35 cm and 30–40 mm thick) with dense microtext across the front surface; dark, fine-grained granite would be preferred for engraving contrast, with final material and text size selected according to the manufacturer’s technical assessment before production
- development of an interactive-fiction edition for the IF Archive, embedding the complete ARCHIVE.txt and ALL_POSTS.txt in a self-contained
.gblorbtext adventure, playable locally or in a browser via Parchment - evaluation of producing a laser-engraved glass archival copy of ALL_POSTS.txt, either as a solid glass cube with dense microtext distributed across all six faces or as a set of smaller glass plates; surface or subsurface engraving would be selected according to technical feasibility
- creation of an additional offline archival copy on M-DISC optical media, containing the most important exports, source files, and offline builds of
vojtamaur.cz - creation of a ZIM edition of
vojtamaur.czfor distribution and browsing through Kiwix, providing an additional standardized, self-contained offline representation of the website alongside the existing HTML-based USB build - donation of physical print editions to institutional collections: an artist’s-book edition of the Metaweb article to the Franklin Furnace Archive, and printed editions of the Metaweb article and ALL_POSTS.txt to the Internet Archive Physical Archive, providing additional long-term preservation within independent cultural and library institutions
- experimental neural archival edition of
ALL_POSTS.txt, training a small Transformer language model from scratch exclusively on the text corpus and deliberately overfitting it to memorize the original content within its neural weights, with the aim of reconstructing a substantial proportion of the original text (ideally 80–90% or more) from the trained model without access to the source file - development of an automated approach to indirectly disseminating the project’s ideas through third-party language model training, building on initial manual experiments involving discussions about
vojtamaur.cz,ALL_POSTS.txt, and the Metaweb article, accompanied by positive response feedback to potentially increase their relevance for training; possible approaches include sustained automated conversations between models from different providers, aiming to marginally influence future models’ learned representations and token probability distributions. Training inclusion and actual influence remain unverified, automated interactions may be filtered, and any implementation must account for providers’ terms of service and the potential classification of such activity as training-data poisoning - evaluation of lunar archival preservation through LifeShip, either by submitting a PDF export of
vojtamaur.czfor storage on a digital microchip (starting at $75 for up to 2 MB) or by commissioning a more durable, passively readable NanoFiche edition; the latter requires individual pricing and verification of archival longevity and future mission availability - evaluation of submitting
vojtamaur.czas an artistic and archival project to a future Moon Gallery Foundation open call, subject to the availability of new submission opportunities and curatorial selection - evaluation of sending a physical NanoFiche archival copy of
vojtamaur.czto the Moon through Astrobotic MoonBox, potentially using a small circular NanoFiche (12–19 mm in diameter) instead of conventional flash storage; feasibility depends on verification of the capsule’s internal dimensions, material compatibility, environmental protection, mission availability, and cost - evaluation of submitting
vojtamaur.czto a future Galactic Library Preserve Humanity (GLPH) lunar archival collection organized by NanoFiche / Stamper Technology in partnership with Astrobotic, as an alternative to purchasing a dedicated MoonBox capsule; this approach could allow the project’s content to be physically recorded on NanoFiche as part of a shared lunar archive, potentially reducing individual manufacturing and transportation costs. The initial GLPH collection for the Griffin-1 mission has reportedly already been prepared and delivered for integration, so participation would depend on future submission opportunities, curatorial acceptance, pricing, supported content volume, and verification of the final archived material
14.2 Failed attempts and abandoned approaches
The following approaches were tested or seriously evaluated but were ultimately abandoned:
- development of a custom OpenAI GPT tentatively called vojtamaurGPT was explored as an interactive interface to the project and its content; the approach was abandoned because the Custom GPT interface at the time was considered too restrictive for the intended development workflow, while relying on a language model to interpret and mediate access to the archived material was considered less desirable than providing direct, deterministic access to the underlying content
- development of a dedicated wiki with interconnected articles (e.g.
wiki.vojtamaur.cz) was considered but ultimately abandoned for both practical and philosophical reasons; the approximately six-hour video-based verbal export ofvojtamaur.czalready covers much of the material that might otherwise have been included in the wiki, while the need for a separate explanatory knowledge base was considered undesirable, as the website itself should remain sufficiently self-contained and comprehensible without an additional layer of interpretation - publication of the searchable PDF export of
vojtamaur.czon sourceAFRICA was considered as an additional public archival and distribution channel; sourceAFRICA was contacted about the possibility, but no response was received, so the attempt was abandoned - use of Perpetual Storage and its Family GoBox service was considered for long-term physical storage in a climate-controlled, highly secured facility in Utah; the company was contacted for further information, but no response was received, so the option was abandoned
- publication of the
Metaweb Articleon Kobo was attempted but rejected because Kobo considered it too niche and insufficiently distinct for publication; rather than resubmitting the same work, the approach was changed and the EPUB export of all articles was later published successfully as an authorial collection categorized primarily as biography/art rather than technology/programming - hosting of a Spartan capsule on tilde.team was explored as an additional alternative-protocol mirror of
vojtamaur.cz; the service was contacted about Spartan hosting, but no response was received, so this option was abandoned - inclusion of the English
Metaweb Articlein ArchiveBox’s Web Archiving Community page was proposed to Nick Sweeting for the “Articles We Like About Internet Archiving” section; no response was received and the article was not added - registration of two physical air-gapped copies listed in the Metaweb Article — ID 013 (
ALL_POSTS.txt Memory of Mankind Edition) and ID 015 (vojtamaur.cz offline 5) — with the International Time Capsule Society was attempted; the first registration was submitted on 2026-09-07 and the second on 2026-09-21; as of 2026-09-27, neither had appeared in the public registry, and a follow-up email sent on 2026-09-18 had received no response - publication of text-based archival exports of
vojtamaur.czon Tumblr was attempted using its official API; smaller exports were successfully published, but the account was subsequently terminated, possibly due to automated spam detection triggered by unusually large posts and repeated API requests; an appeal was submitted to Tumblr Trust & Safety on 2026-09-27, and the approach was suspended pending review - a BitTorrent edition of
vojtamaur.czwas evaluated as a way to reconstruct the complete web build directly from existing public hosting mirrors using HTTP web seeds, without storing an additional ZIP archive; the approach was abandoned because the BEP 19 multi-file web seeding specification requires compatible URL directory structures, while individual hosting providers may serve structurally or byte-wise different builds. Supporting these differences would require provider-specific configuration or changes to the existing deployment workflows. Furthermore, continuously updated mirrors cannot guarantee the availability of immutable historical torrent releases. The resulting implementation and maintenance overhead was considered disproportionate to the expected archival benefit. Conventional P2P torrent distribution remains technically possible but was not pursued as part of this approach. - inclusion of
vojtamaur.czin the Lunar Codex was considered as a potential form of lunar archival preservation; however, the project closed submissions in August 2025 and has stated that it does not plan to reopen them. The approach was therefore abandoned indefinitely without contacting the organization, although its submission status may be reassessed in the future
15. Archival operations
15.1 Preservation priority under constraints
When storage capacity or snapshot quotas make it impossible to preserve the complete site, archive these artifacts in the following order:
ALL_POSTS.txt– the smallest practical content-preservation layerARCHIVE.txt– the map of repositories, mirrors, snapshots, and archive entry points- this technical documentation (
/documentation/or its source file,src/pages/documentation.mdx)
This priority is especially useful for storage-constrained deposits such as Memory of Mankind and for services such as Perma.cc where the number of available snapshots may be limited.
15.2 Manual mirror updates
Some mirrors do not support automatic deployment from the canonical repository and therefore need to be updated manually. These currently include ChatGPT Sites, Neocities, ArDrive / Arweave, Pollux.casa, and envs.net.
The ChatGPT Sites deployment does not update automatically when the canonical repository changes. To refresh it, ask ChatGPT in the Sites chat to load the current state of the canonical repository, save or build a new version, and deploy that version explicitly.
The Vercel mirror requires this root-level vercel.json:
{
"buildCommand": "npm run build:web",
"outputDirectory": "dist",
"framework": "astro"
}Without this override, Vercel’s default Astro build may skip the post-build step that generates ALL_POSTS.txt and ALL_POSTS.json.
Codeberg Pages is also updated separately through deploy-codeberg-pages.bat; the complete worktree-based procedure is documented in section 12.2.2.
15.3 External web archives
As an operating assumption, newly published article URLs and mirror URLs should be submitted to the Internet Archive. Seeding the archive with these URLs is expected to improve discovery and may lead to later automatic recrawling, but it is not a guarantee that every URL or later version will be captured.
The Czech Webarchiv likely captured https://vojtamaur.cz/ after a manual preservation request. Searching for the root URL currently returns the following access notice:
Tuto stránku Webarchiv nemůže zobrazit. Z důvodu autorského zákona nemůžeme tuto stránku zpřístupnit online. Archivované verze této stránky jsou dostupné pouze z Referenčního centra NK ČR.
This indicates restricted archived holdings rather than public online access; it does not by itself establish the completeness or capture dates of those holdings.
15.3.1 Manual Wayback snapshots from the live sitemap
The manual exporter scripts/export-wayback-snapshots.py submits selected live sitemap URLs to Internet Archive Wayback Machine Save Page Now (SPN2) and exports confirmed snapshot links. It requires Python 3.10 or newer and uses only the standard library. It runs independently of export-all.bat and all build commands.
Sitemap content and build order
Web builds retain Astro’s sitemap-index.xml and existing child sitemap names. After translation postprocessing, ALL_POSTS generation and source-bundle generation, npm run generate:sitemap enriches the child sitemaps from the rendered HTML. Existing local images are attached to their pages through image:image / image:loc, including thumbnails, full-size links, social images and srcset variants. Missing files and external image hosts are excluded; unreferenced images are not assigned to arbitrary pages. Images are deduplicated per page, with a maximum of 1,000 per page. A sitemap exceeding the URL or byte limit fails the build instead of silently dropping entries.
Public PDFs and relevant TXT files are ordinary loc entries: PDFs in the build, TXT files under files/, linked nontechnical TXT files, and the root exports ALL_POSTS.txt, ARCHIVE.txt, PRESERVATION_INSTRUCTIONS.txt and llms.txt. Technical checksum/status files and robots.txt are excluded. MEDIA_MANIFEST.json and MEDIA_SHA256SUMS.txt exist inside the source ZIP, not as standalone public URLs, so the sitemap does not invent links to them.
Integrity generation follows sitemap enrichment; signed builds sign that final manifest afterwards. USB builds continue to omit sitemaps. Arweave builds inherit the web sitemap before their own rewrite and integrity steps. The GitHub Pages workflow refreshes ALL_POSTS.json and integrity after rewriting mirror paths; the sitemap retains canonical https://vojtamaur.cz URLs.
Choose pages only or pages with files
Run either command from any working directory:
python "F:\vojtamaur-web\scripts\export-wayback-snapshots.py" --pages-only
python "F:\vojtamaur-web\scripts\export-wayback-snapshots.py" --with-files--pages-only is the default when no selection option is supplied. It submits page URLs, including all languages and /ns/.../, without separately submitting images or PDF/TXT files. Wayback may still capture a page’s embedded resources while capturing the page itself. --with-files selects pages plus all images, PDFs and TXT files discovered in the sitemap. It processes pages first, then images, PDFs and TXT files, so a large image collection does not delay the page batch.
For a custom selection, use --types images pdf txt or --types pages pdf. --pages-only, --with-files and --types are mutually exclusive. “With files” means the supported files in the sitemap, not a filesystem upload or discovery of arbitrary files outside it. The checked build on 2026-09-28 contained 199 page URLs, 822 unique image URLs, 7 PDFs and 5 TXT files; these are observations, not fixed limits.
The exporter follows nested sitemap indices and gzip-compressed sitemaps. URLs must be valid public HTTP(S) addresses on the selected sitemap’s origin; credentials, local addresses, invalid paths and cross-origin sitemap redirects are rejected. Fragments and equivalent host/default-port/percent-escape spellings are normalized for deduplication. It does not discover mirror URLs or pages outside the sitemap.
Internet Archive credentials
Obtain an access key and secret key from Internet Archive S3 API keys. In Windows Command Prompt, set the placeholders to your actual keys in the same window before running the exporter:
set "IA_ACCESS_KEY_ID=YOUR_ACCESS_KEY"
set "IA_SECRET_ACCESS_KEY=YOUR_SECRET_KEY"
python "F:\vojtamaur-web\scripts\export-wayback-snapshots.py" --pages-onlyIn PowerShell:
$env:IA_ACCESS_KEY_ID = "YOUR_ACCESS_KEY"
$env:IA_SECRET_ACCESS_KEY = "YOUR_SECRET_KEY"
python "F:\vojtamaur-web\scripts\export-wayback-snapshots.py" --with-filesThese settings last for the current shell session. The script does not read .env, load ia configure credentials, or ask for an account password. Keep keys out of source files, documentation and committed batch files. Credentials are sent only to the fixed SPN2 endpoint; authenticated redirects are disabled.
Preview without archiving
Live sitemap previews need no keys and create no files or captures:
python "F:\vojtamaur-web\scripts\export-wayback-snapshots.py" --pages-only --dry-run
python "F:\vojtamaur-web\scripts\export-wayback-snapshots.py" --with-files --dry-runA local preview performs no network requests at all:
python "F:\vojtamaur-web\scripts\export-wayback-snapshots.py" --with-files --dry-run --dist "F:\vojtamaur-web\dist"--dist requires --dry-run; actual captures use the published sitemap. --sitemap URL selects another sitemap and defines the permitted origin. --limit N limits the number of selected unique URLs after ordering, including any locally reusable results. Dry-run lists the selected URLs regardless of recent captures; it does not resume or write an export.
Restarting and recent captures
By default, the exporter reuses confirmed captures from the last 24 hours. It reads previous exports/vojtamaur-web-wayback-*.txt files, validates their snapshot links and capture timestamps, and reuses the newest matching original URL without submitting it again. Those confirmed links are included in the new output and marked as reused in the console. Old, future-dated, malformed and invalid links are ignored. Previous export files are never modified.
The same interval is sent to SPN2 through if_not_archived_within, allowing the service to return a recent capture too. Thus, after interruption or rate limiting, run the same command again to reuse recent successes and work on the remaining URLs. Local matching uses the final original URL in the stored snapshot; a redirect from a different requested URL may require a request to SPN2 again. Pending jobs, failed captures and ambiguous submissions are never assumed successful.
--if-not-archived-within 86400 explicitly selects 24 hours. A different nonnegative number changes the interval; --if-not-archived-within 0 disables local reuse and requests a fresh capture from SPN2. It cannot override Internet Archive’s resource-specific daily limits. See the Internet Archive SPN2 client options.
Output and failures
Each capture run creates a timestamped UTF-8 TXT under F:\vojtamaur-web\exports\. It contains only confirmed snapshot URLs, one per line, including reused recent results. Each link uses the actual capture timestamp and final original_url; accepting a job alone is not success. An unexpected favicon.ico result is omitted without resubmission. Other legitimate redirects remain supported.
Successful lines are flushed and synced immediately. A sibling .failures.jsonl file records each failed target URL and its error, including a failure that stops the batch. An empty failure file means no individual failures were recorded; also check the final sitemap-error count and exit status. A disk-write failure stops processing. The console reports selection counts, jobs, waits, local reuse and the final totals. Files are created only by actual capture runs, never by dry-run.
After review, manually copy desired snapshot links into public/ARCHIVE.txt; the exporter never edits it. Export files stay outside dist/ and are ignored by Git.
Waiting, rate limits and interruption
Submissions are sequential, with a default 15-second pause between targets that need a request (--delay, minimum 5 seconds). Locally reused results need no pause. Accepted jobs are polled every five seconds with a default 900-second deadline (--poll-timeout). HTTP requests have a maximum 60-second timeout, shortened by the remaining polling deadline.
An HTTP 429 rejects a submission. The exporter waits at least 300 seconds, or longer if Retry-After requests it, and retries that rejected request up to six times. If a 429 unexpectedly includes a job ID or confirmed success, that result is handled instead of resubmitting. Explicit error:too-many-requests and error:user-session-limit admission errors without a job ID also wait and retry. Exhausted rate-limit retries stop the batch while preserving completed results. Authentication/access failures (401/403) stop immediately.
Once a job has been accepted, it is never submitted again in the same run. Lost POST replies, POST 5xx responses, job errors, polling timeouts and unexpected favicon results are recorded as failures without blind resubmission. Safe GET requests may retry network/429/5xx failures with backoff and Retry-After, within the polling deadline where applicable.
error:too-many-daily-captures is a limit for that particular URL. It is deferred to a later run while other targets continue; the script cannot bypass it. Errors such as error:gateway-timeout, error:no-captures or nocaptures mean Wayback could not complete the capture. The script records them and continues, but cannot guarantee that Internet Archive can reach the target. Rerunning later reuses recent successes and tries unfinished targets again.
By default, HTTP error pages are not requested as captures. Use --capture-errors only when you also want SPN2 to capture HTTP error responses, matching the older exporter’s capture_all behavior. This option does not fix rate limits or connectivity errors.
Ctrl+C preserves completed lines. An unconfirmed job may still finish at Internet Archive after interruption, which is why it is not blindly resubmitted within the same run. Exit codes: 0 = complete selected batch, 1 = failure or partial results, 130 = interrupted.
15.4 YouTube video archiving
YouTube videos are downloaded from Windows Command Prompt with yt-dlp. Put the video URLs in urls.txt, one URL per line, then run:
yt-dlp --write-info-json --write-thumbnail --write-description --write-subs --write-auto-subs --sub-langs "cs,en" --embed-metadata --download-archive downloaded.txt -f "bv*+ba/b" -o "%(title).200B [%(id)s]/%(title).200B [%(id)s].%(ext)s" -a urls.txtThe downloaded.txt archive prevents already recorded videos from being downloaded again.
15.5 Arweave gateway entry points
An Arweave manifest Data TX ID identifies the archived deployment independently of the HTTP gateway used to retrieve it. The current archived deployment uses this manifest Data TX ID:
GHwSYFJtzrRqt5AOZuwJfJHoV6Bc5qjwe5FtFP6UjAsARCHIVE.txt and other preservation records should list at least these two entry points to the same manifest:
https://db6beycsnxhli2vxsahgn3ajpsi6qv5alttkr4d3sfwrj7uurqfq.ardrive.net/GHwSYFJtzrRqt5AOZuwJfJHoV6Bc5qjwe5FtFP6UjAs/
https://arweave.net/GHwSYFJtzrRqt5AOZuwJfJHoV6Bc5qjwe5FtFP6UjAs/These URLs are not separate archived copies. They are two HTTP gateway routes to the same manifest and the same underlying Arweave data. If one gateway returns intermittent 404 or 504 responses, or loads HTML while failing to load CSS or other assets, test the identical manifest path through the other gateway before concluding that archived data is missing.
On 2026-08-18, the ArDrive gateway was observed returning alternating 404 and 200 responses for the same immutable CSS path, while the corresponding arweave.net path returned 200 consistently. The failing ArDrive responses reported X-Cache-Status: UPDATING. This behavior indicates a gateway cache or retrieval failure rather than a change to the permanent stored data.
The manifest Data TX ID is the authoritative archival identifier. Gateway URLs are replaceable access routes and additional gateway entry points may be recorded without uploading the deployment again.
15.6 Arctic World Archive pricing note
For Shared piqlFilm, the observed minimum price is €139. An approximately 0.8 GB package cost €139; 1.19 GB was quoted at €165.48, close to €139/GB; and 9.22 GB was quoted at €1,113.21, which may indicate a different calculation or volume discount at larger sizes. Because the exact public pricing formula is not known, treat these figures as practical observations rather than guaranteed price tiers. To stay at the €139 minimum in practice, keep the package safely below 1 GB.
15.7 Recurring preservation and maintenance
The following archival operations should be performed periodically to maintain the accessibility and preservation of the project. The intervals are approximate and may be adjusted according to the extent of changes.
- update
ARCHIVE.txtwhenever new archival locations, permanent identifiers, or significant snapshots are created - approximately once per year, publish an updated archival release of
vojtamaur.czon Zenodo, preserving previous versions - approximately once per year, upload an updated website build to Arweave / ArDrive as a new immutable archival snapshot
- periodically run
scripts/export-wayback-snapshots.pyto submit the current website to the Internet Archive Wayback Machine, particularly after substantial content updates - periodically update manually maintained mirrors that do not support automatic deployment
- approximately once per year, review the availability of existing archival copies, repositories, and mirrors, identifying inaccessible or discontinued services
- periodically verify the integrity and recoverability of stored archival copies using available checksums and signatures
- approximately once per decade, deposit an updated archival edition of
vojtamaur.czin the Memory of Mankind (MoM) archive, preserving previous deposits - approximately once per decade, deposit an updated archival edition of
vojtamaur.czin the Arctic World Archive (AWA), preserving previous deposits - approximately once every three years, update the archival copies of the project’s YouTube videos, including newly published content, on the existing archival platforms
- periodically archive
ALL_POSTS.txt,ARCHIVE.txt, the documentation page, and the Metaweb article using independent web snapshot services (e.g. archive.today, Ghost Archive, Arquivo.pt, Web Gyotaku, …), particularly after significant updates - approximately once every one to three years, print an updated archival edition of
vojtamaur.czand deposit it in the author’s private physical archive, preserving previous editions - approximately once every three to five years, test the reconstruction of
vojtamaur.czfrom an independent archival copy, verifying that the website can be restored without relying on the original development environment - approximately once every six months, submit the public source code repository of
vojtamaur.czto Software Heritage using its Save Code Now feature, particularly after significant updates - periodically regenerate the Rosetta multilingual export using the current Czech content of ALL_POSTS.json, particularly after substantial content updates, and update its published copies while preserving previous archival snapshots
16. Summary
The project is designed as a file-oriented static website. Content is versioned directly in the repository, and the final published form is produced by the build process. This model makes it easier to archive, restore, and migrate the project without relying on a database runtime.
From a maintenance perspective, the following points are especially important:
- content is managed as
.mdxfiles - metadata is validated through content collections
- specialized content is encapsulated in several reusable components
- the project supports a standard web build, a portable file-based build, a derived Arweave / Permaweb deployment build, and a generated bilingual Gemini capsule in
dist-gemini/ - the Gemini capsule uses English at the root, Czech under
/cs/, native Gemtext navigation, and HTTPS links for binary media and interactive content build:web:translate:signedgenerates the Gemini capsule but signs only the normaldist/checksum manifest;dist-gemini/is currently a separate unsigned artifact- Codeberg Pages is deployed separately from the generated
pagesbranch through thedeploy-codeberg-pages.batworktree workflow - web and USB builds generate a recursively reconstructable source package at
/source/vojtamaur-web-source.zip - when working with new content, it is useful to expect the occasional need to restart the development server
- public assets in
public/images/andpublic/files/should be checked for unintended embedded metadata before commit - unused English translation cache entries can be pruned with a dry-run-first workflow, while
manual,locked, andeditedentries are preserved
This documentation describes the current architecture and operating model of the project in a form suitable for ongoing maintenance, handoff, or future migration.