PyPotteryLens
v0.3.2Mines published monographs and excavation report PDFs, detecting and segmenting pottery drawings with Computer Vision.
YOLO Detection · Interactive Review · LLM Metadata · Cards Export
Manually inking vessel drawings, building catalogue entries by hand from old publications, retyping the notes off a stack of inventory cards, working out reduction scales by hand: this is where most of the time in ceramic documentation actually goes, and it's a large part of why so much excavated material sits in a drawer for years before it ever gets published.
PyPottery is a suite of five open-source tools that takes on each of these tasks computationally: extracting figures from monograph PDFs, digitizing field drawings, turning pencil sketches into publication-ready ink, tracing vessel profiles into clean vector curves, and assembling typological plates in seconds.
Everything runs on your own machine (You can use remote models to simplify certain tasks). The full source code is public, so anyone can check exactly how it works - and improve it.
Five modules in three phases, taking raw field data and legacy monographs to publication-grade plates and structured data. Use the whole chain, or only the phase you need.
Automates the tedious hours spent screenshotting monograph tables or hand-cropping drawing plates and retyping handwritten field notes, cutting transcription time by roughly 90%.
.xlsx or .csv), structured catalog records
Mines published monographs and excavation report PDFs, detecting and segmenting pottery drawings with Computer Vision.
YOLO Detection · Interactive Review · LLM Metadata · Cards Export
Digitizes handwritten drawing tables on index cards: segments multi-sherd plates, reads context notes via OCR, and outputs relational Excel tables.
Handwriting OCR · Auto Vessel Cropping · Relational Excel
Replaces the manual ordeal of tracing paper and point-by-point Bézier retracing in vector software.
Transforms pencil sketches into publication-grade ink drawings in seconds using a dedicated diffusion architecture.
Diffusion · Diagnostic Line Weight · Batch Processing · CUDA / MPS / CPU
Interactive profile vectorization powered by SAM 2: click on drawing elements to generate editable, smoothed Bézier curves.
SAM 2 Interactive · Bézier Spline Editor · Layered SVG Export
Eliminates hand-recalculated reduction scales, hand-drawn millimeter bars, and tedious manual alignment of figures on the page.
Composes multi-vessel publication plates with automatic metric scaling, calibrated scale bars, typological grouping, and editable vector PDF/SVG export.
Automatic Layout · Calibrated Metric Scale Bars · Automated Metadata Placement & Numbers · Vector PDF & SVG
PyPottery is designed as an entirely modular suite rather than an all-or-nothing pipeline: you choose exactly how to use it according to your workflow, research needs, and personal documentation preferences:
For environment setup on Windows, macOS, and Linux, see the Suite Installation Guide.
The idea for PyPottery grew out of a set of needs the author first ran into during university: let’s be honest, documenting pottery isn’t boring work, but it is work that takes a great deal of time. That “lost” time takes away from the other stages of research like typological analysis, comparison with other contexts, writing up results. PyPottery exists to drastically cut the time spent on repetitive, mechanical tasks, leaving more room for interpretation and research.
Using PyPottery has been shown to cut documentation time by roughly 60 times compared to the traditional routine of tracing paper, technical pens, and manually re-tracing everything in vector software.
That doesn’t mean it gets everything right on the first pass: a transcribed word sometimes needs a correction, a traced profile sometimes needs a nudge. PyPottery is built around a concept called Human-in-the-Loop: every stage gives you the chance to check and correct before moving on, so what comes out the other end is something you stand behind, not something you’re asked to trust blindly.
None of this replaces the archaeologist’s eye — the interpretive drawing itself, deciding what a fragment tells you, still starts on paper. What changes is everything that happens after that: the hours of repetitive work standing between a first pencil sketch and a page ready to send to a publisher.
The routine
If you’re starting from a stack of hand-written excavation cards, the routine is usually: scan each card, open Photoshop or a similar program and crop out every drawing by hand, then read the context, stratigraphic unit, and inventory number off the card — often in decades-old handwriting — and type them into a spreadsheet, one row at a time. If you’re starting from a published monograph instead, it typically means taking screenshots of PDF pages one by one and re-cataloguing the plates yourself.
With PyPottery
PyPotteryScan and PyPotteryLens take over this stage. Scan reads the drawings sheets and fills in the spreadsheet for you, so you’re checking and correcting rather than transcribing from scratch. Lens goes straight through a PDF and pulls out every figure automatically, no screenshotting involved.
The routine
This is traditionally where most of the hours disappear. A pencil drawing gets taped over with tracing film and gone over by hand with a technical pen to produce clean, publishable ink lines — and if you also need an editable digital version, it gets traced a second time, point by point, in Illustrator, Inkscape or similar software.
With PyPottery
PyPotteryInk takes the pencil scan and produces the finished ink version directly. PyPotteryTrace then turns that into an editable vector drawing, with the section, building-lines, decoration kept as separate elements you can still adjust — not one flat outline.
The routine
Finally, individual drawings have to become a plate: arranging images on the page, working out the correct reduction scale for each one by hand, drawing scale bars, and typing captions and catalogue numbers.
With PyPottery
PyPotteryLayout arranges the plate automatically — consistent scale, positioning, and captions across the whole page — and still hands you back a file you can open and fine-tune in your usual software afterward.
Automation stops where judgment starts
PyPottery automates the mechanical parts — cropping, transcribing, inking, laying out — never the archaeological judgment. Every output is something to check and confirm, not something asked to be trusted blindly.
A toolkit, not a rigid pipeline
PyPottery adapts to your workflow, not the other way around. Don't need profile vectorization? Skip Trace. Prefer manual plate layout to aid typological classification? Skip Layout. Automate only what you need and retain full research control.
Your data stays on your machine
Unpublished excavation data, stratigraphic coordinates, drawings, notes: all of it stays on your computer, with no account and no telemetry. The one exception is metadata extraction in PyPotteryLens, which currently calls a cloud LLM — the best option available today for that specific task.
Open source, not a black box
No license fees, no proprietary format locking your work in. The full source code is public, so anyone can check exactly how it works — and improve it.
Built for consumer hardware
No research-grade GPU required. PyPottery is developed and tested on ordinary laptops and desktops, with a CPU-only fallback for every module. Just a gaming laptop or a MacBook is enough to run the whole suite.
If you use PyPottery tools in academic publications, excavation reports, or museum catalogues, please cite the peer-reviewed article of each module you used, and the software itself.
The repository ships a CITATION.cff file, so GitHub’s Cite this repository button on the PyPottery page gives you ready-made APA and BibTeX entries, and reference managers such as Zotero can import it directly. The same entry as BibTeX:
@software{cardarelli2025pypottery,
title = {{PyPottery Suite: Digitizing Archaeological Pottery Documentation}},
author = {Cardarelli, Lorenzo},
year = {2025},
url = {https://github.com/lrncrd/PyPottery}
}PyPottery is developed as an independent open-source project for the archaeological community. Source code, pretrained models, and issue tracking are on GitHub.