Warré Archive
An interest in keeping bees in a way that respects their needs led me to start with Warré hives and join the international warrebeekeeping group. I still had a great deal to learn, but I also wanted to contribute. I had little to offer on the finer points of beekeeping, whereas software development was something I knew. That gave me the idea of making the members’ accumulated experience accessible again through an archive.
A community’s experience
The warrebeekeeping group has been discussing Warré hives since 2007. Its conversations cover practical work with bees, observations, difficulties and different approaches. Over the years, more than 33,000 messages have accumulated, along with photo albums and attachments.
A question often drew several answers based on different experiences. The archive therefore preserves complete discussions, including follow-up questions and disagreements.
When Yahoo discontinued its groups service, this collection was at risk of being lost. Access to the old content ended in late 2019, and the service closed entirely in 2020. Without a separate backup, years of discussion would no longer have been accessible to the group. The material we preserved forms the basis of today’s archive.
The group has continued on Google Groups since October 2020. The archive brings the saved Yahoo posts together with the newer discussions, keeping the shared history accessible across the change of platform.
From the first archive to a rebuild
Around 2020, I had already built a first website for the archive, programming it myself through painstaking manual work. Agentic software development has since given me much greater scope. I could return to the archive and implement features well beyond the way the original site displayed the posts.
I took that opportunity to rebuild the application while retaining the existing data. The work now extends from processing individual messages to multilingual search and thematic digests with traceable sources.
Making the messages readable again
Much of a mailing list consists of repeated text. Replies contain earlier messages, often with further quotations and signatures nested inside them. If a search treated all of this alike, it would find the same statement repeatedly in other people’s replies.
The processing therefore separates each author’s own contribution from quoted text and signatures. Original contributions account for only about 41 per cent of the raw material. A reply written below a quotation must not disappear in the process, so the original messages are retained and the cleaned-up presentation can be checked against them.
Matching names and pictures took further work. Some members wrote under different names, while photo albums often lacked a reliable link to the discussions. I left those gaps open rather than guess at a connection.
Finding knowledge across languages
The archive combines conventional full-text search with semantic search. One finds specific terms; the other finds passages related in meaning. A question in German can therefore lead to a relevant English discussion whose author used quite different words. Both searches feed into a single result list.
The interface is available in English, German and French. Translations help with reading, while the original English text remains accessible. A dedicated beekeeping glossary keeps technical terms consistent, including the names of hive components across different posts.
Thematic digests in preparation
As a next step, I am preparing thematic digests, which have not yet been published. They bring together experience from several discussions and retain differences of opinion. The underlying posts remain traceable as sources. An account of someone’s experience should remain recognisable as such; summarising it does not turn it into generally applicable evidence.
Before publishing the digests, I want to read through them all myself and discuss with the members whether, and in what form, they could be made public. Anonymisation alone does not answer the question of what the community wants to share beyond the group. Original posts, photographs and attachments remain restricted to members.
Technical implementation
I built the archive with PHP, a relational database and a small amount of JavaScript. Pages are rendered on the server, without a frontend framework or build step. The search vectors also live in the existing database: a fast initial comparison narrows down the candidates before a more precise assessment.
AI is used for search, translation and content processing, while the original data is kept separate. Proposed corrections are reviewed before being applied, and changes must be traceable to the source text. It must still be possible to read what the original message said after processing.
Glimpses.
-
Preview of the proposed public digests -
Draft digest on choosing a Warré hive