Tools for Crawling and Archiving Your Own Website
A website can change quietly: an old article disappears, a navigation path breaks, an image is replaced, or a domain is transferred to a new host. Regular crawling and archiving create a dependable record of what was published, how pages connected, and whether visitors could still access important material at a particular point in time.
The right approach depends on the site’s size and purpose. A small community resource may need a monthly mirror and a list of preserved URLs, while a large Australian organisation may require scheduled scans, storage controls, privacy reviews, and evidence that archived files have not been altered. The best tools support both discovery and long-term stewardship.
A useful preservation workflow usually has four stages: crawl the live site, save the resources, inspect the captured copy, and document the process. Keeping these stages separate makes it easier to identify missing pages, resolve redirects, remove sensitive material where necessary, and publish a clear route through historical content.
| Tool | Best use | Strengths | Points to watch |
|---|---|---|---|
| Wget | Small static websites | Free, scriptable, easy to schedule | Limited support for complex JavaScript |
| HTTrack | A straightforward offline mirror | Friendly interface and broad format support | Dynamic features may not replay correctly |
| SiteOne Crawler | Technical audits and link discovery | Detailed reports, modern crawling features | Primarily an audit tool rather than a complete archive |
| Screaming Frog SEO Spider | Structured site inventories | Excellent metadata, status and redirect analysis | Free version has a crawl limit |
| Browsertrix or Webrecorder | Interactive and JavaScript-heavy sites | Captures browser behaviour and replayable sessions | Requires more setup and storage planning |
Choosing A Crawler For Your Site
Wget is often enough for a brochure site, historical archive, blog, or simple resource hub built from HTML, CSS, images, documents, and a modest amount of JavaScript. Its command-line options allow an administrator to limit the crawl to one domain, follow internal links, preserve directory structures, and create a repeatable job. That repeatability matters when a site is hosted on a small Australian provider and backups need to run outside business hours.
HTTrack offers a more approachable alternative for people who would rather configure a project through a desktop interface. It can create an offline copy that is convenient for checking page relationships and browsing older material without an internet connection. Neither tool should be treated as a perfect representation of a web application: login areas, search results, shopping baskets, embedded video, maps, and content loaded through APIs often need separate treatment.
For a technical inventory, SiteOne Crawler and Screaming Frog are valuable companions to a preservation tool. They can identify broken links, redirect chains, duplicate titles, missing descriptions, oversized images, orphan pages, and server errors. Their reports help establish what should be archived before a mirror is created. A crawler report is evidence about the live site; it is not, by itself, a preserved copy.
Capturing Dynamic And Media-Rich Content
Modern websites frequently assemble pages in the browser. A basic downloader may save the initial HTML while missing menus, articles, comments, captions, or files that appear only after JavaScript runs. Browsertrix and Webrecorder are designed for this environment. They use browser-based capture and produce WARC files, a widely used container for web archive data that can be replayed through compatible systems.
Interactive capture should be planned rather than left to chance. Record the important paths through the site, including a search page, category navigation, a representative article, a downloadable PDF, and any media player that forms part of the public record. Save screenshots or screen recordings for elements that cannot be replayed reliably, such as live feeds, third-party widgets, or content protected by expiring tokens.
Audio and video deserve extra checks because a downloaded file may exist while its player, transcript, subtitles, or surrounding explanation has vanished. File names, MIME types, duration, dimensions, and checksums provide useful technical evidence. The same careful verification mindset used in broadcast workflows can be seen in real-time loudness checks, where a recording is assessed against defined requirements instead of being assumed correct because it was captured.
Building A Reliable Archive Workflow
Begin with a crawl policy. Define the domains and subdomains that belong in scope, the URL patterns to exclude, the maximum crawl depth, and the file types worth retaining. Decide how the crawler should handle query strings, calendar pages, print views, tracking parameters, and duplicate URLs. A written policy prevents one archive from following an endless chain of filters while another misses key sections.
Run a discovery crawl before the final capture. Review the sitemap, server logs, analytics landing pages, editorial lists, and known external references. This can reveal pages that are not linked from the current navigation. For a site with historical community reporting, an archive such as the news collection can act as a useful starting point for checking whether older stories, dates, and category paths remain discoverable.
Store the resulting files in at least two locations, preferably with one copy separated from the main hosting account. A local drive in Melbourne and an encrypted cloud copy in an Australian region can offer practical resilience, although location does not replace a tested backup. Generate SHA-256 checksums for important files, keep a manifest of captured URLs, and record the crawler version, date, time zone, scope, exclusions, and any manual actions.
Preservation also requires a review layer. Open a selection of pages from the captured copy, compare them with the live site, test internal links, inspect images, and confirm that documents can be downloaded. Keep the original capture unchanged, then create a working copy for repairs or redactions. Clear naming conventions make future migration easier: include the domain, capture date, collection name, and version rather than relying on generic folders such as “backup-final”.
Managing Privacy, Copyright And Access
Archiving a website does not remove legal responsibilities. In Australia, personal information may be subject to the Privacy Act 1988 and the Australian Privacy Principles, especially where a site records names, email addresses, photographs, community submissions, or contact details. Before publishing a public mirror, review forms, user profiles, private correspondence, location information, and old staff pages. A private preservation copy can have a different access policy from a public historical collection.
Copyright also needs attention under the Copyright Act 1968. The person who operates a site may not own every photograph, article, logo, video, map, or embedded item displayed there. Check licences and permissions, retain attribution, and document whether a capture is for internal continuity, research, public access, or a transfer to a cultural institution. Australian organisations should also consider whether records contain material covered by contracts, confidentiality clauses, or sector-specific obligations.
Respecting crawl controls is part of responsible practice. Read robots.txt, follow published terms where appropriate, keep request rates modest, and avoid placing unnecessary load on a small host. A site receiving visitors over an NBN connection in regional Queensland may have different capacity from a large data-centre deployment in Sydney. Schedule intensive scans overnight, use caching where possible, and contact the owner of an external service before capturing a high-volume resource.
Access controls should be documented alongside the archive. Mark collections as public, restricted, embargoed, or preservation-only. If a page contains a person’s phone number or an outdated address, record the reason for limiting access rather than silently changing the historical file. This approach supports transparency while reducing the chance that an old publication creates a current privacy or safety problem.
Practical Recommendations For Long-Term Stewardship
A sustainable programme is easier to maintain when each capture answers a clear question: what existed, when was it available, and can another person verify the record? Small sites should avoid collecting data merely because a tool makes it possible. A focused archive with accurate metadata is more useful than a large, unexamined folder of duplicates.
Use the following practices when establishing a recurring crawl and preservation routine:
- Schedule monthly or quarterly captures according to how often the site changes, with an extra run before a redesign, domain transfer, or hosting migration.
- Keep a URL inventory that includes canonical addresses, redirects, important files, page titles, publication dates, and notes about missing material.
- Combine a static mirror with browser-based capture when JavaScript, embedded media, forms, or interactive navigation are important to the visitor experience.
- Save WARC files, downloaded assets, checksums, crawl logs, screenshots, and documentation in a consistent folder structure.
- Test the archive on another computer and network, including a mobile device, because Australian visitors commonly browse community resources on phones.
- Review privacy, copyright, and access settings before every public release, particularly when older contact details or user contributions are included.
- Maintain two independent copies and perform a restoration test at least once a year rather than assuming that a successful backup job proves recoverability.
Long-term stewardship also benefits from a human-readable landing page. Explain the collection dates, source domain, preservation method, known gaps, and how visitors can report a damaged file or request information. This context helps future editors distinguish an original historical page from a later explanatory note and makes the archive easier to navigate for researchers, former residents, journalists, and community groups.
Preserving A Website Beyond The First Capture
A crawl is a snapshot, not a permanent guarantee. Domains expire, hosting accounts are closed, file formats become harder to open, and external services remove content without notice. Revisit the archive at planned intervals and compare new captures with earlier manifests. Changes in page count, file size, redirects, and status codes can reveal that an important section has moved or disappeared.
Keep the domain and archive responsibilities clear. Domain registration details, renewal dates, DNS records, hosting credentials, software licences, and backup locations should be held in a secure handover record rather than inside one person’s email account. For an Australian community organisation, this can be especially important when volunteers change, committees rotate, or a project moves between cities such as Brisbane, Adelaide, Perth, and Hobart.
When a website is retired, publish a stable notice explaining what happened and where preserved material can be found. Maintain redirects for high-value URLs when practical, but do not redirect every historical address to a generic homepage; that removes useful context and makes verification harder. A dated archive index, clear page titles, and consistent links give future visitors a better path through the material.
Start with a defined collection, run a small test crawl, and inspect the result before committing to a larger capture. Then place the files, checksums, legal notes, and access rules under reliable stewardship so the record remains usable after the original website has changed. Properly maintained archives protect community memory, support accountable publishing, and give a website’s history a durable place to be found.