<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Track Awesome Web Archiving Updates Daily</title>
  <id>https://www.trackawesomelist.com/iipc/awesome-web-archiving/feed.xml</id>
  <updated>2026-04-28T04:03:50.911Z</updated>
  <link rel="self" type="application/atom+xml" href="https://www.trackawesomelist.com/iipc/awesome-web-archiving/feed.xml"/>
  <link rel="alternate" type="application/json" href="https://www.trackawesomelist.com/iipc/awesome-web-archiving/feed.json"/>
  <link rel="alternate" type="text/html" href="https://www.trackawesomelist.com/iipc/awesome-web-archiving/"/>
  <generator uri="https://github.com/bcomnes/jsonfeed-to-atom#readme" version="1.2.2">jsonfeed-to-atom</generator>
  <icon>https://www.trackawesomelist.com/favicon.ico</icon>
  <logo>https://www.trackawesomelist.com/icon.png</logo>
  <subtitle>An Awesome List for getting started with web archiving</subtitle>
  <entry>
    <id>https://www.trackawesomelist.com/2026/04/28/</id>
    <title>Awesome Web Archiving Updates on Apr 28, 2026</title>
    <updated>2026-04-28T04:03:50.911Z</updated>
    <published>2026-04-28T04:03:50.911Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/midwork-finds-jobs/duckdb_warc" rel="noopener noreferrer">duckdb_warc (⭐7)</a> - DuckDB extension to query WARC files. <em>(In Development)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2026/04/28/"/>
    <summary>1 awesome projects updated on Apr 28, 2026</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2026/04/22/</id>
    <title>Awesome Web Archiving Updates on Apr 22, 2026</title>
    <updated>2026-04-22T14:06:11.524Z</updated>
    <published>2026-04-22T14:06:11.502Z</published>
    <content type="html"><![CDATA[<h3><p>Training/Documentation / Introductions to Web Archiving Concepts</p>
</h3>
<ul>
<li><a href="https://en.wikipedia.org/wiki/List_of_Web_archiving_initiatives" rel="noopener noreferrer">Wikipedia's List of Web Archiving Initiatives</a></li>
</ul>

<ul>
<li><a href="https://support.archive-it.org/hc/en-us/articles/208111686-Glossary-of-Archive-It-and-Web-Archiving-Terms" rel="noopener noreferrer">Glossary of Archive-It and Web Archiving Terms</a></li>
</ul>

<ul>
<li><a href="https://archive-it.org/blog/post/announcing-the-web-archiving-life-cycle-model/" rel="noopener noreferrer">The Web Archiving Lifecycle Model</a> - An attempt to incorporate the technological and programmatic arms of the web archiving into a framework that will be relevant to any organization seeking to archive content from the web. Archive-It, the web archiving service from the Internet Archive, developed the model based on its work with memory institutions around the world.</li>
</ul>

<ul>
<li><a href="https://kit.exposingtheinvisible.org/en/web-archive.html/" rel="noopener noreferrer">Retrieving and Archiving Information from Websites by Wael Eskandar and Brad Murray</a></li>
</ul>
<h3><p>Training/Documentation / Training Materials</p>
</h3>
<ul>
<li><a href="https://github.com/vphill/web-archiving-course" rel="noopener noreferrer">UNT Web Archiving Course (⭐24)</a></li>
</ul>

<ul>
<li><a href="https://cedwarc.github.io/" rel="noopener noreferrer">Continuing Education to Advance Web Archiving (CEDWARC)</a></li>
</ul>

<ul>
<li><a href="https://github.com/commoncrawl/whirlwind-python/" rel="noopener noreferrer">A Whirlwind Tour of Common Crawl's Datasets using Python (⭐45)</a></li>
</ul>

<ul>
<li><a href="https://github.com/commoncrawl/whirlwind-python-notebook" rel="noopener noreferrer">A Whirlwind Tour of Common Crawl's Datasets as a Python notebook (⭐4)</a></li>
</ul>

<ul>
<li><a href="https://github.com/commoncrawl/whirlwind-java/" rel="noopener noreferrer">A Whirlwind Tour of Common Crawl's Datasets using Java (⭐3)</a></li>
</ul>
<h3><p>Training/Documentation / The WARC Standard</p>
</h3>
<ul>
<li><a href="https://iipc.github.io/warc-specifications/" rel="noopener noreferrer">The warc-specifications</a> - A community HTML version of the official specification and hub for new proposals.</li>
</ul>

<ul>
<li><a href="http://bibnum.bnf.fr/WARC/" rel="noopener noreferrer">Offical ISO 28500 WARC specification homepage</a></li>
</ul>
<h3><p>Training/Documentation / For Researchers using Web Archives</p>
</h3>
<ul>
<li><a href="https://aut.docs.archivesunleashed.org/" rel="noopener noreferrer">Archives Unleashed Toolkit documentation</a></li>
</ul>

<ul>
<li><a href="https://sobre.arquivo.pt/en/tutorial-for-humanities-researchers-about-how-to-use-arquivo-pt/" rel="noopener noreferrer">Tutorial for Humanities researchers about how to explore Arquivo.pt</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2026/04/22/"/>
    <summary>13 awesome projects updated on Apr 22, 2026</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2026/03/18/</id>
    <title>Awesome Web Archiving Updates on Mar 18, 2026</title>
    <updated>2026-03-18T03:18:30.154Z</updated>
    <published>2026-03-18T03:18:30.154Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/internetarchive/bagnabit2warc" rel="noopener noreferrer">bagnabit2warc (⭐1)</a> - Convert a <a href="https://github.com/harvard-lil/bag-nabit" rel="noopener noreferrer">bag-nabit (⭐40)</a> dataset stored in a ZIP into a full-content WARC.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2026/03/18/"/>
    <summary>1 awesome projects updated on Mar 18, 2026</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2026/03/04/</id>
    <title>Awesome Web Archiving Updates on Mar 04, 2026</title>
    <updated>2026-03-04T13:20:50.391Z</updated>
    <published>2026-03-04T13:20:50.385Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Quality Assurance</p>
</h3>
<ul>
<li><a href="https://chromewebstore.google.com/detail/check-my-links/ojkcdipcgfaekbeaelaapakgnjflfglf" rel="noopener noreferrer">Chrome Check My Links</a> - Browser extension: a link checker with more options.</li>
</ul>

<ul>
<li><a href="https://chromewebstore.google.com/detail/link-checker/aibjbgmpmnidnmagaefhmcjhadpffaoi" rel="noopener noreferrer">Chrome link checker</a> - Browser extension: basic link checker.</li>
</ul>

<ul>
<li><a href="https://chromewebstore.google.com/detail/bpjdkodgnbfalgghnbeggfbfjpcfamkf/publish-accepted?hl=en-US&amp;gl=US" rel="noopener noreferrer">Chrome link gopher</a> - Browser extension: link harvester on a page.</li>
</ul>

<ul>
<li><a href="https://chromewebstore.google.com/detail/open-multiple-urls/oifijhaokejakekmnjmphonojcfkpbbh?hl=de" rel="noopener noreferrer">Chrome Open Multiple URLs</a> - Browser extension: opens multiple URLs and also extracts URLs from text.</li>
</ul>

<ul>
<li><a href="https://chromewebstore.google.com/detail/revolver-tabs/dlknooajieciikpedpldejhhijacnbda" rel="noopener noreferrer">Chrome Revolver</a> - Browser extension: switches between browser tabs.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2026/03/04/"/>
    <summary>5 awesome projects updated on Mar 04, 2026</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2026/02/08/</id>
    <title>Awesome Web Archiving Updates on Feb 08, 2026</title>
    <updated>2026-02-08T09:11:59.921Z</updated>
    <published>2026-02-08T09:11:59.915Z</published>
    <content type="html"><![CDATA[<h3><p>Public Data / Hosted, Closed Source</p>
</h3>
<ul>
<li><a href="https://data.commoncrawl.org/" rel="noopener noreferrer">Common Crawl files</a> - WARCs, CDX files, parquet url index, parquet host index, etc.</li>
</ul>

<ul>
<li><a href="https://index.commoncrawl.org/" rel="noopener noreferrer">Common Crawl CDX API</a> - Search Common Crawl's CDX URL index.</li>
</ul>

<ul>
<li><a href="https://eotarchive.org/" rel="noopener noreferrer">End of Term Archive</a> - WARCs, CDX files, parquet url index.</li>
</ul>

<ul>
<li><a href="https://govarchive.us/" rel="noopener noreferrer">Webrecorder US GovArchive</a> - High-fidelity replay.</li>
</ul>

<ul>
<li><a href="https://www.nationalarchives.gov.uk/webarchive/" rel="noopener noreferrer">UK Government Web Archive</a> - Main page for the UKGWA.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2026/02/08/"/>
    <summary>5 awesome projects updated on Feb 08, 2026</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2026/01/09/</id>
    <title>Awesome Web Archiving Updates on Jan 09, 2026</title>
    <updated>2026-01-09T02:27:56.230Z</updated>
    <published>2026-01-09T02:27:56.171Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/midwork-finds-jobs/duckdb-web-archive" rel="noopener noreferrer">duckdb-web-archive-cdx (⭐22)</a> - DuckDB extension to query the Internet Archive and CommonCrawl CDX APIs directly from SQL. <em>(In Development)</em></li>
</ul>
<h3><p>Tools &amp; Software / WARC I/O Libraries</p>
</h3>
<ul>
<li><a href="https://github.com/jedireza/warc" rel="noopener noreferrer">warc (⭐61)</a> - A Rust library for reading and writing WARC files. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2026/01/09/"/>
    <summary>2 awesome projects updated on Jan 09, 2026</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2026/01/08/</id>
    <title>Awesome Web Archiving Updates on Jan 08, 2026</title>
    <updated>2026-01-08T02:27:29.281Z</updated>
    <published>2026-01-08T02:27:29.279Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Discord</p>
</h3>
<ul>
<li><a href="https://discord.gg/njaVFh7avF" rel="noopener noreferrer">Common Crawl Foundation</a></li>
</ul>
<h3><p>Community Resources / Twitter</p>
</h3>
<ul>
<li><a href="https://twitter.com/commoncrawl" rel="noopener noreferrer">@commoncrawl</a> - Official Common Crawl Foundation handle.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2026/01/08/"/>
    <summary>2 awesome projects updated on Jan 08, 2026</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2025/11/24/</id>
    <title>Awesome Web Archiving Updates on Nov 24, 2025</title>
    <updated>2025-11-24T16:12:58.278Z</updated>
    <published>2025-11-24T16:12:58.095Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/ArchiveBox/ArchiveBox" rel="noopener noreferrer">ArchiveBox (⭐28k)</a> - A tool which maintains an additive archive from RSS feeds, bookmarks, and links using wget, Chrome headless, and other methods (formerly <code>Bookmark Archiver</code>). <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="https://github.com/PromyLOPh/crocoite" rel="noopener noreferrer">crocoite (⭐45)</a> - Crawl websites using headless Google Chrome/Chromium and save resources, static DOM snapshot and page screenshots to WARC files. <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="https://github.com/DO-SAY-GO/dn" rel="noopener noreferrer">DiskerNet (⭐3.9k)</a> - A non-WARC-based tool which hooks into the Chrome browser and archives everything you browse making it available for offline replay. <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="https://github.com/DocNow/twarc" rel="noopener noreferrer">twarc (⭐1.4k)</a> - A command line tool and Python library for archiving Twitter JSON data. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/ArchiveTeam/wpull" rel="noopener noreferrer">Wpull (⭐613)</a> - A Wget-compatible (or remake/clone/replacement/alternative) web downloader and crawler. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Search &amp; Discovery</p>
</h3>
<ul>
<li><a href="https://github.com/machawk1/Mink" rel="noopener noreferrer">Mink (⭐58)</a> - A <a href="https://www.google.com/intl/en/chrome/" rel="noopener noreferrer">Google Chrome</a> extension for querying Memento aggregators while browsing and integrating live-archived web navigation. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/harvard-lil/warcbench" rel="noopener noreferrer">warcbench (⭐22)</a> - A tool for exploring, analyzing, transforming, recombining, and extracting data from WARC (Web ARChive) files.</li>
</ul>
<h3><p>Tools &amp; Software / Quality Assurance</p>
</h3>
<ul>
<li><a href="https://github.com/flameshot-org/flameshot" rel="noopener noreferrer">FlameShot (⭐31k)</a> - Screen capture and annotation on Ubuntu.</li>
</ul>
<h3><p>Community Resources / Other Awesome Lists</p>
</h3>
<ul>
<li><a href="https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-Community" rel="noopener noreferrer">Web Archiving Community (⭐28k)</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2025/11/24/"/>
    <summary>9 awesome projects updated on Nov 24, 2025</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2025/04/10/</id>
    <title>Awesome Web Archiving Updates on Apr 10, 2025</title>
    <updated>2025-04-10T02:03:34.010Z</updated>
    <published>2025-04-10T02:03:34.002Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Search &amp; Discovery</p>
</h3>
<ul>
<li><a href="https://github.com/ukwa/shine" rel="noopener noreferrer">Shine (⭐43)</a> - A prototype web archives exploration UI, developed with researchers as part of the <a href="https://buddah.projects.history.ac.uk/" rel="noopener noreferrer">Big UK Domain Data for the Arts and Humanities project</a>. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/netarchivesuite/solrwayback" rel="noopener noreferrer">SolrWayback (⭐146)</a> - A backend Java and frontend VUE JS project with freetext search and a build in playback engine. Require Warc files has been index with the Warc-Indexer. The web application also has a wide range of data visualization tools and data export tools that can be used on the whole webarchive. <a href="https://github.com/netarchivesuite/solrwayback/releases" rel="noopener noreferrer">SolrWayback 4 Bundle release (⭐146)</a> contains all the software and dependencies in an out-of-the box solution that is easy to install.</li>
</ul>

<ul>
<li><a href="https://github.com/archivesunleashed/warclight" rel="noopener noreferrer">Warclight (⭐50)</a> - A Project Blacklight based Rails engine that supports the discovery of web archives held in the WARC and ARC formats. <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="https://github.com/webis-de/wasp" rel="noopener noreferrer">Wasp (⭐28)</a> - A fully functional prototype of a personal <a href="http://ceur-ws.org/Vol-2167/paper6.pdf" rel="noopener noreferrer">web archive and search system</a>. <em>(In Development)</em></li>
</ul>

<ul>
<li>Other possible options for builting a front-end are listed on in the <code>webarchive-discovery</code> wiki, <a href="https://github.com/ukwa/webarchive-discovery/wiki/Front-ends" rel="noopener noreferrer">here (⭐133)</a>.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2025/04/10/"/>
    <summary>5 awesome projects updated on Apr 10, 2025</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2025/02/12/</id>
    <title>Awesome Web Archiving Updates on Feb 12, 2025</title>
    <updated>2025-02-12T01:52:58.803Z</updated>
    <published>2025-02-12T01:52:58.803Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://www.community-archive.org/" rel="noopener noreferrer">Community Archive</a> - Open Twitter Database and API with tools and resources for building on archived Twitter data.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2025/02/12/"/>
    <summary>1 awesome projects updated on Feb 12, 2025</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2025/01/31/</id>
    <title>Awesome Web Archiving Updates on Jan 31, 2025</title>
    <updated>2025-01-31T15:10:15.546Z</updated>
    <published>2025-01-31T15:10:15.522Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Search &amp; Discovery</p>
</h3>
<ul>
<li><a href="https://github.com/medialab/hyphe" rel="noopener noreferrer">hyphe (⭐385)</a> - A webcrawler built for research uses with a graphical user interface in order to build web corpuses made of lists of web actors and maps of links between them. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/Guillaume-Levrier/PANDORAE" rel="noopener noreferrer">PANDORÆ (⭐16)</a> - A desktop research software to be plugged on a Solr endpoint to query, retrieve, normalize and visually explore web archives. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://httpreserve.info" rel="noopener noreferrer">httpreserve.info</a> - Service to return the status of a web page or save it to the Internet Archive. HTTPreserve includes disambiguation of well-known short link services. It returns JSON via the browser or command line via CURL using GET. Describes web sites using earliest and latest dates in the Internet Archive and demonstrates the construction of Robust Links in its output using that range. (Golang). <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2025/01/31/"/>
    <summary>3 awesome projects updated on Jan 31, 2025</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2025/01/22/</id>
    <title>Awesome Web Archiving Updates on Jan 22, 2025</title>
    <updated>2025-01-22T06:39:59.616Z</updated>
    <published>2025-01-22T06:39:59.489Z</published>
    <content type="html"><![CDATA[<h3><p>Resources for Web Publishers / For Researchers using Web Archives</p>
</h3>
<ul>
<li><a href="https://nullhandle.org/web-archivability/index.html" rel="noopener noreferrer">Definition of Web Archivability</a> - This describes the ease with which web content can be preserved. (<a href="https://web.archive.org/web/20230728211501/https://library.stanford.edu/projects/web-archiving/archivability" rel="noopener noreferrer">Archived version from the Stanford Libraries</a>)</li>
</ul>
<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://theunarchiver.com/" rel="noopener noreferrer">The Unarchiver</a> - Program to extract the contents of many archive formats, inclusive of WARC, to a file system. Free variant of The Archive Browser (macOS only, Proprietary app).</li>
</ul>
<h3><p>Tools &amp; Software / WARC I/O Libraries</p>
</h3>
<ul>
<li><a href="https://github.com/chfoo/warcat-rs" rel="noopener noreferrer">Warcat-rs (⭐34)</a> - Command-line tool and Rust library for handling Web ARChive (WARC) files. <em>(In Development)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2025/01/22/"/>
    <summary>3 awesome projects updated on Jan 22, 2025</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2024/11/12/</id>
    <title>Awesome Web Archiving Updates on Nov 12, 2024</title>
    <updated>2024-11-12T12:55:28.730Z</updated>
    <published>2024-11-12T12:55:28.536Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://webrecorder.net/archivewebpage/" rel="noopener noreferrer">ArchiveWeb.Page</a> - A plugin for Chrome and other Chromium based browsers that lets you interactively archive web pages, replay them, and export them as WARC &amp; WACZ files. Also available as an Electron based desktop application.</li>
</ul>
<h3><p>Tools &amp; Software / Replay</p>
</h3>
<ul>
<li><a href="https://webrecorder.net/replaywebpage/" rel="noopener noreferrer">ReplayWeb.page</a> - A browser-based, fully client-side replay engine for both local and remote WARC &amp; WACZ files. Also available as an Electron based desktop application. <em>(Stable)</em></li>
</ul>
<h3><p>Community Resources / Blogs and Scholarship</p>
</h3>
<ul>
<li><a href="https://commoncrawl.org/blog" rel="noopener noreferrer">Common Crawl Foundation Blog</a> - <a href="http://commoncrawl.org/blog/rss.xml" rel="noopener noreferrer">rss</a></li>
</ul>
<h3><p>Community Resources / Slack</p>
</h3>
<ul>
<li><a href="https://ccfpartners.slack.com/" rel="noopener noreferrer">Common Crawl Foundation Partners</a> (ask greg zat commoncrawl zot org for an invite)</li>
</ul>
<h3><p>Web Archiving Service Providers / Self-hostable, Open Source</p>
</h3>
<ul>
<li><a href="https://webrecorder.net/browsertrix/" rel="noopener noreferrer">Browsertrix</a> - From <a href="https://webrecorder.net/" rel="noopener noreferrer">Webrecorder</a>, source available at <a href="https://github.com/webrecorder/browsertrix" rel="noopener noreferrer">https://github.com/webrecorder/browsertrix (⭐458)</a>.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2024/11/12/"/>
    <summary>5 awesome projects updated on Nov 12, 2024</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2024/10/18/</id>
    <title>Awesome Web Archiving Updates on Oct 18, 2024</title>
    <updated>2024-10-18T01:54:06.918Z</updated>
    <published>2024-10-18T01:54:06.918Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="http://mementoweb.github.io/SiteStory/" rel="noopener noreferrer">SiteStory</a> - A transactional archive that selectively captures and stores transactions that take place between a web client (browser) and a web server. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2024/10/18/"/>
    <summary>1 awesome projects updated on Oct 18, 2024</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2024/05/09/</id>
    <title>Awesome Web Archiving Updates on May 09, 2024</title>
    <updated>2024-05-09T01:31:26.124Z</updated>
    <published>2024-05-09T01:31:26.122Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / WARC I/O Libraries</p>
</h3>
<ul>
<li><a href="https://github.com/netarchivesuite/jwat" rel="noopener noreferrer">Jwat (⭐4)</a> - Libraries for reading/writing/validating WARC/ARC/GZIP files (Java). <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/netarchivesuite/jwat-tools" rel="noopener noreferrer">Jwat-Tools (⭐5)</a> - Tools for reading/writing/validating WARC/ARC/GZIP files (Java). <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2024/05/09/"/>
    <summary>2 awesome projects updated on May 09, 2024</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2024/05/06/</id>
    <title>Awesome Web Archiving Updates on May 06, 2024</title>
    <updated>2024-05-06T12:39:56.782Z</updated>
    <published>2024-05-06T12:39:56.782Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/natliblux/warc-safe" rel="noopener noreferrer">warc-safe (⭐18)</a> - Automatic detection of viruses and NSFW content in WARC files.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2024/05/06/"/>
    <summary>1 awesome projects updated on May 06, 2024</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2024/04/25/</id>
    <title>Awesome Web Archiving Updates on Apr 25, 2024</title>
    <updated>2024-04-25T12:32:06.150Z</updated>
    <published>2024-04-25T12:32:06.150Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Replay</p>
</h3>
<ul>
<li><a href="https://github.com/webrecorder/pywb" rel="noopener noreferrer">PYWB (⭐1.7k)</a> - A Python 3 implementation of web archival replay tools, sometimes also known as 'Wayback Machine'. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2024/04/25/"/>
    <summary>1 awesome projects updated on Apr 25, 2024</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2024/01/19/</id>
    <title>Awesome Web Archiving Updates on Jan 19, 2024</title>
    <updated>2024-01-19T01:34:03.829Z</updated>
    <published>2024-01-19T01:34:03.821Z</published>
    <content type="html"><![CDATA[<h3><p>Web Archiving Service Providers / Self-hostable, Open Source</p>
</h3>
<ul>
<li><a href="https://conifer.rhizome.org/" rel="noopener noreferrer">Conifer</a> - From <a href="https://rhizome.org/" rel="noopener noreferrer">Rhizome</a>, source available at <a href="https://github.com/Rhizome-Conifer" rel="noopener noreferrer">https://github.com/Rhizome-Conifer</a>.</li>
</ul>
<h3><p>Web Archiving Service Providers / Hosted, Closed Source</p>
</h3>
<ul>
<li><a href="https://archive-it.org/" rel="noopener noreferrer">Archive-It</a> - From the Internet Archive.</li>
</ul>

<ul>
<li><a href="https://arkiwera.se/wp/websites/" rel="noopener noreferrer">Arkiwera</a></li>
</ul>

<ul>
<li><a href="https://www.hanzo.co/chronicle" rel="noopener noreferrer">Hanzo</a></li>
</ul>

<ul>
<li><a href="https://www.mirrorweb.com/solutions/capabilities/website-archiving" rel="noopener noreferrer">MirrorWeb</a></li>
</ul>

<ul>
<li><a href="https://www.pagefreezer.com/" rel="noopener noreferrer">PageFreezer</a></li>
</ul>

<ul>
<li><a href="https://www.smarsh.com/platform/compliance-management/web-archive" rel="noopener noreferrer">Smarsh</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2024/01/19/"/>
    <summary>7 awesome projects updated on Jan 19, 2024</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/12/26/</id>
    <title>Awesome Web Archiving Updates on Dec 26, 2023</title>
    <updated>2023-12-26T23:14:46.443Z</updated>
    <published>2023-12-26T23:14:46.443Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/jjjake/internetarchive" rel="noopener noreferrer">Internet Archive Library (⭐1.9k)</a> - A command line tool and Python library for interacting directly with <a href="https://archive.org" rel="noopener noreferrer">archive.org</a>. (Python). <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/12/26/"/>
    <summary>1 awesome projects updated on Dec 26, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/10/17/</id>
    <title>Awesome Web Archiving Updates on Oct 17, 2023</title>
    <updated>2023-10-17T01:23:22.130Z</updated>
    <published>2023-10-17T01:23:22.130Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/florents-Tselai/warcdb" rel="noopener noreferrer">warcdb (⭐407)</a> - A command line utility (Python) for importing WARC files into a SQLite database. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/10/17/"/>
    <summary>1 awesome projects updated on Oct 17, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/08/30/</id>
    <title>Awesome Web Archiving Updates on Aug 30, 2023</title>
    <updated>2023-08-30T12:38:46.010Z</updated>
    <published>2023-08-30T12:38:46.010Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/httpreserve/linkstat" rel="noopener noreferrer">HTTPreserve linkstat (⭐10)</a> - Command line implementation of <a href="https://httpreserve.info" rel="noopener noreferrer">httpreserve.info</a> to describe the status of a web page. Can be easily scripted and provides JSON output to enable querying through tools like JQ. HTTPreserve Linkstat describes current status, and earliest and latest links on <a href="https://archive.org/" rel="noopener noreferrer">archive.org</a>. (Golang). <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/08/30/"/>
    <summary>1 awesome projects updated on Aug 30, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/07/14/</id>
    <title>Awesome Web Archiving Updates on Jul 14, 2023</title>
    <updated>2023-07-14T12:47:06.786Z</updated>
    <published>2023-07-14T12:47:06.786Z</published>
    <content type="html"><![CDATA[<h3><p>Training/Documentation / Training Materials</p>
</h3>
<ul>
<li><a href="https://netpreserve.org/web-archiving/training-materials/" rel="noopener noreferrer">IIPC and DPC Training materials: module for beginners (8 sessions)</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/07/14/"/>
    <summary>1 awesome projects updated on Jul 14, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/07/05/</id>
    <title>Awesome Web Archiving Updates on Jul 05, 2023</title>
    <updated>2023-07-05T02:04:52.400Z</updated>
    <published>2023-07-05T02:04:52.395Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Analysis</p>
</h3>
<ul>
<li><a href="https://commoncrawl.org/tag/columnar-index/" rel="noopener noreferrer">Common Crawl Columnar Index</a> - SQL-queryable index, with CDX info plus language classification. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://commoncrawl.org/category/web-graph/" rel="noopener noreferrer">Common Crawl Web Graph</a> - A host or domain-level graph of the web, with ranking information. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/commoncrawl/cc-notebooks" rel="noopener noreferrer">Common Crawl Jupyter notebooks (⭐67)</a> - A collection of notebooks using Common Crawl's various datasets. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/07/05/"/>
    <summary>3 awesome projects updated on Jul 05, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/07/04/</id>
    <title>Awesome Web Archiving Updates on Jul 04, 2023</title>
    <updated>2023-07-04T12:49:09.093Z</updated>
    <published>2023-07-04T12:49:08.895Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://pypi.org/project/cdx-toolkit/" rel="noopener noreferrer">cdx-toolkit</a> - Library and CLI to consult cdx indexes and create WARC extractions of subsets. Abstracts away Common Crawl's unusual crawl structure. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Analysis</p>
</h3>
<ul>
<li><a href="http://webdatacommons.org/" rel="noopener noreferrer">Web Data Commons</a> - Structured data extracted from Common Crawl. <em>(Stable)</em></li>
</ul>
<h3><p>Community Resources / Mailing Lists</p>
</h3>
<ul>
<li><a href="https://groups.google.com/g/common-crawl" rel="noopener noreferrer">Common Crawl</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/07/04/"/>
    <summary>3 awesome projects updated on Jul 04, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/06/02/</id>
    <title>Awesome Web Archiving Updates on Jun 02, 2023</title>
    <updated>2023-06-02T02:00:06.738Z</updated>
    <published>2023-06-02T02:00:06.738Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/karust/gogetcrawl" rel="noopener noreferrer">Go Get Crawl (⭐184)</a> - Extract web archive data using <a href="https://web.archive.org/" rel="noopener noreferrer">Wayback Machine</a> and <a href="https://commoncrawl.org/" rel="noopener noreferrer">Common Crawl</a>. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/06/02/"/>
    <summary>1 awesome projects updated on Jun 02, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/05/26/</id>
    <title>Awesome Web Archiving Updates on May 26, 2023</title>
    <updated>2023-05-26T06:01:47.672Z</updated>
    <published>2023-05-26T06:01:47.672Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/turicas/crau" rel="noopener noreferrer">crau (⭐64)</a> - A lightweight command-line tool for archiving the Web and playing archives: you just need a list of URLs. The name "crau" stems from the Brazilian pronounciation of "crawl". <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/05/26/"/>
    <summary>1 awesome projects updated on May 26, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/04/27/</id>
    <title>Awesome Web Archiving Updates on Apr 27, 2023</title>
    <updated>2023-04-27T01:42:04.749Z</updated>
    <published>2023-04-27T01:42:04.749Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/harvard-lil/scoop" rel="noopener noreferrer">Scoop (⭐209)</a> - High-fidelity, browser-based, single-page web archiving library and CLI for witnessing the web. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/04/27/"/>
    <summary>1 awesome projects updated on Apr 27, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/04/19/</id>
    <title>Awesome Web Archiving Updates on Apr 19, 2023</title>
    <updated>2023-04-19T01:42:10.072Z</updated>
    <published>2023-04-19T01:42:10.072Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://gitlab.com/taricorp/warcdedupe" rel="noopener noreferrer">warcdedupe</a> - WARC deduplication tool (and WARC library) written in Rust. <em>(In Development)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/04/19/"/>
    <summary>1 awesome projects updated on Apr 19, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2023/04/13/</id>
    <title>Awesome Web Archiving Updates on Apr 13, 2023</title>
    <updated>2023-04-13T01:38:14.326Z</updated>
    <published>2023-04-13T01:38:14.322Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://nlnwa.github.io/warchaeology/" rel="noopener noreferrer">Warchaeology</a> - A collection of tools for inspecting, manipulating, deduplicating and validating WARC-files. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/arcalex/warcrefs" rel="noopener noreferrer">warcrefs (⭐10)</a> - Web archive deduplication tools. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2023/04/13/"/>
    <summary>2 awesome projects updated on Apr 13, 2023</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2022/09/27/</id>
    <title>Awesome Web Archiving Updates on Sep 27, 2022</title>
    <updated>2022-09-27T14:46:37.000Z</updated>
    <published>2022-09-27T14:46:37.000Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Blogs and Scholarship</p>
</h3>
<ul>
<li><a href="https://ws-dl.blogspot.com/" rel="noopener noreferrer">WS-DL Blog</a> - Web Science and Digital Libraries Research Group blogs about various Web archiving related topics, scholarly work, and academic trip reports.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2022/09/27/"/>
    <summary>1 awesome projects updated on Sep 27, 2022</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2022/09/24/</id>
    <title>Awesome Web Archiving Updates on Sep 24, 2022</title>
    <updated>2022-09-24T02:38:51.000Z</updated>
    <published>2022-09-24T02:38:51.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/bellingcat/auto-archiver" rel="noopener noreferrer">Auto Archiver (⭐1.1k)</a> - Python script to automatically archive social media posts, videos, and images from a Google Sheets document. Read the <a href="https://www.bellingcat.com/resources/2022/09/22/preserve-vital-online-content-with-bellingcats-auto-archiver-tool/" rel="noopener noreferrer">article about Auto Archiver on bellingcat.com</a>.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2022/09/24/"/>
    <summary>1 awesome projects updated on Sep 24, 2022</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2022/08/23/</id>
    <title>Awesome Web Archiving Updates on Aug 23, 2022</title>
    <updated>2022-08-23T16:10:15.000Z</updated>
    <published>2022-08-23T16:10:15.000Z</published>
    <content type="html"><![CDATA[<h3><p>Training/Documentation / For Researchers using Web Archives</p>
</h3>
<ul>
<li><a href="https://glam-workbench.github.io/web-archives/" rel="noopener noreferrer">GLAM Workbench: Web Archives</a> - See also <a href="https://netpreserveblog.wordpress.com/2020/05/28/asking-questions-with-web-archives/" rel="noopener noreferrer">this related blog post on 'Asking questions with web archives'</a>.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2022/08/23/"/>
    <summary>1 awesome projects updated on Aug 23, 2022</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2022/05/25/</id>
    <title>Awesome Web Archiving Updates on May 25, 2022</title>
    <updated>2022-05-25T21:30:37.000Z</updated>
    <published>2022-05-25T21:30:37.000Z</published>
    <content type="html"><![CDATA[<h3><p>Training/Documentation / Introductions to Web Archiving Concepts</p>
</h3>
<ul>
<li><a href="https://youtu.be/ubDHY-ynWi0" rel="noopener noreferrer">What is a web archive?</a> - A video from <a href="https://www.youtube.com/channel/UCJukhTSw8VRj-VNTpBcqWkw" rel="noopener noreferrer">the UK Web Archive YouTube Channel</a></li>
</ul>
<h3><p>Tools &amp; Software / WARC I/O Libraries</p>
</h3>
<ul>
<li><a href="https://github.com/internetarchive/Sparkling" rel="noopener noreferrer">Sparkling (⭐17)</a> - Internet Archive's Sparkling Data Processing Library. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Analysis</p>
</h3>
<ul>
<li><a href="https://github.com/internetarchive/arch" rel="noopener noreferrer">Archives Research Compute Hub (⭐20)</a> - Web application for distributed compute analysis of Archive-It web archive collections. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2022/05/25/"/>
    <summary>3 awesome projects updated on May 25, 2022</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2022/03/03/</id>
    <title>Awesome Web Archiving Updates on Mar 03, 2022</title>
    <updated>2022-03-03T17:19:52.000Z</updated>
    <published>2022-03-03T17:19:52.000Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Blogs and Scholarship</p>
</h3>
<ul>
<li><a href="https://www.uclpress.co.uk/products/84010" rel="noopener noreferrer">The Web as History</a> - An open-source book that provides a conceptual overview to web archiving research, as well as several case studies.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2022/03/03/"/>
    <summary>1 awesome projects updated on Mar 03, 2022</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2022/01/23/</id>
    <title>Awesome Web Archiving Updates on Jan 23, 2022</title>
    <updated>2022-01-23T04:51:23.000Z</updated>
    <published>2022-01-23T04:51:23.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/akamhy/waybackpy" rel="noopener noreferrer">Waybackpy (⭐600)</a> -  Wayback Machine Save, CDX and availability API interface in Python and a command-line tool  <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2022/01/23/"/>
    <summary>1 awesome projects updated on Jan 23, 2022</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2022/01/05/</id>
    <title>Awesome Web Archiving Updates on Jan 05, 2022</title>
    <updated>2022-01-05T15:32:05.000Z</updated>
    <published>2022-01-05T15:32:05.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / WARC I/O Libraries</p>
</h3>
<ul>
<li><a href="https://github.com/emmadickson/unwarcit" rel="noopener noreferrer">Unwarcit (⭐13)</a> - Command line interface to unzip WARC and WACZ files (Python).</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2022/01/05/"/>
    <summary>1 awesome projects updated on Jan 05, 2022</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/12/13/</id>
    <title>Awesome Web Archiving Updates on Dec 13, 2021</title>
    <updated>2021-12-13T16:30:53.000Z</updated>
    <published>2021-12-13T16:30:53.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / WARC I/O Libraries</p>
</h3>
<ul>
<li><a href="https://github.com/chatnoir-eu/chatnoir-resiliparse" rel="noopener noreferrer">FastWARC (⭐144)</a> - A high-performance WARC parsing library (Python).</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/12/13/"/>
    <summary>1 awesome projects updated on Dec 13, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/11/08/</id>
    <title>Awesome Web Archiving Updates on Nov 08, 2021</title>
    <updated>2021-11-08T05:53:24.000Z</updated>
    <published>2021-11-08T05:53:24.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Replay</p>
</h3>
<ul>
<li><a href="https://github.com/iipc/warc2html" rel="noopener noreferrer">warc2html (⭐59)</a> - Converts WARC files to static HTML suitable for browsing offline or rehosting.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/11/08/"/>
    <summary>1 awesome projects updated on Nov 08, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/10/07/</id>
    <title>Awesome Web Archiving Updates on Oct 07, 2021</title>
    <updated>2021-10-07T12:45:48.000Z</updated>
    <published>2021-10-07T12:45:48.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/wabarc/wayback" rel="noopener noreferrer">Wayback (⭐2.2k)</a> - A toolkit for snapshot webpage to Internet Archive, archive.today, IPFS and beyond. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/10/07/"/>
    <summary>1 awesome projects updated on Oct 07, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/07/20/</id>
    <title>Awesome Web Archiving Updates on Jul 20, 2021</title>
    <updated>2021-07-20T13:15:59.000Z</updated>
    <published>2021-07-20T13:15:59.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/nlnwa/gowarcserver" rel="noopener noreferrer">gowarcserver (⭐18)</a> - <a href="https://github.com/dgraph-io/badger" rel="noopener noreferrer">BadgerDB (⭐16k)</a>-based capture index (CDX) and WARC record server, used to index and serve WARC files (Go).</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/07/20/"/>
    <summary>1 awesome projects updated on Jul 20, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/07/13/</id>
    <title>Awesome Web Archiving Updates on Jul 13, 2021</title>
    <updated>2021-07-13T12:33:08.000Z</updated>
    <published>2021-07-13T12:33:08.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://webcuratortool.org" rel="noopener noreferrer">Web Curator Tool</a> - Open-source workflow management for selective web archiving. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Curation</p>
</h3>
<ul>
<li><a href="https://robustlinks.mementoweb.org/zotero/" rel="noopener noreferrer">Zotero Robust Links Extension</a> - A <a href="https://www.zotero.org/" rel="noopener noreferrer">Zotero</a> extension that submits to and reads from web archives. Source <a href="https://github.com/lanl/Zotero-Robust-Links-Extension" rel="noopener noreferrer">on GitHub (⭐22)</a>. Supercedes <a href="https://github.com/leonkt/zotero-memento" rel="noopener noreferrer">leonkt/zotero-memento (⭐357)</a>.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/07/13/"/>
    <summary>2 awesome projects updated on Jul 13, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/06/22/</id>
    <title>Awesome Web Archiving Updates on Jun 22, 2021</title>
    <updated>2021-06-22T12:39:37.000Z</updated>
    <published>2021-06-22T12:39:37.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/machawk1/wail" rel="noopener noreferrer">WAIL (⭐400)</a> - A graphical user interface (GUI) atop multiple web archiving tools intended to be used as an easy way for anyone to preserve and replay web pages; <a href="https://machawk1.github.io/wail/" rel="noopener noreferrer">Python</a>, <a href="https://github.com/n0tan3rd/wail" rel="noopener noreferrer">Electron (⭐128)</a>. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/internetarchive/warcprox" rel="noopener noreferrer">Warcprox (⭐461)</a> - WARC-writing MITM HTTP/S proxy. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/06/22/"/>
    <summary>2 awesome projects updated on Jun 22, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/05/28/</id>
    <title>Awesome Web Archiving Updates on May 28, 2021</title>
    <updated>2021-05-28T18:45:29.000Z</updated>
    <published>2021-05-28T18:45:29.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/webrecorder/browsertrix-crawler" rel="noopener noreferrer">Browsertrix Crawler (⭐1.1k)</a> - A Chromium based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/internetarchive/brozzler" rel="noopener noreferrer">Brozzler (⭐813)</a> - A distributed web crawler (爬虫) that uses a real browser (Chrome or Chromium) to fetch pages and embedded urls and to extract links. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/05/28/"/>
    <summary>2 awesome projects updated on May 28, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/04/27/</id>
    <title>Awesome Web Archiving Updates on Apr 27, 2021</title>
    <updated>2021-04-27T18:46:18.000Z</updated>
    <published>2021-04-27T18:46:18.000Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Twitter</p>
</h3>
<ul>
<li><a href="https://twitter.com/WebSciDL" rel="noopener noreferrer">@WebSciDL</a> - ODU Web Science and Digital Libraries Research Group.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/04/27/"/>
    <summary>1 awesome projects updated on Apr 27, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/04/24/</id>
    <title>Awesome Web Archiving Updates on Apr 24, 2021</title>
    <updated>2021-04-24T17:36:54.000Z</updated>
    <published>2021-04-24T17:36:54.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Search &amp; Discovery</p>
</h3>
<ul>
<li><a href="https://github.com/wabarc/playback" rel="noopener noreferrer">playback (⭐13)</a> - A toolkit for searching archived webpages from <a href="https://web.archive.org" rel="noopener noreferrer">Internet Archive</a>, <a href="https://archive.today" rel="noopener noreferrer">archive.today</a>, <a href="http://timetravel.mementoweb.org" rel="noopener noreferrer">Memento</a> and beyond. <em>(In Development)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/04/24/"/>
    <summary>1 awesome projects updated on Apr 24, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2021/04/16/</id>
    <title>Awesome Web Archiving Updates on Apr 16, 2021</title>
    <updated>2021-04-16T13:08:07.000Z</updated>
    <published>2021-04-16T13:08:07.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/nla/httrack2warc" rel="noopener noreferrer">httrack2warc (⭐35)</a> - Convert HTTrack archives to WARC format (Java).</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2021/04/16/"/>
    <summary>1 awesome projects updated on Apr 16, 2021</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/11/06/</id>
    <title>Awesome Web Archiving Updates on Nov 06, 2020</title>
    <updated>2020-11-06T18:41:21.000Z</updated>
    <published>2020-11-06T18:41:21.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/wabarc/cairn" rel="noopener noreferrer">Cairn (⭐51)</a> - A npm package and CLI tool for saving webpages. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/go-shiori/obelisk" rel="noopener noreferrer">Obelisk (⭐320)</a> - Go package and CLI tool for saving web page as single HTML file. <em>(Stable)</em></li>
</ul>
<h3><p>Community Resources / Mailing Lists</p>
</h3>
<ul>
<li><a href="https://groups.google.com/g/openwayback-dev" rel="noopener noreferrer">OpenWayback</a></li>
</ul>

<ul>
<li><a href="https://groups.google.com/g/wasapi-community" rel="noopener noreferrer">WASAPI</a></li>
</ul>
<h3><p>Community Resources / Slack</p>
</h3>
<ul>
<li><a href="https://iipc.slack.com/" rel="noopener noreferrer">IIPC Slack</a> - Ask <a href="https://twitter.com/NetPreserve?s=20" rel="noopener noreferrer">@netpreserve</a> for access.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/11/06/"/>
    <summary>5 awesome projects updated on Nov 06, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/09/18/</id>
    <title>Awesome Web Archiving Updates on Sep 18, 2020</title>
    <updated>2020-09-18T01:01:26.000Z</updated>
    <published>2020-09-18T01:01:26.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Replay</p>
</h3>
<ul>
<li><a href="https://github.com/iipc/openwayback/" rel="noopener noreferrer">OpenWayback (⭐525)</a> - The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. <em>(Stable)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/09/18/"/>
    <summary>1 awesome projects updated on Sep 18, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/09/16/</id>
    <title>Awesome Web Archiving Updates on Sep 16, 2020</title>
    <updated>2020-09-16T21:37:20.000Z</updated>
    <published>2020-09-16T21:37:20.000Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Blogs and Scholarship</p>
</h3>
<ul>
<li><a href="https://blogs.bl.uk/webarchive/" rel="noopener noreferrer">UK Web Archive Blog</a></li>
</ul>
<h3><p>Community Resources / Twitter</p>
</h3>
<ul>
<li><a href="https://twitter.com/hashtag/webarchivewednesday" rel="noopener noreferrer">#WebArchiveWednesday</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/09/16/"/>
    <summary>2 awesome projects updated on Sep 16, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/06/05/</id>
    <title>Awesome Web Archiving Updates on Jun 05, 2020</title>
    <updated>2020-06-05T10:40:01.000Z</updated>
    <published>2020-06-05T10:40:01.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/internetarchive/heritrix3/wiki" rel="noopener noreferrer">Heritrix (⭐3.3k)</a> - An open source, extensible, web-scale, archival quality web crawler. <em>(Stable)</em><ul>
<li><a href="https://github.com/internetarchive/heritrix3/discussions/categories/q-a" rel="noopener noreferrer">Heritrix Q&amp;A (⭐3.3k)</a> - A discussion forum for asking questions and getting answers about using Heritrix.</li>
<li><a href="https://github.com/web-archive-group/heritrix-walkthrough" rel="noopener noreferrer">Heritrix Walkthrough (⭐10)</a> <em>(In Development)</em></li>
</ul>
</li>
</ul>

<ul>
<li><a href="https://github.com/WebMemex" rel="noopener noreferrer">WebMemex</a> - Browser extension for Firefox and Chrome which lets you archive web pages you visit. <em>(In Development)</em></li>
</ul>
<h3><p>Community Resources / Other Awesome Lists</p>
</h3>
<ul>
<li><a href="https://github.com/machawk1/awesome-memento" rel="noopener noreferrer">Awesome Memento (⭐121)</a></li>
</ul>

<ul>
<li><a href="http://www.archiveteam.org/index.php?title=The_WARC_Ecosystem" rel="noopener noreferrer">The WARC Ecosystem</a></li>
</ul>

<ul>
<li><a href="http://coptr.digipres.org/Category:Web_Crawl" rel="noopener noreferrer">The Web Crawl section of COPTR</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/06/05/"/>
    <summary>5 awesome projects updated on Jun 05, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/03/26/</id>
    <title>Awesome Web Archiving Updates on Mar 26, 2020</title>
    <updated>2020-03-26T21:00:22.000Z</updated>
    <published>2020-03-26T21:00:22.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Acquisition</p>
</h3>
<ul>
<li><a href="https://github.com/oduwsdl/archivenow" rel="noopener noreferrer">archivenow (⭐434)</a> - A <a href="http://ws-dl.blogspot.com/2017/02/2017-02-22-archive-now-archivenow.html" rel="noopener noreferrer">Python library</a> to push web resources into on-demand web archives. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/CGamesPlay/chronicler" rel="noopener noreferrer">Chronicler (⭐92)</a> - Web browser with record and replay functionality. <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="https://git.autistici.org/ale/crawl" rel="noopener noreferrer">Crawl</a> - A simple web crawler in Golang. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/justinlittman/fbarc" rel="noopener noreferrer">F(b)arc (⭐78)</a> - A commandline tool and Python library for archiving data from <a href="https://www.facebook.com/" rel="noopener noreferrer">Facebook</a> using the <a href="https://developers.facebook.com/docs/graph-api" rel="noopener noreferrer">Graph API</a>. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/WebMemex/freeze-dry" rel="noopener noreferrer">freeze-dry (⭐305)</a> - JavaScript library to turn page into static, self-contained HTML document; useful for browser extensions. <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="https://github.com/ArchiveTeam/grab-site" rel="noopener noreferrer">grab-site (⭐1.6k)</a> - The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/steffenfritz/html2warc" rel="noopener noreferrer">html2warc (⭐24)</a> - A simple script to convert offline data into a single WARC file. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="http://www.httrack.com/" rel="noopener noreferrer">HTTrack</a> - An open source website copying utility. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/Y2Z/monolith" rel="noopener noreferrer">monolith (⭐15k)</a> - CLI tool to save a web page as a single HTML file. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/gildas-lormeau/SingleFile" rel="noopener noreferrer">SingleFile (⭐22k)</a> - Browser extension for Firefox/Chrome and CLI tool to save a faithful copy of a complete page as a single HTML file. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://gwu-libraries.github.io/sfm-ui/" rel="noopener noreferrer">Social Feed Manager</a> - Open source software that enables users to create social media collections from Twitter, Tumblr, Flickr, and Sina Weibo public APIs. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/N0taN3rd/Squidwarc" rel="noopener noreferrer">Squidwarc (⭐178)</a> - An <a href="http://ws-dl.blogspot.com/2017/07/2017-07-24-replacing-heritrix-with.html" rel="noopener noreferrer">open source, high-fidelity, page interacting</a> archival crawler that uses Chrome or Chrome Headless directly. <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="http://stormcrawler.net/" rel="noopener noreferrer">StormCrawler</a> - A collection of resources for building low-latency, scalable web crawlers on Apache Storm. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="http://matkelly.com/warcreate/" rel="noopener noreferrer">WARCreate</a> - A <a href="https://www.google.com/intl/en/chrome/browser/" rel="noopener noreferrer">Google Chrome</a> extension for archiving an individual webpage or website to a WARC file. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/peterk/warcworker" rel="noopener noreferrer">Warcworker (⭐62)</a> - An open source, dockerized, queued, high fidelity web archiver based on Squidwarc with a simple web GUI. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/helgeho/Web2Warc" rel="noopener noreferrer">Web2Warc (⭐26)</a> - An easy-to-use and highly customizable crawler that enables anyone to create their own little Web archives (WARC/CDX). <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="http://www.gnu.org/software/wget/" rel="noopener noreferrer">Wget</a> - An open source file retrieval utility that of <a href="http://www.archiveteam.org/index.php?title=Wget_with_WARC_output" rel="noopener noreferrer">version 1.14 supports writing warcs</a>. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/alard/wget-lua" rel="noopener noreferrer">Wget-lua (⭐24)</a> - Wget with Lua extension. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Search &amp; Discovery</p>
</h3>
<ul>
<li><a href="http://tempas.L3S.de/v1" rel="noopener noreferrer">Tempas v1</a> - Temporal web archive search based on <a href="https://en.wikipedia.org/wiki/Delicious_(website)" rel="noopener noreferrer">Delicious</a> tags. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="http://tempas.L3S.de/v2" rel="noopener noreferrer">Tempas v2</a> - Temporal web archive search based on links and anchor texts extracted from the German web from 1996 to 2013 (results are not limited to German pages, e.g., <a href="http://tempas.l3s.de/v2/query?q=obama&amp;from=2005&amp;to=2009" rel="noopener noreferrer">Obama@2005-2009 in Tempas</a>). <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/oduwsdl/MementoMap" rel="noopener noreferrer">MementoMap (⭐12)</a> - A Tool to Summarize Web Archive Holdings (Python). <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="https://github.com/oduwsdl/MemGator" rel="noopener noreferrer">MemGator (⭐80)</a> - A Memento Aggregator CLI and Server (Golang). <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/N0taN3rd/node-cdxj" rel="noopener noreferrer">node-cdxj (⭐2)</a> - <a href="https://github.com/oduwsdl/ORS/wiki/CDXJ" rel="noopener noreferrer">CDXJ (⭐15)</a> file parser (Node.js). <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/nla/outbackcdx" rel="noopener noreferrer">OutbackCDX (⭐43)</a> - RocksDB-based capture index (CDX) server supporting incremental updates and compression. Can be used as backend for OpenWayback, PyWb and <a href="https://github.com/ukwa/ukwa-heritrix/blob/master/src/main/java/uk/bl/wap/modules/uriuniqfilters/OutbackCDXRecentlySeenUriUniqFilter.java" rel="noopener noreferrer">Heritrix (⭐11)</a>. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/unt-libraries/py-wasapi-client" rel="noopener noreferrer">py-wasapi-client (⭐16)</a> - Command line application to download crawls from WASAPI (Python). <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/httpreserve/tikalinkextract" rel="noopener noreferrer">tikalinkextract (⭐11)</a> - Extract hyperlinks as a seed for web archiving from folders of document types that can be parsed by Apache Tika (Golang, Apache Tika Server). <em>(In Development)</em></li>
</ul>

<ul>
<li><a href="https://github.com/sul-dlss/wasapi-downloader" rel="noopener noreferrer">wasapi-downloader (⭐7)</a> - Java command line application to download crawls from WASAPI. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/helgeho/WarcPartitioner" rel="noopener noreferrer">WarcPartitioner (⭐1)</a> - Partition (W)ARC Files by MIME Type and Year. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/WikiTeam/wikiteam" rel="noopener noreferrer">wikiteam (⭐859)</a> - Tools for downloading and preserving wikis. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / WARC I/O Libraries</p>
</h3>
<ul>
<li><a href="https://github.com/helgeho/HadoopConcatGz" rel="noopener noreferrer">HadoopConcatGz (⭐9)</a> - A Splitable Hadoop InputFormat for Concatenated GZIP Files (and <code>*.warc.gz</code>). <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/N0taN3rd/node-warc" rel="noopener noreferrer">node-warc (⭐105)</a> - Parse WARC files or create WARC files using either <a href="https://electron.atom.io/" rel="noopener noreferrer">Electron</a> or <a href="https://github.com/cyrus-and/chrome-remote-interface" rel="noopener noreferrer">chrome-remote-interface (⭐4.6k)</a> (Node.js). <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/chfoo/warcat" rel="noopener noreferrer">Warcat (⭐166)</a> - Tool and library for handling Web ARChive (WARC) files (Python). <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Analysis</p>
</h3>
<ul>
<li><a href="https://github.com/helgeho/ArchiveSpark" rel="noopener noreferrer">ArchiveSpark (⭐162)</a> - An Apache Spark framework (not only) for Web Archives that enables easy data processing, extraction as well as derivation. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/archivesunleashed/notebooks" rel="noopener noreferrer">Archives Unleashed Notebooks (⭐26)</a> - Notebooks for working with web archives with the Archives Unleashed Toolkit, and derivatives generated by the Archives Unleashed Toolkit. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/archivesunleashed/aut" rel="noopener noreferrer">Archives Unleashed Toolkit (⭐158)</a> - An open-source platform for analyzing web archives with Apache Spark. <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/archivesunleashed/twut" rel="noopener noreferrer">Tweet Archvies Unleashed Toolkit (⭐10)</a> - An open-source toolkit for analyzing line-oriented JSON Twitter archives with Apache Spark. <em>(In Development)</em></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/03/26/"/>
    <summary>36 awesome projects updated on Mar 26, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/03/10/</id>
    <title>Awesome Web Archiving Updates on Mar 10, 2020</title>
    <updated>2020-03-10T16:33:17.000Z</updated>
    <published>2020-03-10T16:33:17.000Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Blogs and Scholarship</p>
</h3>
<ul>
<li><a href="https://blog.dshr.org/" rel="noopener noreferrer">DSHR's Blog</a> - David Rosenthal regularly reviews and summarizes work done in the Digital Preservation field.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/03/10/"/>
    <summary>1 awesome projects updated on Mar 10, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/03/03/</id>
    <title>Awesome Web Archiving Updates on Mar 03, 2020</title>
    <updated>2020-03-03T20:19:44.000Z</updated>
    <published>2020-03-03T20:19:44.000Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Blogs and Scholarship</p>
</h3>
<ul>
<li><a href="https://webarchivingrt.wordpress.com/" rel="noopener noreferrer">Web Archiving Roundtable</a> - Unofficial blog of the Web Archiving Roundtable of the <a href="https://www2.archivists.org/" rel="noopener noreferrer">Society of American Archivists</a> maintained by the members of the Web Archiving Roundtable.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/03/03/"/>
    <summary>1 awesome projects updated on Mar 03, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/02/27/</id>
    <title>Awesome Web Archiving Updates on Feb 27, 2020</title>
    <updated>2020-02-27T02:31:21.000Z</updated>
    <published>2020-02-27T02:31:21.000Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Twitter</p>
</h3>
<ul>
<li><a href="https://twitter.com/NetPreserve" rel="noopener noreferrer">@NetPreserve</a> - Official IIPC handle.</li>
</ul>

<ul>
<li><a href="https://twitter.com/search?q=%23webarchiving" rel="noopener noreferrer">#WebArchiving</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/02/27/"/>
    <summary>2 awesome projects updated on Feb 27, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/02/26/</id>
    <title>Awesome Web Archiving Updates on Feb 26, 2020</title>
    <updated>2020-02-26T16:49:53.000Z</updated>
    <published>2020-02-26T16:49:53.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Replay</p>
</h3>
<ul>
<li><a href="https://github.com/oduwsdl/ipwb" rel="noopener noreferrer">InterPlanetary Wayback (ipwb) (⭐655)</a> - Web Archive (WARC) indexing and replay using <a href="https://ipfs.io/" rel="noopener noreferrer">IPFS</a>.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/02/26/"/>
    <summary>1 awesome projects updated on Feb 26, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2020/02/25/</id>
    <title>Awesome Web Archiving Updates on Feb 25, 2020</title>
    <updated>2020-02-25T18:50:06.000Z</updated>
    <published>2020-02-25T14:47:58.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Replay</p>
</h3>
<ul>
<li><a href="https://oduwsdl.github.io/Reconstructive/" rel="noopener noreferrer">Reconstructive</a> - A Service Worker module for client-side reconstruction of composite mementos by rerouting resource requests to corresponding archived copies (JavaScript).</li>
</ul>
<h3><p>Tools &amp; Software / Search &amp; Discovery</p>
</h3>
<ul>
<li><a href="https://github.com/ukwa/webarchive-discovery" rel="noopener noreferrer">webarchive-discovery (⭐133)</a> - WARC and ARC full-text indexing and discovery tools, with a number of associated tools capable of using the index shown below. <em>(Stable)</em></li>
</ul>
<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/recrm/ArchiveTools" rel="noopener noreferrer">ArchiveTools (⭐81)</a> - Collection of tools to extract and interact with WARC files (Python).</li>
</ul>

<ul>
<li><a href="https://github.com/webrecorder/har2warc" rel="noopener noreferrer">har2warc (⭐55)</a> - Convert HTTP Archive (HAR) -&gt; Web Archive (WARC) format (Python).</li>
</ul>
<h3><p>Tools &amp; Software / WARC I/O Libraries</p>
</h3>
<ul>
<li><a href="https://github.com/iipc/jwarc" rel="noopener noreferrer">jwarc (⭐60)</a> - Read and write WARC files with a type safe API (Java).</li>
</ul>

<ul>
<li><a href="https://github.com/webrecorder/warcio" rel="noopener noreferrer">warcio (⭐468)</a> - Streaming WARC/ARC library for fast web archive IO (Python). <em>(Stable)</em></li>
</ul>

<ul>
<li><a href="https://github.com/internetarchive/warctools" rel="noopener noreferrer">warctools (⭐176)</a> - Library to work with ARC and WARC files (Python).</li>
</ul>

<ul>
<li><a href="https://github.com/richardlehane/webarchive" rel="noopener noreferrer">webarchive (⭐20)</a> - Golang readers for ARC and WARC webarchive formats (Golang).</li>
</ul>
<h3><p>Tools &amp; Software / Quality Assurance</p>
</h3>
<ul>
<li><a href="https://www.playonlinux.com/en/" rel="noopener noreferrer">PlayOnLinux</a> - For running Xenu and Notepad++ on Ubuntu.</li>
</ul>

<ul>
<li><a href="https://www.playonmac.com/en/" rel="noopener noreferrer">PlayOnMac</a> - For running Xenu and Notepad++ on macOS.</li>
</ul>

<ul>
<li><a href="https://support.microsoft.com/en-gb/help/13776/windows-use-snipping-tool-to-capture-screenshots" rel="noopener noreferrer">Windows Snipping Tool</a> - Windows built-in for partial screen capture and annotation. On macOS you can use Command + Shift + 4 (keyboard shortcut for taking partial screen capture).</li>
</ul>

<ul>
<li><a href="http://winebottler.kronenberg.org/" rel="noopener noreferrer">WineBottler</a> - For running Xenu and Notepad++ on macOS.</li>
</ul>

<ul>
<li><a href="https://github.com/jordansissel/xdotool" rel="noopener noreferrer">xDoTool (⭐3.8k)</a> - Click automation on Ubuntu.</li>
</ul>

<ul>
<li><a href="http://home.snafu.de/tilman/xenulink.html" rel="noopener noreferrer">Xenu</a> - Desktop link checker for Windows.</li>
</ul>
<h3><p>Community Resources / Slack</p>
</h3>
<ul>
<li><a href="https://archivesunleashed.slack.com/" rel="noopener noreferrer">Archives Unleashed Slack</a> - <a href="http://slack.archivesunleashed.org/" rel="noopener noreferrer">Fill out this request form</a> for access to a researcher group of people working with web archives.</li>
</ul>

<ul>
<li><a href="https://archivers.slack.com" rel="noopener noreferrer">Archivers Slack</a> - <a href="https://archivers-slack.herokuapp.com/" rel="noopener noreferrer">Invite yourself</a> to a multi-disciplinary effort for archiving projects run in affiliation with <a href="https://envirodatagov.org/archiving/" rel="noopener noreferrer">EDGI</a> and <a href="http://datatogether.org/" rel="noopener noreferrer">Data Together</a>.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2020/02/25/"/>
    <summary>16 awesome projects updated on Feb 25, 2020</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2018/10/16/</id>
    <title>Awesome Web Archiving Updates on Oct 16, 2018</title>
    <updated>2018-10-16T11:27:37.000Z</updated>
    <published>2018-10-16T11:27:37.000Z</published>
    <content type="html"><![CDATA[<h3><p>Resources for Web Publishers / For Researchers using Web Archives</p>
</h3>
<ul>
<li>The <a href="http://archiveready.com/" rel="noopener noreferrer">Archive Ready</a> tool, for estimating how likely a web page will be archived successfully.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2018/10/16/"/>
    <summary>1 awesome projects updated on Oct 16, 2018</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2018/03/14/</id>
    <title>Awesome Web Archiving Updates on Mar 14, 2018</title>
    <updated>2018-03-14T12:51:20.000Z</updated>
    <published>2018-03-14T12:51:20.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Search &amp; Discovery</p>
</h3>
<ul>
<li><a href="https://securitytrails.com/" rel="noopener noreferrer">SecurityTrails</a> - Web based archive for WHOIS and DNS records. REST API available free of charge.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2018/03/14/"/>
    <summary>1 awesome projects updated on Mar 14, 2018</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2017/06/26/</id>
    <title>Awesome Web Archiving Updates on Jun 26, 2017</title>
    <updated>2017-06-26T20:38:34.000Z</updated>
    <published>2017-06-26T20:38:34.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / For Researchers using Web Archives</p>
</h3>
<ul>
<li><a href="https://github.com/archivers-space/research/tree/master/web_archiving" rel="noopener noreferrer">Comparison of web archiving software (⭐100)</a></li>
</ul>

<ul>
<li><a href="https://github.com/edgi-govdata-archiving/awesome-website-change-monitoring" rel="noopener noreferrer">Awesome Website Change Monitoring (⭐513)</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2017/06/26/"/>
    <summary>2 awesome projects updated on Jun 26, 2017</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2017/06/21/</id>
    <title>Awesome Web Archiving Updates on Jun 21, 2017</title>
    <updated>2017-06-21T20:01:56.000Z</updated>
    <published>2017-06-21T20:01:56.000Z</published>
    <content type="html"><![CDATA[<h3><p>Tools &amp; Software / Utilities</p>
</h3>
<ul>
<li><a href="https://github.com/ikreymer/webarchive-indexing" rel="noopener noreferrer">webarchive-indexing (⭐46)</a> - Tools for bulk indexing of WARC/ARC files on Hadoop, EMR or local file system.</li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2017/06/21/"/>
    <summary>1 awesome projects updated on Jun 21, 2017</summary>
  </entry>
  <entry>
    <id>https://www.trackawesomelist.com/2017/06/16/</id>
    <title>Awesome Web Archiving Updates on Jun 16, 2017</title>
    <updated>2017-06-16T16:35:14.000Z</updated>
    <published>2017-06-16T14:04:02.000Z</published>
    <content type="html"><![CDATA[<h3><p>Community Resources / Blogs and Scholarship</p>
</h3>
<ul>
<li><a href="https://netpreserveblog.wordpress.com/" rel="noopener noreferrer">IIPC Blog</a></li>
</ul>
<h3><p>Community Resources / Mailing Lists</p>
</h3>
<ul>
<li><a href="http://netpreserve.org/about-us/iipc-mailing-list/" rel="noopener noreferrer">IIPC</a></li>
</ul>
]]></content>
    <link rel="alternate" href="https://www.trackawesomelist.com/2017/06/16/"/>
    <summary>2 awesome projects updated on Jun 16, 2017</summary>
  </entry>
</feed>