Eternaltwin

Home | Contributing

Archive

This document gives a high-level description of how Eternaltwin archives data from the Motion Twin websites.

The archive currently covers Hammerfest, Dinoparc and Twinoid. crates/core also carries types for Dinorpg and Popotamo, which have no store and no endpoint yet.

The model

For each website we archive, there is a client and a store.

The role of the client is to handle the communication with the website: convert queries to HTTP requests, send them, and process the response. Most Motion Twin websites only return the response as an HTTP page (no API), so the client has a component converting the HTML response to a more suitable format: the scraper.

The store saves the server responses and persists them in a database. It stores all the responses, so it creates a history of changes. The store is responsible for building the latest known state so it can be exported, ready to be consumed by any other website (usually one of our game websites).

Archive schema

What the repository contains today

The picture above still describes the design, but only part of it is implemented here.

The clients were removed.crates/dinoparc_client, crates/hammerfest_client, crates/dinorpg_client and crates/popotamo_client, each with its own http/scraper.rs, were deleted in 0.13.0. What is left of that half is:

  • the traits in crates/core: DinoparcClient (src/dinoparc.rs), HammerfestClient (src/hammerfest/mod.rs) and TwinoidClient (src/twinoid/client/). They define what a client must offer, and nothing in the repository implements them;
  • the jobs written against those traits, in crates/services/src/job/: archive_twinoid_users, reliable_hammerfest_acquire_session, scrape_all_hammerfest_profiles, scrape_hammerfest_theme, scrape_hammerfest_thread and scrape_hammerfest_thread_list;
  • crates/scraper_tools, helpers for the scraper HTML crate. It is still a workspace member, but no crate depends on it;
  • the fixtures under test-resources/scraping/, one directory per website, each pairing an input.html with the value.json a scraper should produce.

crates/system, which assembles the server from the configuration, builds no archive client, and the only job spawn in crates/cli/src/cmd/job.rs is commented out. In practice the archive is read-only today: it serves what was collected before, and nothing in this repository fetches more.

If you want to bring a client back, Scraping describes how one is written and how the fixtures are used.

The stores are current.crates/dinoparc_store, crates/hammerfest_store and crates/twinoid_store each implement the matching core trait twice, in mem.rs and pg.rs; twinoid_store also has sqlite.rs, and a trace.rs wrapper for telemetry. test.rs holds the suite every implementation must pass, so the memory and Postgres versions are held to the same behaviour.

The Postgres implementations are where the history actually lives. Responses are stored as snapshots with a validity period and a set of retrieval timestamps, inserted through the upsertion query built by the macro in crates/postgres_tools. Archive describes that model in detail: snapshots, point-in-time uniqueness, shards and pagination.

The read API is current.crates/rest/src/archive/ mounts three routers under /api/v1/archive:

EndpointReturns
/archive/dinoparc/:server/users/:user_idan archived Dinoparc user
/archive/dinoparc/:server/dinoz/:dinoz_idan archived dinoz
/archive/hammerfest/:server/users/:user_idan archived Hammerfest user
/archive/twinoid/users/:user_idan archived Twinoid user
/archive/twinoid/users/:user_id/scoresthat user's scores
/archive/twinoid/users/:user_id/rewards/:site_idthat user's rewards on one site

These are the export side of the diagram: this is how a game website reads what Eternaltwin has archived.

See also

  • Archive: the snapshot model and the SQL behind it.
  • Scraping: how to write a scraping client.
  • Overview: where each crate sits.