Eternaltwin

Home

How to write a scraping client for Eternaltwin

⚠ There is no scraping client in this repository any more. Release 0.13.0 (December 2023) deleted crates/dinoparc_client, crates/dinorpg_client, crates/hammerfest_client, crates/popotamo_client and crates/twinoid_client. Nothing in the tree implements HammerfestClient, DinoparcClient or DinorpgClient today, and scraper itself is used by a single file: crates/scraper_tools/src/lib.rs.

What is still here, and what this page is really about:

  • the client traits and the scraped data types, in crates/core (crates/core/src/hammerfest/, crates/core/src/dinoparc.rs, crates/core/src/dinorpg/, crates/core/src/twinoid/, crates/core/src/popotamo.rs);
  • the archive stores that persist what a client returns (crates/dinoparc_store, crates/hammerfest_store, crates/twinoid_store) — see Archive;
  • the shared scraping helpers of crates/scraper_tools;
  • the DNS overrides of crates/mt_dns, which point the well-known Motion Twin domains at the right addresses;
  • the captured pages under test-resources/scraping;
  • the jobs that drive a client, in crates/services/src/job/ (scrape_all_hammerfest_profiles, scrape_hammerfest_theme, scrape_hammerfest_thread, scrape_hammerfest_thread_list, archive_twinoid_users, reliable_hammerfest_acquire_session).

Those jobs are generic over a context that supplies a client (HammerfestClientRef, TwinoidClientRef). Because no such client ships here, JobRuntimeExtra in crates/system/src/lib.rs carries none, and the matching job_runtime.register::<…>() calls are commented out: the scraping jobs are currently registered nowhere.

So treat what follows as a design guide, not as a walkthrough you can copy from a neighbouring crate. The code samples describe the shape a client had; they are illustrations, not extracts that still compile.

Setup a dev environment

Scraping clients are written in Rust. Follow Rust for the installation: the workspace requires Rust 1.89.0 or above (see rust-version in the root Cargo.toml) and uses the 2024 edition.

Capture HTML example files and define the mapping

A scraper is tested against pages captured from the real server, stored under test-resources/scraping. There is one directory per site (dinoparc, dinorpg, hammerfest, popotamo, twinoid), then one directory per kind of page, then one directory per captured case:

test-resources/scraping/
    hammerfest/
        shop/
            en-user158159/
                input.html     -- the captured page
                options.json   -- extra scraper input (`null` when there is none)
                expected.json  -- the value the scraper must produce
            fr-user778923/
                ...

The file names are not uniform across sites, because each client picked its own when it was written:

  • hammerfest uses input.html + options.json + expected.json;
  • dinorpg uses input.html + value.json;
  • dinoparc keeps the raw capture as main.html and its transcoded copy as main.utf8.html, and reads value.json; a handful of directories hold only a capture, with no expected value recorded yet;
  • popotamo uses main.html + value.json;
  • twinoid is mostly an API, so most of its cases are input.json + value.json, with a handful of input.html pages.

A new site should pick one of these conventions and stay consistent inside its own directory. The expected value is the JSON serialisation of the response struct the scraper returns, for example:

{
  "user": {
    "id": "476256",
    "username": "Evian"
  },
  "cities": [
    {
      "id": "1",
      "name": "Creux infernal",
      "nb_survived_days": 0
    },
    {
      "id": "2",
      "name": "Trou des cousins déformés",
      "nb_survived_days": 59
    }
  ]
}

Create the client crate

Crates are discovered by the crates/* glob in the members list of the root Cargo.toml, so creating the directory is enough — there is no list of crates to extend. Only the shared dependency versions live in the root [workspace.dependencies] table.

cd crates
cargo new --lib hordes_client

Give the crate the package name eternaltwin_hordes_client and inherit the workspace settings, the way every other crate does:

[package]
name = "eternaltwin_hordes_client"
version = "0.16.4"
license.workspace = true
edition.workspace = true
rust-version.workspace = true

[dependencies]
scraper = { workspace = true }
thiserror = { workspace = true }
eternaltwin_scraper_tools = { workspace = true }

Client files

The deleted clients all shared this layout:

crates/hordes_client/
    Cargo.toml    -- client dependencies
    README.md     -- client documentation
    src/
        lib.rs    -- feature-gated re-exports
        mem.rs    -- in-memory implementation, for tests
        http/
            errors.rs   -- all errors the client can return
            mod.rs      -- high-level client (session handling, requests)
            scraper.rs  -- the functions that turn HTML into data
            url.rs      -- builders for the URLs to fetch

errors.rs

An enum of every error the client can return: a div with the expected id was missing, a value did not parse, a session expired, and so on.

use thiserror::Error;

#[derive(Debug, Error)]
pub enum ScraperError {
  #[error("DivWithIdNotFound: {0}")]
  DivWithIdNotFound(String),
  #[error("InvalidHordesUserCityId: {0}")]
  InvalidHordesUserCityId(String),
}

mod.rs

The high-level client:

  • a resolver that finds the address of the server to query (see crates/mt_dns);
  • a method that opens a session with game or Twinoid credentials;
  • one method per structure of interest, each fetching a page and handing it to the matching function from scraper.rs.

scraper.rs

The functions that turn HTML into data. They use the scraper crate (version 0.24.0, pinned in [workspace.dependencies]), whose Selector type runs CSS selectors over a parsed Html document.

Parsing a selector is not free, so do not call Selector::parse inside a loop. crates/scraper_tools exports a selector! macro that parses each literal once and caches it in a OnceLock:

use eternaltwin_scraper_tools::{ElementRefExt, selector};
use scraper::{Html, Selector};

pub(crate) fn scrape_user_cities(doc: &Html) -> Result<Vec<HordesUserCity>, ScraperError> {
  let root = doc.root_element();
  let cities = root
    .select(selector!("div#cities"))
    .next()
    .ok_or_else(|| ScraperError::DivWithIdNotFound("cities".to_string()))?;

  let mut user_cities = Vec::new();

  for city in cities.select(selector!("div.city")) {
    let id: u32 = city
      .value()
      .attr("data-city-id")
      .ok_or_else(|| ScraperError::DivWithIdNotFound("data-city-id".to_string()))?
      .parse()
      .map_err(|_| ScraperError::InvalidHordesUserCityId("data-city-id".to_string()))?;
    let name = city
      .select(selector!("div.city-name"))
      .next()
      .ok_or_else(|| ScraperError::DivWithIdNotFound("city-name".to_string()))?
      .get_one_text()
      .map_err(|_| ScraperError::DivWithIdNotFound("city-name".to_string()))?;
    let name = HordesUserCityName::from_str(name)?;

    user_cities.push(HordesUserCity { id, name });
  }

  Ok(user_cities)
}

crates/scraper_tools is small and worth reading in full. Besides selector! it exports:

  • get_one_text / get_opt_text, and the ElementRefExt trait that hangs them off ElementRef: they extract the single text node of an element and fail loudly when there is more than one, instead of silently concatenating;
  • FlashVars, an iterator over the a=1&b=2 payload found in the flashvars attribute of the old Flash embeds.

url.rs

An enum of the base URLs of the site, plus builders that add sub-paths and query parameters.

lib.rs

Feature-gated re-exports, so that a consumer can pull in only the implementation it needs:

#[cfg(feature = "http")]
pub mod http;
#[cfg(feature = "mem")]
pub mod mem;

mem.rs

An in-memory implementation of the same client trait, so that services and tests can run without touching the network.

1. Defining the data structures

The scraped types belong to crates/core, not to the client: the stores, the services and the REST layer all need them, and the client is only one producer. Add a module under crates/core/src — a single file for a small site (popotamo.rs, dinoparc.rs), a directory for a large one (dinorpg/, hammerfest/, twinoid/).

#[derive(Debug, Clone, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub struct HordesUserCity {
  pub id: u32,
  pub name: HordesUserCityName,
  pub level: u32,
  pub nb_survived_days: u32,
}

Scraped strings get their own validated newtype rather than a bare String. The declare_new_string! macro (crates/core/src/types.rs) generates the parser, the error type and the SQL mapping:

declare_new_string! {
  pub struct HordesUserCityName(String);
  pub type ParseError = HordesUserCityNameParseError;
  const PATTERN = r"^[a-zA-Z]{1,32}$";
  const SQL_NAME = "hordes_user_city_name";
}

crates/core/src/types.rs also provides declare_new_int!, declare_new_uuid! and declare_new_enum!. The latter maps each variant to the exact string found on the page:

declare_new_enum!(
  pub enum HordesUserCityDeathCause {
    #[str("Disparu dans l'Outre-Monde pendant la nuit !")]
    LostInTheOuterMonde,
    #[str("Lacéré, dévoré... pendant l'attaque de la nuit !")]
    DeadAtTown,
  }
  pub type ParseError = HordesUserCityDeathCauseParseError;
  const SQL_NAME = "hordes_user_city_death_cause";
);

The same file is where the response structs live:

#[cfg_attr(feature = "serde", derive(Serialize, Deserialize))]
#[derive(Clone, Debug, PartialEq, Eq)]
pub struct HordesUserCitiesResponse {
  pub profile: HordesUserProfile,
  pub cities: Option<Vec<HordesUserCity>>,
}

crates/core/src/dinorpg is the most complete example: it splits server enums, session keys, per-entity modules and the client trait across mod.rs, client.rs and one file per domain concept.

2. Testing the scraper

Each scraper function gets a test that reads a captured page, runs the function, and compares the result with the recorded expectation.

⚠ The old tests used the test_resources attribute from the test-generator crate to generate one test per directory. That dependency was removed along with the clients and is no longer in Cargo.lock; a new client has to either re-introduce it deliberately or iterate over the directories by hand:

#[test]
fn test_scrape_user_cities() {
  for entry in std::fs::read_dir("../../test-resources/scraping/hordes/user").unwrap() {
    let path = entry.unwrap().path();
    let raw_html = std::fs::read_to_string(path.join("input.html")).unwrap();
    let doc = Html::parse_document(&raw_html);

    let actual = scrape_user_cities(&doc).unwrap();
    // Written back so a failing run leaves a diffable artefact next to the input.
    std::fs::write(
      path.join("rs.actual.json"),
      format!("{}\n", serde_json::to_string_pretty(&actual).unwrap()),
    )
    .unwrap();

    let expected: HordesUserCitiesResponse =
      serde_json::from_str(&std::fs::read_to_string(path.join("value.json")).unwrap()).unwrap();
    assert_eq!(actual, expected);
  }
}

Note the relative path: tests run with the crate directory as the working directory, so test-resources is reached through ../...

Run the tests with cargo test -p eternaltwin_hordes_client.

3. Reaching the servers

Most Motion Twin domains no longer resolve to a server we can query, so Eternaltwin ships its own address table in crates/mt_dns. MtDnsResolver answers from dead.rs first (domains whose public DNS record is gone) and falls back to live.rs.

⚠ crates/mt_dns/src/dead.rs and crates/mt_dns/src/live.rs are generated — do not edit them by hand. Their header says so, and cargo run -p xtask -- dns rewrites them from the four text files in the dns/ directory:

  • dns/live-domains.txt and dns/dead-domains.txt: the domains to resolve;
  • dns/live-records.txt and dns/dead-records.txt: the recorded answers.

Adding a site therefore means:

  1. declaring the server enum in crates/core, next to the other types of the site, with each variant mapped to its host name:

    declare_new_enum!(
      pub enum HordesServer {
        #[str("hordes.fr")]
        HordesFr,
        #[str("die2nite.com")]
        Die2NiteCom,
        #[str("www.zombinoia.com")]
        ZombinoiaCom,
      }
      pub type ParseError = HordesServerParseError;
      const SQL_NAME = "hordes_server";
    );
    
  2. adding the domains to dns/live-domains.txt (or dns/dead-domains.txt) and re-running cargo run -p xtask -- dns to regenerate the records;

  3. forwarding the new type in crates/mt_dns/src/lib.rs, which is the only hand-written file of the crate. It already does this for str, HammerfestServer and DinorpgServer:

    impl DnsResolver<HordesServer> for MtDnsResolver {
      fn resolve4(&self, domain: &HordesServer) -> Option<Ipv4Addr> {
        dead::DnsClient
          .resolve4(domain)
          .or_else(|| live::DnsClient.resolve4(domain))
      }
    
      fn resolve6(&self, domain: &HordesServer) -> Option<Ipv6Addr> {
        dead::DnsClient
          .resolve6(domain)
          .or_else(|| live::DnsClient.resolve6(domain))
      }
    }
    

    SystemDnsResolver, in the same file, is the opt-out: it resolves nothing and lets the operating system decide.

4. Wiring the client into the server

A client is useless on its own. To make it reachable from the running server:

  • store what it returns. The archive stores keep the full history of every response instead of just the latest state; the model and the upsertion query are described in Archive.
  • drive it from a job in crates/services/src/job/, following scrape_all_hammerfest_profiles.rs or scrape_hammerfest_thread.rs. A job is generic over its context and requires the client through a …ClientRef bound.
  • add the client to JobRuntimeExtra in crates/system/src/lib.rs and register the job there. This is the step that is currently missing for every scraping job in the tree.
  • expose the archived data through crates/rest/src/archive/, which is mounted at /api/v1/archive.

The eternaltwin binary (bin/, built from crates/cli) has no per-game demo subcommand. Its only archive-related command is eternaltwin tidsave, which runs the ArchiveTwinoidUsersSingleToken job over a range of Twinoid user ids using one access token.