How to write a scraping client for Eternaltwin
⚠ There is no scraping client in this repository any more. Release 0.13.0
(December 2023) deleted crates/dinoparc_client, crates/dinorpg_client,
crates/hammerfest_client, crates/popotamo_client and
crates/twinoid_client. Nothing in the tree implements HammerfestClient,
DinoparcClient or DinorpgClient today, and scraper itself is used by a
single file: crates/scraper_tools/src/lib.rs.
What is still here, and what this page is really about:
- the client traits and the scraped data types, in
crates/core(crates/core/src/hammerfest/,crates/core/src/dinoparc.rs,crates/core/src/dinorpg/,crates/core/src/twinoid/,crates/core/src/popotamo.rs); - the archive stores that persist what a client returns
(
crates/dinoparc_store,crates/hammerfest_store,crates/twinoid_store) — see Archive; - the shared scraping helpers of
crates/scraper_tools; - the DNS overrides of
crates/mt_dns, which point the well-known Motion Twin domains at the right addresses; - the captured pages under
test-resources/scraping; - the jobs that drive a client, in
crates/services/src/job/(scrape_all_hammerfest_profiles,scrape_hammerfest_theme,scrape_hammerfest_thread,scrape_hammerfest_thread_list,archive_twinoid_users,reliable_hammerfest_acquire_session).
Those jobs are generic over a context that supplies a client
(HammerfestClientRef, TwinoidClientRef). Because no such client ships here,
JobRuntimeExtra in crates/system/src/lib.rs carries none, and the matching
job_runtime.register::<…>() calls are commented out: the scraping jobs are
currently registered nowhere.
So treat what follows as a design guide, not as a walkthrough you can copy from a neighbouring crate. The code samples describe the shape a client had; they are illustrations, not extracts that still compile.
Setup a dev environment
Scraping clients are written in Rust. Follow Rust for the
installation: the workspace requires Rust 1.89.0 or above (see rust-version
in the root Cargo.toml) and uses the 2024 edition.
Capture HTML example files and define the mapping
A scraper is tested against pages captured from the real server, stored under
test-resources/scraping.
There is one directory per site (dinoparc, dinorpg, hammerfest,
popotamo, twinoid), then one directory per kind of page, then one directory
per captured case:
test-resources/scraping/
hammerfest/
shop/
en-user158159/
input.html -- the captured page
options.json -- extra scraper input (`null` when there is none)
expected.json -- the value the scraper must produce
fr-user778923/
...
The file names are not uniform across sites, because each client picked its own when it was written:
hammerfestusesinput.html+options.json+expected.json;dinorpgusesinput.html+value.json;dinoparckeeps the raw capture asmain.htmland its transcoded copy asmain.utf8.html, and readsvalue.json; a handful of directories hold only a capture, with no expected value recorded yet;popotamousesmain.html+value.json;twinoidis mostly an API, so most of its cases areinput.json+value.json, with a handful ofinput.htmlpages.
A new site should pick one of these conventions and stay consistent inside its own directory. The expected value is the JSON serialisation of the response struct the scraper returns, for example:
{
"user": {
"id": "476256",
"username": "Evian"
},
"cities": [
{
"id": "1",
"name": "Creux infernal",
"nb_survived_days": 0
},
{
"id": "2",
"name": "Trou des cousins déformés",
"nb_survived_days": 59
}
]
}
Create the client crate
Crates are discovered by the crates/* glob in the members list of the root
Cargo.toml, so creating the directory is enough — there is no list of crates
to extend. Only the shared dependency versions live in the root
[workspace.dependencies] table.
cd crates cargo new --lib hordes_client
Give the crate the package name eternaltwin_hordes_client and inherit the
workspace settings, the way every other crate does:
[package]
name = "eternaltwin_hordes_client"
version = "0.16.4"
license.workspace = true
edition.workspace = true
rust-version.workspace = true
[dependencies]
scraper = { workspace = true }
thiserror = { workspace = true }
eternaltwin_scraper_tools = { workspace = true }
Client files
The deleted clients all shared this layout:
crates/hordes_client/
Cargo.toml -- client dependencies
README.md -- client documentation
src/
lib.rs -- feature-gated re-exports
mem.rs -- in-memory implementation, for tests
http/
errors.rs -- all errors the client can return
mod.rs -- high-level client (session handling, requests)
scraper.rs -- the functions that turn HTML into data
url.rs -- builders for the URLs to fetch
errors.rs
An enum of every error the client can return: a div with the expected id
was missing, a value did not parse, a session expired, and so on.
use thiserror::Error;
#[derive(Debug, Error)]
pub enum ScraperError {
#[error("DivWithIdNotFound: {0}")]
DivWithIdNotFound(String),
#[error("InvalidHordesUserCityId: {0}")]
InvalidHordesUserCityId(String),
}
mod.rs
The high-level client:
- a resolver that finds the address of the server to query (see
crates/mt_dns); - a method that opens a session with game or Twinoid credentials;
- one method per structure of interest, each fetching a page and handing it to
the matching function from
scraper.rs.
scraper.rs
The functions that turn HTML into data. They use the
scraper crate (version 0.24.0,
pinned in [workspace.dependencies]), whose Selector type runs CSS selectors
over a parsed Html document.
Parsing a selector is not free, so do not call Selector::parse inside a loop.
crates/scraper_tools exports a selector! macro that parses each literal once
and caches it in a OnceLock:
use eternaltwin_scraper_tools::{ElementRefExt, selector};
use scraper::{Html, Selector};
pub(crate) fn scrape_user_cities(doc: &Html) -> Result<Vec<HordesUserCity>, ScraperError> {
let root = doc.root_element();
let cities = root
.select(selector!("div#cities"))
.next()
.ok_or_else(|| ScraperError::DivWithIdNotFound("cities".to_string()))?;
let mut user_cities = Vec::new();
for city in cities.select(selector!("div.city")) {
let id: u32 = city
.value()
.attr("data-city-id")
.ok_or_else(|| ScraperError::DivWithIdNotFound("data-city-id".to_string()))?
.parse()
.map_err(|_| ScraperError::InvalidHordesUserCityId("data-city-id".to_string()))?;
let name = city
.select(selector!("div.city-name"))
.next()
.ok_or_else(|| ScraperError::DivWithIdNotFound("city-name".to_string()))?
.get_one_text()
.map_err(|_| ScraperError::DivWithIdNotFound("city-name".to_string()))?;
let name = HordesUserCityName::from_str(name)?;
user_cities.push(HordesUserCity { id, name });
}
Ok(user_cities)
}
crates/scraper_tools is small and worth reading in full. Besides selector!
it exports:
get_one_text/get_opt_text, and theElementRefExttrait that hangs them offElementRef: they extract the single text node of an element and fail loudly when there is more than one, instead of silently concatenating;FlashVars, an iterator over thea=1&b=2payload found in theflashvarsattribute of the old Flash embeds.
url.rs
An enum of the base URLs of the site, plus builders that add sub-paths and query parameters.
lib.rs
Feature-gated re-exports, so that a consumer can pull in only the implementation it needs:
#[cfg(feature = "http")] pub mod http; #[cfg(feature = "mem")] pub mod mem;
mem.rs
An in-memory implementation of the same client trait, so that services and tests can run without touching the network.
1. Defining the data structures
The scraped types belong to crates/core, not to the client: the stores, the
services and the REST layer all need them, and the client is only one producer.
Add a module under
crates/core/src
— a single file for a small site (popotamo.rs, dinoparc.rs), a directory
for a large one (dinorpg/, hammerfest/, twinoid/).
#[derive(Debug, Clone, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub struct HordesUserCity {
pub id: u32,
pub name: HordesUserCityName,
pub level: u32,
pub nb_survived_days: u32,
}
Scraped strings get their own validated newtype rather than a bare String.
The declare_new_string! macro (crates/core/src/types.rs) generates the
parser, the error type and the SQL mapping:
declare_new_string! {
pub struct HordesUserCityName(String);
pub type ParseError = HordesUserCityNameParseError;
const PATTERN = r"^[a-zA-Z]{1,32}$";
const SQL_NAME = "hordes_user_city_name";
}
crates/core/src/types.rs also provides declare_new_int!,
declare_new_uuid! and declare_new_enum!. The latter maps each variant to
the exact string found on the page:
declare_new_enum!(
pub enum HordesUserCityDeathCause {
#[str("Disparu dans l'Outre-Monde pendant la nuit !")]
LostInTheOuterMonde,
#[str("Lacéré, dévoré... pendant l'attaque de la nuit !")]
DeadAtTown,
}
pub type ParseError = HordesUserCityDeathCauseParseError;
const SQL_NAME = "hordes_user_city_death_cause";
);
The same file is where the response structs live:
#[cfg_attr(feature = "serde", derive(Serialize, Deserialize))]
#[derive(Clone, Debug, PartialEq, Eq)]
pub struct HordesUserCitiesResponse {
pub profile: HordesUserProfile,
pub cities: Option<Vec<HordesUserCity>>,
}
crates/core/src/dinorpg
is the most complete example: it splits server enums, session keys, per-entity
modules and the client trait across mod.rs, client.rs and one file per
domain concept.
2. Testing the scraper
Each scraper function gets a test that reads a captured page, runs the function, and compares the result with the recorded expectation.
⚠ The old tests used the test_resources attribute from the
test-generator crate to generate one test per directory. That dependency was
removed along with the clients and is no longer in Cargo.lock; a new client
has to either re-introduce it deliberately or iterate over the directories by
hand:
#[test]
fn test_scrape_user_cities() {
for entry in std::fs::read_dir("../../test-resources/scraping/hordes/user").unwrap() {
let path = entry.unwrap().path();
let raw_html = std::fs::read_to_string(path.join("input.html")).unwrap();
let doc = Html::parse_document(&raw_html);
let actual = scrape_user_cities(&doc).unwrap();
// Written back so a failing run leaves a diffable artefact next to the input.
std::fs::write(
path.join("rs.actual.json"),
format!("{}\n", serde_json::to_string_pretty(&actual).unwrap()),
)
.unwrap();
let expected: HordesUserCitiesResponse =
serde_json::from_str(&std::fs::read_to_string(path.join("value.json")).unwrap()).unwrap();
assert_eq!(actual, expected);
}
}
Note the relative path: tests run with the crate directory as the working
directory, so test-resources is reached through ../...
Run the tests with cargo test -p eternaltwin_hordes_client.
3. Reaching the servers
Most Motion Twin domains no longer resolve to a server we can query, so
Eternaltwin ships its own address table in
crates/mt_dns.
MtDnsResolver answers from dead.rs first (domains whose public DNS record is
gone) and falls back to live.rs.
⚠ crates/mt_dns/src/dead.rs and crates/mt_dns/src/live.rs are generated —
do not edit them by hand. Their header says so, and cargo run -p xtask -- dns
rewrites them from the four text files in the
dns/ directory:
dns/live-domains.txtanddns/dead-domains.txt: the domains to resolve;dns/live-records.txtanddns/dead-records.txt: the recorded answers.
Adding a site therefore means:
declaring the server enum in
crates/core, next to the other types of the site, with each variant mapped to its host name:declare_new_enum!( pub enum HordesServer { #[str("hordes.fr")] HordesFr, #[str("die2nite.com")] Die2NiteCom, #[str("www.zombinoia.com")] ZombinoiaCom, } pub type ParseError = HordesServerParseError; const SQL_NAME = "hordes_server"; );adding the domains to
dns/live-domains.txt(ordns/dead-domains.txt) and re-runningcargo run -p xtask -- dnsto regenerate the records;forwarding the new type in
crates/mt_dns/src/lib.rs, which is the only hand-written file of the crate. It already does this forstr,HammerfestServerandDinorpgServer:impl DnsResolver<HordesServer> for MtDnsResolver { fn resolve4(&self, domain: &HordesServer) -> Option<Ipv4Addr> { dead::DnsClient .resolve4(domain) .or_else(|| live::DnsClient.resolve4(domain)) } fn resolve6(&self, domain: &HordesServer) -> Option<Ipv6Addr> { dead::DnsClient .resolve6(domain) .or_else(|| live::DnsClient.resolve6(domain)) } }SystemDnsResolver, in the same file, is the opt-out: it resolves nothing and lets the operating system decide.
4. Wiring the client into the server
A client is useless on its own. To make it reachable from the running server:
- store what it returns. The archive stores keep the full history of every response instead of just the latest state; the model and the upsertion query are described in Archive.
- drive it from a job in
crates/services/src/job/, followingscrape_all_hammerfest_profiles.rsorscrape_hammerfest_thread.rs. A job is generic over its context and requires the client through a…ClientRefbound. - add the client to
JobRuntimeExtraincrates/system/src/lib.rsand register the job there. This is the step that is currently missing for every scraping job in the tree. - expose the archived data through
crates/rest/src/archive/, which is mounted at/api/v1/archive.
The eternaltwin binary (bin/, built from crates/cli) has no per-game demo
subcommand. Its only archive-related command is eternaltwin tidsave, which
runs the ArchiveTwinoidUsersSingleToken job over a range of Twinoid user ids
using one access token.