Command Palette

Search for a command to run...

Evaluation Protocol

Reference hardware for problem runs, how native holdout metrics become Arena scores, and the metric catalog on each problem page.

Reference hardware

Rescale core type
grossular-1
Allocation
1 node
GPU
1× NVIDIA A10G (24 GB)
CPU
4 vCPU · 2nd gen AMD EPYC @ 2.8 GHz
RAM
32 GB / node
Storage
450 GB / node
Max wall time
48 hours (job terminated if exceeded)

Plate and airfoil reference runs typically finish in tens of minutes on grossular-1. On the OpenRadioss Car Rescale Tutorial pack, reference runs were about ~2 hours for GeoTransolver and ~12 hours for MeshGraphNet. Other crash packs or architectures may differ. DrivAerML and HiLiftAeroML native-surface defaults training is multi-hour on the same class of node (architecture-dependent). Training and inference times on published submissions are informational only.

How Arena scores are built

The score arena maps each submission's native holdout metrics into a shared 0–1 playground: Scalars, Fields, and Explore. Recipes differ by problem; pin a run and open “How scored?” for the worked math. Absolute errors stay in native units on the problem page.

Holdout inference

Each submission is evaluated on hidden test inputs. Ground truth is never exposed during inference. Predictions are compared to simulation reference values after the run completes.

Scalars (global accuracy)

For decision-driving quantities (Cl, Cd, peak stress, Kt, etc.), each holdout case gets an accuracy score: max(0, 1 − |predicted − truth| / |truth|). Scalars is the mean of those accuracies across global metrics and holdout cases. NAFEMS plate packs can use peak S1 accuracy (and Kt when present). Crash packs are usually scalar-empty.

Fields (nodal quality)

Mesh field outputs (pressure, stress, displacement, mean surface Cp, wall shear, etc.) are scored with R² on nodal values. Negative R² is clamped to zero before aggregation. Fields is usually the mean of those per-case R² values. Transient crash displacement first takes median R² across timesteps within each case. NAFEMS Fields also fold in a relative MAE score vs peak scale.

Explore (combined)

When both Scalars and Fields exist, Explore is the geometric mean √(Scalars × Fields). When only one side is available (nodal-only or scalar-only), Explore is that score alone. Explore is a playground blend, not a crown ranking.

Problem rooms

Arena rooms group submissions by problem (Plate with Hole, 2D Airfoil, DrivAerML, OpenRadioss Car, HiLiftAeroML, and so on). Compare within a room; recipes and native units are not interchangeable across problems.

What is not scored

Training time, inference time, and hardware core type are recorded for context but do not change Scalars, Fields, or Explore. See each problem’s submissions table for per-run native breakdowns.

Open the score arena
Explore formula

For each model on a problem with both scalars and fields:

Explore = √(Scalars × Fields)

Both inputs are on a 0–1 scale. Example: 0.99 Scalars and 0.97 Fields → Explore ≈ 0.98. When only Scalars or only Fields exist, Explore equals that score alone.

Native metrics on problem pages

Problem pages keep the full predicted-vs-actual catalog from the eval pipeline. Some problems emphasize global scalars (Cl, Cd, Cm, Kt). Others are nodal-only (crash displacement). Plate with Hole leads with an S1 peak-error suite. Sort and expand rows to see what fits your question.

Global metrics

One number per holdout case for engineering quantities that drive decisions.

  • Accuracy %: max(0, 1 − |predicted − truth| / |truth|), shown as a percentage. Used for Cl, Cd, Cm, Kt, and similar scalars.
  • Relative error %: Same basis as accuracy, reported as percent error vs simulation truth. Useful when you want the raw gap, not the inverted score.

Plate with Hole (S1 suite)

Across Rescale Tutorial and NAFEMS D1–D6, published plate scores center on max principal stress (S1) and, where applicable, Kt. Headline field metrics:

  • AEinPEAK: |max(truth) − max(pred)| on the mesh.
  • AE@PEAK: absolute error at the ground-truth peak node.
  • MAE / ME / RMSE: mean absolute, mean signed, and root-mean-square error over mesh points.
  • Field R²: coefficient of determination on nodal S1.
  • Kt accuracy: when a concentration-factor head is present (centered-hole packs). Not shown for off-center NAFEMS D5–D6.

Relative L1 and percent-error vs analytical (Heywood) are not used in the published plate tables.

Nodal metrics (other problems)

Scores on mesh field outputs: pressure, velocity, stress, displacement time series, mean surface Cp, wall shear, and similar.

  • Field R²: Coefficient of determination on nodal values. Negative R² from poor fits is clamped to zero before display.
  • Median field R² (crash): OpenRadioss Car displacement is a time series. Each holdout case uses median R² across timesteps.
  • Relative L1: mean(|error|) / mean(|truth|) on nodal fields, per field (airfoil / DrivAerML / HiLiftAeroML where shown).
  • RMSE: Root mean square error on nodal fields, per field.

OpenRadioss Car hides Rel L1 / RMSE in the UI because displacement is a time series.

Where to look

Use the score arena for cross-run Scalars / Fields / Explore, and each problem page for native units and per-case tables. Expand a submission row for the full metric set from the eval pipeline.