How do you test engines scientifically?
Scientific engine testing puts the engines on the same machine, uses a single testing tool, keeps the test conditions and settings of both sides identical, and uses a sufficient number of varied opening positions. It makes sure moves come purely from engine calculation (some engines may have a built-in opening book for mainstream positions), plays every position as a paired game, preferably without pondering, and tries to exclude possible interference from background programs, as well as game records with rules disagreements between engines or between the GUI and the engine.
It is also worth occasionally looking at the game records, which can reveal problems such as wrong rules settings, wrong time settings, or a limited-strength setting being enabled.
It is best not to use hyper-threading, unless you run only one game at a time rather than several in parallel. Also rule out possible multi-socket CPU scheduling issues.
The number of games must also be large enough to avoid error, for example several thousand, preferably with statistical tools. If the gap is very small, at least tens of thousands of games are needed. (Suppose a beats b in 10% of games and the rest are draws. The probability of 10 draws in a row is still over one third, so you cannot conclude they are equal from those 10 games.)
In addition, it is recommended to test with an increment time control (time plus increment per move). The engine allocates thinking time itself according to position complexity and other information, which reduces unnecessary thinking and also gives stronger play for the same time used, unless the engine's time management is very poor.
The purpose of testing is to magnify the strength differences between engines. So if the engines are close in strength, paired games from unbalanced positions are generally used, since they magnify the strength difference better and reduce the time needed to overcome error.
