Is it reliable to judge which engine is stronger by "feel" without testing?
"I feel xx is not as good as yy", "I feel xx can't beat yy so their strength is about the same", "xx can't solve such a simple position so xx is bad", "xx gives such a high score to a drawn position so xx is bad", "xx calculates this position slower than yy so xx is worse than yy", "xx's scores are too floaty and inflated, so it's worse than yy, whose scores are steadier"... all such claims are extremely one-sided.
Human "feel" is wildly unreliable, and some people can reach all kinds of absurd conclusions from feel. Only when two engines differ enormously can the difference be felt easily.
An engine's strength cannot be measured by its performance in individual positions alone. Every engine has blind spots, and different engines' blind spots do not necessarily coincide. Even if one engine cannot solve some positions that another engine can, the sample is far too small to conclude that it is weaker than the other.
Draws are perfectly normal, so more statistics are needed, and the opening book must be taken into account. Even testing requires a large sample; a test of a few games or a few dozen games can "prove" any so-called "conclusion".
Engines are not gods and have many imperfections. By analogy, an engine is an off-road vehicle: in the vast majority of scenarios it is superior to a human, but faced with a wall that must be climbed, it is worse than a human. In this example, the "wall" corresponds to certain composed endgame puzzles, but that does not mean the engine's playing strength is poor.
Fluctuations in an engine's score say even less about strength. In fact, if you multiply an engine's raw scores by a coefficient (say 3), many people will feel the engine has become weaker and its scores too floaty and inflated, while its strength is exactly the same.
