Why do engines give different scores, even different versions of the same engine?
Engine scoring has no standard today. Nobody has decided that a given position equals a fixed number of points, so scores of different engines and different versions cannot be compared with each other. A score can only be compared with scores from the same engine and version.
So it is perfectly normal for different versions and different engines to give different scores.
An engine's score has nothing to do with so-called "sensitivity" or "inflation". If you multiplied Pikafish's scores by 10 and called it a "new engine", would you think the new engine was more sensitive to scores or had inflated scores? Obviously the two engines are the same; they just look different on the surface. The only indicator of an engine's playing strength is scientific test data.
Win-rate model
There is, however, one unified standard: build a win-rate model from test data and convert the engine's raw output score into a win-rate score.
Currently Pikafish's win-rate score is tied to the Elo difference. For example, an advantage of 200 points represents a 76% win rate (the win rate commonly used in Xiangqi, meaning wins plus half the draws; for example 4 wins, 4 draws and 2 losses is a 60% win rate, and 3 wins, 4 draws and 3 losses is a 50% win rate).
However, the web version and other unofficially released Pikafish engines may not use this kind of score, and instead divide the raw score by a constant.


Pikafish's win-rate model is fitted to the actual win rates of engine self-play tests (equivalent to 1 thread at 60 seconds + 0.6 seconds).
By design, none of an engine's non-mate scores mean a forced win. These scores are just "evaluations": the engine's view of who is better in the current position, similar to a human thinking a position is easy to play, clearly better, or winning, except that the engine subdivides it into numeric scores.
