Position Evaluation and NNUE
Evaluation asks “How does this position look?” Search asks “What happens if both sides keep playing?” NNUE gives search a quick assessment. It does not directly select the best move or replace mate, repetition, and rule checks.
This page was checked against official Pikafish commit 1c66b9b. It explains how this code actually uses the network. The origins and size of the training data, or the training process, cannot be inferred from inference code alone.
From Board to Score: Each Part's Job
A normal static evaluation roughly follows these steps:
- Encode piece locations and some relationships between pieces as network features.
- Reuse and update accumulators, then compute the later network layers to obtain a raw NNUE score.
- Let
evaluate.cppadjust it using material, search preferences, move-limit counts, and other information. - Let search add evaluation corrections as needed and compare continuations to obtain a search result.
These scores are not interchangeable. The score shown by the GUI also goes through an output conversion; see Engine Runtime and Score Output. For how search uses evaluation, see Search Algorithms.
Sources: network evaluation, the evaluation entry point and scaling, evaluation corrections in search.
evaluate.cpp: More Than a Network Call
The current Eval::evaluate() calls NNUE, then applies scaling. It requires the position not to be in check. Search handles checked nodes through the appropriate paths; a normal static score cannot stand in for the result of searching evasions.
simple_eval() provides a rough material balance to compare with NNUE's assessment. It is not a fallback evaluator for a failed network load. The processing mainly does four things:
- Compare whether the material balance and NNUE score point in the same direction. For example, a position favored by the network despite a material deficit differs from a straightforward material advantage.
- Adjust the result using the
optimismpassed in by search. It comes from root-candidate scores, not a random bonus or the user's Elo rating. - Scale the score by the total material on the board and move it toward zero as the move-limit count rises.
- Keep ordinary evaluations outside the mate-score range, so “a large advantage” is not reported as “a mate has been found.”
The raw NNUE score can therefore differ from the static evaluation actually used in search. Scaling for the move limit only adjusts the estimate; rule code still decides whether a draw or another rule has actually been triggered.
Sources: material comparison and scaling, where optimism comes from.
What Does NNUE Look At?
This implementation uses two kinds of input together:
| Feature | How to understand it |
|---|---|
HalfKAv2_hm | Records which pieces occupy which squares, with feature groups selected using king locations, left-right mirroring, and the friendly chariot, horse, and cannon configuration. |
FullThreats | Extracts selected attack and defense relationships between pieces on the current board, adding information about how pieces affect one another. |
The first is not a fixed piece-square table: locations are encoded as network inputs, and the weights combine them into a result. Despite its name, FullThreats does not identify every tactical threat. It builds relationships from current attack ranges and occupied squares, filtering some combinations and duplicate relationships. It does not play out a sequence for both sides, nor does it implement the full rule definition of a chase. Whether a sacrifice works or an attack leads to mate still requires search.
Red and Black each have their own perspective. Inputs to the later layers are arranged with the side to move first and the opponent second. The network ultimately produces one position score from the side-to-move's perspective.
Sources: position features and grouping, filtering and extracting relationship features, transforming the two perspectives.
Current Architecture: A Shared Front End, Later Layers Chosen by Material
Each Network contains a feature transformer and 16 sets of later-layer parameters. Normal evaluation selects one set using configurations such as the numbers of chariots and combined horses and cannons on each side. It does not run 16 complete networks and take a vote for every position.
After feature transformation, the later layers receive 1024 inputs. Two intermediate affine layers each produce 32 values, ultimately yielding one score. But describing this simply as “1024→32→32→1” misses some structure: the code also concatenates clipped and squared activation results and keeps a connection from an earlier layer to the output.
These intermediate values are network channels, not successive moves or hand-assigned terms such as “chariot score” and “king safety score.” Normal evaluation in this version returns one score; descriptions of dual networks or two independent evaluation outputs in other versions do not apply.
Sources: network layers and propagation, selecting among the 16 sets of later layers, the normal evaluation entry point.
Incremental Updates: Recompute Only What Changed
Adjacent positions usually differ by one move. Adding up every input again after each move would be costly. NNUE accumulators store results from the front part of the network: a move subtracts the contribution of the old piece location and adds that of the new one. A capture also removes the captured piece's contribution, while changed attack relationships are updated too.
Incremental updates reuse the front-end results; they do not skip the whole network. The later layers still need to run. The search stack records changes, and only when an evaluation is needed does the engine look for reusable accumulators and apply the missing updates. Undoing moves also returns to the corresponding state. This differs from the transposition table's reuse of search results.
Nor can every move directly reuse the previous accumulator. A move by the friendly king, a capture that changes the material grouping of the features, or a change in mirroring can require a refresh for the affected perspective. A refresh can still use caches organized by king location and other information: it updates position features from the piece differences, then adds the current relationship features again, instead of always rebuilding everything from zero.
Sources: the accumulator stack and updates on demand, feature-difference updates, refresh conditions, refresh caches.
Quantization and SIMD: Keep Each Evaluation Cheap
Current inference mainly uses integer arithmetic: feature weights use 8-bit integers, accumulators use 16-bit integers, and later-layer multiply-add results use 32-bit integers, together with clipping, shifts, and scaling. This is what quantization means here: representing network values on specified integer scales.
The code also provides SIMD paths for different CPUs, processing several values with one instruction, such as the SSE/AVX families on x86 and NEON on ARM. Some layers exploit the many zero values in their inputs to reduce work. These paths accelerate the same evaluation rather than giving one build extra knowledge about the game. Actual speed depends on the CPU, build path, and overall search workload.
Sources: integer types and scales, feature transformation, later-layer multiply-add operations, the sparse-input layer.
Network Files: Matching Names Do Not Guarantee Compatibility
EvalFile specifies the network file; the current default name is pikafish.nnue. The loader tries the relevant paths, reads the format version, architecture identifier, and layer parameters, and checks that the read completes properly. If verification finds that the required network has not loaded successfully, it reports compatibility and path issues and terminates. It does not silently fall back to a simple material evaluation and keep playing.
A replacement network must therefore match the engine architecture. Renaming a file cannot turn an old architecture into a new one. The architecture identifier checks format compatibility; it is not cryptographic authentication of an arbitrary file.
Sources: the default filename, finding and loading the file, format and parameter checks, load verification and error handling.
Reading the eval Debug Output
The engine's eval command shows evaluation details for the current position. It does not perform a normal search to choose a move. Current output includes both NNUE internal units from the side-to-move's perspective and display scores converted to Red's perspective. The retained label white side refers to Red in the Xiangqi implementation.
The final static score in this debug output uses zero optimism and does not reconstruct the correction history of a particular search node. It need not equal the score after a go search. When in check, the debug entry point explicitly reports that no ordinary final static evaluation is available.
Sources: the eval command entry point, debug evaluation output. For the differences between normal search scores, UCI cp, mate, and WDL, see Score Output.
