Imphash, Rich Header Hash and TLSH: Grouping Samples
How imphash, the Rich header hash and TLSH group related Windows executables, how each is computed, and where each one breaks down.
TL;DR. A SHA-256 changes when one byte changes, so it cannot tell you that two files are related. Three other fingerprints can: the imphash (what the program imports), the Rich header hash (which Microsoft toolchain components built it) and TLSH (how similar the bytes are). Each fails in known ways. Use them to form groups, then confirm with other evidence.
Imphash: the import fingerprint
Mandiant introduced the import hash in 2014 (Tracking Malware with Import Hashing). The widely used implementation is get_imphash() in pefile:
- For each imported DLL, in file order, take the lower-case name and drop a
.dll,.ocxor.sysextension. - For each function, take its lower-case name. Functions imported by ordinal are named from pefile's tables for
ws2_32,wsock32andoleaut32, and writtenordNotherwise. - Join
dll.functionpairs with commas and take the MD5.
Because the linker writes imports in the order the code uses them, builds from the same source tree tend to share an imphash while unrelated programs rarely do. In the sample loaded by the PE Parser, the two builds of m64.exe have different SHA-256 values but the same imphash, and the batch table labels them family 1.
Where it breaks: packed files (the imphash describes the packer's stub, shared by thousands of unrelated samples), .NET assemblies (they import only mscoree.dll!_CorExeMain), and very small programs whose few imports are common to everything. Delay-load imports are not part of pefile's imphash. See the imphash glossary entry.
Rich header hash: the toolchain fingerprint
Microsoft's linker writes an undocumented block into the DOS stub: a list of comp.id values (product id and build number of each compiler, assembler, linker and import library that contributed objects) with a count for each, XOR-masked with a key and closed by the word Rich. The key is not random: it is a checksum of the DOS header and of the entries, as described by Daniel Pistelli in Microsoft's Rich Signature (undocumented).
That gives two things:
- a Rich header hash (pefile's
get_rich_header_hash(): MD5 of the decoded header), shared by programs built with the same toolchain mix — often the same project; - a consistency check: if the stored key does not equal the recomputed checksum, the header was edited or transplanted.
Forgery is real. Kaspersky showed that the Olympic Destroyer sample carried a Rich header copied from another family to mislead attribution (The devil's in the Rich header). Research such as Webster et al., Finding the Needle: A Study of the PE32 Rich Header and Respective Malware Triage (DIMVA 2017), shows how well it clusters in practice. The Rich header entry summarises the format; the PE Parser decodes comp.ids to Visual Studio releases using the public list maintained by the richprint project.
Where it breaks: non-Microsoft toolchains (MinGW, Go, Rust with the GNU target, Delphi) write no Rich header, and it can be stripped or zeroed.
TLSH: the byte-similarity fingerprint
TLSH (Trend Micro Locality Sensitive Hash) summarises the distribution of byte trigrams into a 70-hex-digit digest, prefixed T1. Two digests are compared with a distance: 0 means identical, and smaller means more similar. Close variants of one program usually land at small distances; there is no universal threshold, so calibrate on known-related and known-unrelated files from your own collection. TLSH needs at least 50 bytes with enough variety to produce a digest.
The PE Parser's comparison view shows the TLSH distance between any two loaded files. ssdeep, the other common fuzzy hash, is not included: its reference implementation is GPL-licensed, and we did not find a permissively licensed one.
Where it breaks: packing or compression changes the bytes entirely; large shared resources (icons, embedded runtimes) can make unrelated files look close.
Putting them together
| Fingerprint | Groups by | Survives recompilation | Survives packing | Forgeable |
|---|---|---|---|---|
| SHA-256 | exact bytes | no | no | no |
| Imphash | import list | often | no (hashes the stub) | yes, easily |
| Rich header hash | toolchain mix | often | sometimes (stub may keep it) | yes |
| TLSH | byte distribution | partly | no | with effort |
A sound workflow: cluster a collection by imphash and Rich hash, check each cluster's TLSH distances and timestamps (PE timestamps), and only then write "same family" in a report — with the evidence listed.
FAQ
What is an imphash?
An MD5 of the ordered list of imported functions, written as lower-case dll.function pairs joined by commas. Two builds of the same program often share it even when every byte of their code differs.
Is the Rich header reliable for attribution?
It is a useful clue but can be copied or forged: the Olympic Destroyer malware carried a Rich header taken from another family. Check whether its checksum is consistent and weigh it with other evidence.
Why is there no ssdeep in the PE Parser?
The reference ssdeep implementation is licensed under the GPL and we found no permissively licensed implementation, so the tool uses TLSH (Apache-2.0) for similarity instead.