This moves locking to the inodes themselves which allows reducing lock times significantly. Main inodes (ext2 and tmpfs) still do contain a single big mutex that gets locked during operations but now we have the architecture to optimize these.
rdrand
I don't really want to be working with i386 since it doesn't support compare exchange instruction