apgsearch v3.1
Posted: March 19th, 2016, 9:23 am
The latest version of apgsearch (v3.x, codenamed apgmera) is available from:
https://gitlab.com/apgoucher/apgmera/
It's basically a faster version of 2.x with plenty of extra bonus features borrowed from 1.x. This hybrid nature is the reason for the choice of codename (on the pattern of 'chimera'). It also has a penchant for self-modifying and recompiling its code whenever necessary (such as changing the rule and symmetry).
The interface is very similar to that of 2.x (with which I assume you're familiar). The only noticeable difference is that to compile the program, you run the helper script:
instead of the usual 'make'. Also, there are --rule and --symmetry options, so you can run (for example):
and it will perform the necessary code generation, recompilation and self-execution. This relies heavily on the Python script rule2asm.py, which generates highly-optimised assembly code for running the specified rule.
Advantages over v2.x:
https://gitlab.com/apgoucher/apgmera/
It's basically a faster version of 2.x with plenty of extra bonus features borrowed from 1.x. This hybrid nature is the reason for the choice of codename (on the pattern of 'chimera'). It also has a penchant for self-modifying and recompiling its code whenever necessary (such as changing the rule and symmetry).
The interface is very similar to that of 2.x (with which I assume you're familiar). The only noticeable difference is that to compile the program, you run the helper script:
Code: Select all
bash recompile.shCode: Select all
./apgmera --rule b3s238 --symmetry D2_+1 -n 5000000Moreover, rule2asm.py also creates versions of the rule simulation algorithm in three different instruction sets (SSE2, AVX and AVX2); apgmera chooses the highest version supported by the CPU on which it's running. This gives a particularly noticeable speed boost on machines with AVX2 support, since it processes 256 bits simultaneously instead of just 128.an aside wrote:(Specifically, it consults a lookup table of minimal Boolean circuits for all 32768** falsity-preserving* 4-input Boolean functions, which was precomputed in 20 minutes by another Python script.)
* That is to say, f(0000) = 0. These are the only functions that can be implemented using the assembly instructions vpor, vpand, vpandn and vpxor. Restricting to falsity-preserving functions simply means that B0 rules are forbidden (as has always been the case).
** You may be wondering that, since there are 131072 non-B0 rules but only 32768 falsity-preserving 4-input Boolean functions, from whence the extra factor of 4 is obtained. There is, in addition to the main 4-input circuit f(c, b0, b1, b2), an optional 2-input auxiliary circuit g(c, b3) whose output is XOR'd with that of the main circuit. The variable c is the current state of the centre cell; the variables b3 b2 b1 b0 are the bits of the live neighbour count (which includes the centre cell due to the way that it's implemented).
Advantages over v2.x:
- Up to 35% faster, depending on the CPU.
- Supports arbitrary outer-totalistic rules.
- Supports three different symmetry types (C1, D2_+2 and D2_+1).
- Can upgrade itself to the latest online version by including the --update option!
- Requires Python 2 (for rule2asm.py) and Bash (for recompile.sh).
- Between 5x and 100x faster, depending on the rule, symmetry, and CPU.
- Does not depend on Golly, and can run on headless machines.
- Better handling of large objects.
- Supports fewer symmetries (3 compared with 16).