Running an A/B test is easy. Running one whose result you should believe is mostly a documentation problem, and the documents have to exist in a particular order.
The reason is uncomfortable but simple. Every experiment involves dozens of small analytical choices — which metric counts, which users are included, when to stop, which segments to examine, what size of effect is worth shipping. Made before the data arrives, those choices are neutral. Made after, every one of them is quietly informed by which option produces the answer somebody wants. Nobody involved experiences this as dishonesty, which is exactly why it needs a procedural fix rather than a cultural one.
So the deliverables split cleanly. Six are written while you are still ignorant of the outcome, and they are the ones that give the other six their authority. The rest describe what happened and what you did about it.
A useful test of any experimentation programme: pick a shipped result at random and ask whether the primary metric and the ship threshold were written down before launch. If nobody can produce that document, the experiment measured something, but it did not decide anything.