Before you commit engineering time to an AI feature, find out which direction is actually useful and earns trust.
The directions usually differ in more than appearance. One might answer in conversation, another inside the screen the person is already on. One might do the work, another suggest what to do. These are different products for the same problem, and people respond to them differently. Choosing between them by discussion tends to settle on whoever argues most confidently.
In some work a plausible but wrong answer is worse than no answer, because the person acts on it. Where that's true, people need to see where an answer came from before they will use it. What earns that confidence is specific and can be tested, and it is usually not what the team assumes it is.
An AI feature generally costs more to build and more to run than the feature next to it on the list. A direction that people were observed choosing, together with the reasons they gave for choosing it, is a different kind of argument than a preference. It also gives the team something to check the finished build against.
A within-subjects test that leaves you with a defensible build decision:
Within-subjects means the same person sees every direction, so they can compare them directly and say which they would rather use and why. Showing each direction to a different group tells you less, because the direction and the people who saw it can't be separated afterwards. Where the workflow is specialist, the participants need to be specialists too.
Prerequisitetwo or three concept directions defined and mocked to a comparable level of finish, using consistent example data. The comparison is only as good as the parity between them.
A concept test produces more than a winner. It shows the conditions under which people accepted an answer at all — what they needed to see, and at what point they needed to see it. Written down, those become requirements the build can be checked against, rather than something everyone remembers differently six months later.
One direction, the reasoning behind it, and what would have to be true for it to be wrong. The reasoning matters because the recommendation gets repeated in rooms you are not in.