Badly, is the short answer, which is why fifty years has not settled it.
The problems, roughly in order of how much damage they do:
Self-report is nearly worthless here. Perceived sleep quality correlates poorly with measured sleep. People routinely report sleeping well on nights that polysomnography says were fragmented, and badly on nights that were fine.
The act of measuring changes it. Sleeping in a lab with electrodes is not sleeping.
Enormous night-to-night variance. Which means a small effect needs a lot of nights or a lot of people, and almost no study of this kind has either.
Expectation is unusually powerful. Sleep is close to the maximally suggestible endpoint. You take something at bedtime believing it will help you sleep. Also: you have now established a bedtime routine and are thinking about sleep on purpose, both of which are actual sleep hygiene interventions.
So a compound taken before bed has to beat placebo and beat the behavioural intervention of taking a compound before bed.