A week, because a day proves nothing
The first day is unrepresentative in both directions — everything is novel, and nothing has accumulated. Five days shows you the actual shape.
Run the daily rhythm for five working days and come out with a calibrated queue and a written list of what needs fixing.
The point of the practice week is not to get through your queue, you would do that anyway. It is to build the rhythm while you are still deliberate about it, and to gather evidence about what is arriving badly while you can still see it clearly.
Both halves have a shelf life. The rhythm is easy to establish in a week when you are paying attention to it and hard to establish later by resolution alone. And the clarity about what is wrong is highest in the first days, before familiarity turns a recurring annoyance into just how the queue is.
Work the queue at the two times you picked in the earlier lesson, and resist working it continuously in between. Continuous queue-working feels responsive and is the actual failure mode: it fragments the day, and it means every item is handled in whatever state of attention you happen to be in when it arrives.
Start each pass with overdue items (they carry the red rail for a reason) then work today. Aim to leave nothing overdue overnight, which is an achievable standard for most roles and a useful one because it keeps the backlog from becoming the thing you manage instead of the work.
Keep a running note as you go: one line for every item you rejected or edited, with the reason. It takes seconds per item and it is the entire deliverable of the week.

On the fifth day, read your note rather than your queue. What you are looking for is repetition: the same kind of item edited the same way four times is not four judgment calls; it is one defect showing up four times. Those repeats are your fix list, and they are worth more than any impression of how the week felt.
Judge the week on Friday and not before. Early days are noisier than settled ones, instructions are still rough, thresholds are untuned, and the system has not yet been corrected by the feedback you are in the middle of generating. Reading Tuesday as a verdict on the model is how teams talk themselves out of something that was two days from working.
Do this in the product
You should be able to answer each of these from memory before opening it. Recalling the answer is what makes it stick; recognizing it when you read it does not.
Continuous queue-working feels responsive and is the failure mode. It fragments the day and means each item is handled in whatever state of attention you happen to be in when it arrives.
The running note: one line per rejection or edit, with the reason. On Friday the repeats in that note are the fix list, and they are worth more than any impression of how the week felt.
Early days are noisier because instructions are rough and thresholds are untuned. Reading Tuesday as a verdict is how teams talk themselves out of something that was two days from working.
The first day is unrepresentative in both directions — everything is novel, and nothing has accumulated. Five days shows you the actual shape.
One line per item you rejected or edited, with why. By Friday the repeats in that list are your fix list.
Week one is noisier than week four. Reading Tuesday as a verdict is the single most common way teams abandon something that was working.
See agents carry the repeatable work of GTM across sales, marketing, customer success, and RevOps. Every action prepared, reviewed, and recorded. Fictional data, real product.
Explore the demo