Arriving at the office to find nothing ran
A friend running a device room in Shenzhen complained about this to me last year.
He got in at half nine and found the publishing job scheduled overnight had not run on a single device. He suspected the devices had dropped — half an hour of checking, all online. He suspected the script — rolled back a version, still nothing.
Eventually he found the status in the task panel reading “paused”: a step in the flow waiting for a human to click it. The job had stalled at 2am. He found out at half nine. Seven and a half hours in which nothing happened.
He said one thing that stuck: I do not mind it failing. I mind finding out by accident.
That is the single thing this article fixes. We have an article on device inspection elsewhere, and that one covers you going to look — opening the console, checking online status, reading execution records. This one covers the opposite direction: making the system come to you.
Why keep both? Because they surface different classes of problem.
Inspection is good at the slow kind: a few devices that keep dropping offline, an account whose numbers have been sliding, a cable starting to make intermittent contact. None urgent; all of it becomes an incident if left alone.
Alerting handles the sudden kind: the scheduled job that failed wholesale at 2am, the device offline since yesterday afternoon, the script that stopped halfway. None of that can wait.
Step one: make it visible first
Alerting needs something to alert on. So the first job is not wiring notifications, it is knowing the task panel.
Each entry in the task panel shows four fields — status, source, how many devices ran, and success/failure counts — plus the time. The source tells you whether the task came from an AI session, a workflow editor test run, or a scheduled trigger; named schedules show their name too.
Two of the statuses matter most:
- Partially successful: the flow ran end to end but devices dropped out. This one gets mistaken for success, and it is exactly the earliest signal — 4 failures out of 30 every single time usually means those 4 have a cable or OS version problem.
- Paused: usually a step waiting for a human. If the flow was meant to run unattended, a pause means nobody advances it and it sits in the background indefinitely.
Open a task and you see each device’s individual status; failed devices carry an error description. This is how you size the impact: the whole batch, or a cluster within it.
As for running overnight at all, that is what scheduled tasks are for: they run a saved workflow on your time rules, and each trigger produces an ordinary task in the same panel. So as long as the schedule is right, what ran at night is inspectable by day.
One trap worth knowing up front: if every target device is skipped at trigger time — because they are busy, say — the trigger counts as failed and the reason is written into the recent-errors column of the scheduled task list. No task appears in the task panel at all, because none was created. When you are chasing “the job didn’t run”, check both places.
Step two: make it come to you
Being able to look is not enough. You want it to knock.
A script can post straight to a DingTalk bot. EC USB 6.16.0 and above provides the function: pass the bot webhook URL, the secret and the message, and you can @ named people or @ everyone.
function main() {
let url = "https://oapi.dingtalk.com/robot/send?access_token=YOUR_TOKEN";
let secret = "YOUR_SECRET"; // optional — keyword filtering can be used instead
var res = sendDingDingMsg(url, secret, "Overnight publishing job failed — check the task panel", "", true);
logd("alert result: " + res);
}
main();
It returns a JSON string where errcode of 0 means sent and anything else is an error. Log that return value — otherwise you will not notice when the channel itself breaks.
The minimum useful version puts this call in the script’s error branch: a failure posts a message. A better version reports at milestones too — “all 30 devices finished”, “4 failed” — so you learn both that something went wrong and what state it ended in.
The companion to alerting is the run log. After a failure, open the log to find which step it stopped on: element not found, load timed out, a permission dialog in the way. The log opens from the task panel, and replaying a failed step in the editor marks the reason.
Step three: decide which events to cover
A single “notify on task failure” is not enough. Cover at least four — the first two business-level, the last two infrastructure-level:
- Whole-task failure: the batch did not run, today’s delivery is affected.
- One device failing several times in a row: one failure is noise; the same device repeatedly means the machine, not the script.
- A device offline beyond a set duration: two hours, say. This catches devices that dropped on their own overnight, which the task panel may not reflect.
- Licence or environment problems: the nastiest, because the script never started, so no alert can possibly fire. This needs an independent scheduled check rather than waiting for a script error.
Step four: grade it, or it becomes noise
This is the step most people skip, and then they are immune to their own alerting within a week.
Round-the-clock push means three wakes in one night, two of them a single device. By day three you have muted notifications. On day four the real incident arrives and you never saw it.
Two tiers work well:
| Tier | What | How |
|---|---|---|
| Immediate | Whole batch failed, a critical in-hours task did not run, wide offline | Push to phone and @ someone |
| Digest | Single-device failure, small out-of-hours anomalies, partially successful | Collected into one morning summary |
The key is reserving “@ someone” for the tier that genuinely needs you out of bed. @ everyone on every alert is the same as @ nobody.
One boundary worth stating: alerting buys you time, it does not solve the problem. If you have not decided what to do when woken — rerun first, or stop first — then getting up only means staring at a screen. Settle that during working hours and write it into the process.
Do not let the alerting die quietly
A broken alert channel is invisible: messages never send, the script reports nothing, and everything looks fine.
Three common failure modes:
- Bot configuration lapses. Secret or URL expiring, the bot removed from the group, keyword filters changed so nothing matches.
- Licence expiry. As above: the script never starts, so nothing fires.
- The console machine itself. Shut down, network dropped, restarted by an OS update. Easiest to overlook, because the entire alerting chain runs on it.
The countermeasure is a reverse self-check: send a test message on a fixed schedule and confirm the channel is alive. Check licence validity the same way — the function that returns licence information includes an expiry time, so make it a scheduled task that warns you ahead of the date.
Order of work once an alert arrives
Four steps:
- Size it. Whole-task failure or partial? Scope determines everything after.
- Open the task for detail. Which devices, what the errors say, whether they share a trait — one batch of cables, one OS version.
- Locate it in the run log. Which step it stopped on, whether a screenshot was kept, whether earlier steps were normal.
- Only then act. Rerun, skip the problem devices, or change the script.
The most common result of reversing that order is rerunning on sight, repeating the same failure and destroying the scene. Judge first, act second.
Finally
Back to that friend in Shenzhen. He wired the alerting in, and he says the biggest change was not fewer incidents — it was sleeping properly.
Unattended running does not rest on scripts never failing — that is unachievable. It rests on you knowing when they do, and knowing what to do about it.
Build the smallest useful version first: create the scheduled task, learn the four task-panel fields, add one DingTalk alert call to the error branch. Usually an afternoon. The remaining time belongs to grading and channel self-checks, which is worth more than a longer alert list.
If you are already doing manual inspection, read that alongside this: how to run device inspection for Apple cluster control. For building the schedule itself, see AI agent scheduled tasks.
About EasyClick: A phone automation AI-agent platform covering Android no-root, iOS no-jailbreak (proxy / Bluetooth HID / OTG HID) and HarmonyOS Next, offering script development, Apple cluster control, local central control & mirroring, and cloud control systems. → Explore all products
Ready to build it for real?
Every approach in this article can be built with EasyClick capabilities on iEasyClick — full documentation, developer tools and automation products, free to try.