1. Initialize Before Anything Else
In HarmonyOS automation scripts, every image-related capability hangs off the image and ocr objects, and there is one step you cannot skip: initializing OpenCV. Forget it and matches return empty without raising an error, which can keep you stuck for a long time.
After initialization, switch image storage to mat format. Vendor measurements are concrete: memory down by half to eighty percent, CPU down twenty to thirty percent, and speed up one to two times. In HarmonyOS cluster control, where the same script runs across a rack, the device count multiplies that difference.
// Initialize OpenCV first; image and color matching depend on it
image.initOpenCV();
// Capture, read, match, and color operations move to mat format
image.useOpencvMat(1);
Recognition usually starts from a full-screen capture. If a flow needs many consecutive frames, a streaming capture avoids rebuilding the connection each time. Recycle image objects when you are done with them, because a long-running script will steadily eat memory otherwise.
2. Image Match in Practice
Image matching takes a template and searches the current screenshot for its position. The return value carries the matched coordinates, which is all you need to tap.
let screen = image.captureFullScreen();
let tpl = image.readBitmap("/sdcard/ec/tpl_confirm_btn.png");
let pos = image.findImage(screen, tpl, { threshold: 0.9 });
if (pos) {
logd("confirm button found: " + JSON.stringify(pos));
} else {
logw("confirm button not found");
}
image.recycle(screen);
Three practical rules. Crop the template on the target device, because a screenshot from one model may render differently on another and the match rate drops. Do not crop too small, because the template then resembles everything and a loose threshold misfires. And keep changing text out, leaving it to OCR, so the template contains only stable visual features.
For the threshold, start from a sensible default, read the similarity the script actually reports, then tighten until false hits stop. Different resolutions may need different values.
3. Color Comparison in Practice
Color comparison reads color values, which makes it much lighter than image matching.
Single-point comparison suits checking the state of one fixed position. If a button looks different when it becomes tappable, comparing that pixel is enough and far faster than matching an image.
// Single-point comparison: is the color at this coordinate within range
let hit = image.cmpColor(screen, "#FF5252", 0.9, 0, x, y);
Multi-point comparison solves uniqueness. One color may appear all over the screen, so binding several points with relative offsets into a group narrows it sharply.
There is also a find-by-color-then-match pattern: narrow the region by color first, then run image matching inside it. That is faster and more accurate than scanning the whole screen.
The weakness of color comparison is change. A new accent color, dark mode, or a different background image can break it, so treat it as a fast filter and a fallback rather than the only check.
4. OCR in Practice
OCR has one more choice to make than the other paths: the model type.
The built-in PP-OCRv6_small is the speed-oriented route and handles upright interfaces well. Available types also include ocrLite, tesseract, and the Onnx series V4 and V5, with different trade-offs between accuracy and speed. The snippet below initializes on the PPOCR V6 path.
ocr.releaseAll();
let engine = ocr.newOcr();
let ok = engine.initOcr({
type: "paddleOcrOnnxV6",
numThread: 2,
matMode: 1,
maxSideLen: 640,
doAngleFlag: 0,
mostAngleFlag: 0
});
if (!ok) {
loge("OCR init failed: " + engine.getErrorMsg());
return;
}
Tune in this order. Padding first, because expanding a white border around text boxes fixes partial captures immediately. Then the detection box confidence threshold, where raising it reduces recall but improves precision. Then maxSideLen, capped at 640 for speed, at the cost of very small text. Then text direction detection and angle voting, needed only for rotated images.
Results come back as an array where each entry carries text, a confidence value, and a coordinate range. During debugging, log both the label and the confidence so you can tell a miss from a wrong read.
5. When YOLO Is Worth It
YOLO covers what the other paths cannot: a target whose shape and position vary but which you can describe as a category.
It adds a model file, so initialization and tuning are heavier than the rest, and training follows the same approach as on Android. Avoid it unless nothing else works. The test is simple: if image match and color comparison plus OCR recognition together keep the flow stable, you do not need a model.
6. The Order to Debug In
Follow this and you will usually locate the problem.
Confirm the screenshot was captured. If capture returns nothing, everything downstream fails, so check this first.
Confirm the coordinate systems match. Taps and recognition on HarmonyOS Next both use screenshot coordinates, so if the screen size was set midway, verify both sides agree.
Confirm OpenCV is initialized and the mat switch succeeded.
Then log the matched coordinates alongside the screenshot. Many “element not found” cases are actually elements found at the wrong position, which is obvious the moment you look at the image.
One habit worth building: save the screenshot before the script exits on failure. Looking at the scene beats reading logs afterwards.
About EasyClick: A phone automation AI-agent platform covering Android no-root, iOS no-jailbreak (proxy / Bluetooth HID / OTG HID) and HarmonyOS Next, offering script development, Apple cluster control, local central control & mirroring, and cloud control systems. → Explore all products
Ready to build it for real?
Every approach in this article can be built with EasyClick capabilities on iEasyClick — full documentation, developer tools and automation products, free to try.