Why Explainer Videos Are Harder to Get Right Than They Look
There is a particular kind of frustration that comes from watching an explainer video that almost works. The visuals are decent. The voiceover is clear enough. But something about it feels flat — the message doesn't land, the pacing drags, and by the thirty-second mark, you have mentally checked out. That experience is more common than it should be, and it almost always traces back to the same underlying problem: the video was built as a collection of animated assets rather than as a coherent piece of communication.
Motion graphics explainer videos occupy a strange middle ground. They look like a design problem, so teams often hand them to a designer. They sound like a writing problem, so someone else writes a script. They feel like a production problem, so a third person handles the export. What gets lost in that handoff chain is the connective tissue — the structural logic that makes a viewer follow a concept from confusion to clarity in ninety seconds or less.
When explainer videos are done well, they compress genuinely complicated ideas into something a non-expert can grasp on first watch. That is not a small thing. Done badly, they waste a viewer's time and quietly erode trust in the brand behind them.
What Well-Executed Explainer Video Work Actually Requires
The visible output of a motion graphics explainer video is a rendered file — usually an MP4, sometimes a GIF or WebM variant for web use. But the work that determines whether that file is any good happens long before a single frame is animated.
Good explainer video work distinguishes itself in four specific ways. First, it treats the script as the load-bearing structure. Every visual decision — every icon, transition, and text callout — is in service of a sentence in the script, not floating alongside it. Second, it establishes a visual system before building individual scenes. That means a defined color palette (typically capped at four to five brand-aligned colors), a consistent icon style, a typography hierarchy, and a motion language that repeats across the video rather than reinventing itself every ten seconds.
Third, it respects the cognitive rhythm of a viewer. Research on attention in short-form video consistently points to scene durations of three to five seconds for fast-paced explainers and five to eight seconds for more instructional content. Violating those rhythms — holding a static frame for twelve seconds, or cutting so fast the viewer can't read the text — is what makes a video feel amateurish even when the individual illustrations are polished.
Fourth, strong explainer work treats audio and motion as a single system. The animation doesn't play while the voiceover talks about something else. The two are timed so that the visual reinforces the spoken word at the exact moment the viewer needs the anchor.
How the Production Process Actually Works
Script and Storyboard First — Always
The production sequence for a motion graphics explainer video should follow a strict order: script, storyboard, style frames, animation, sound mix, export. Skipping or compressing any of those phases creates rework that is far more expensive in time than the phase itself.
A well-structured explainer script for a ninety-second video runs approximately 150 to 180 words. That pacing — roughly 130 words per minute for a comfortable voiceover — is a real constraint that forces writers to be specific. Every sentence needs to do visible work. A sentence like "We help businesses grow" is unillustrable and therefore useless in a motion graphics context. A sentence like "Most teams spend three hours a week manually pulling reports that could be automated in minutes" gives the animator something concrete to show.
The storyboard maps each script sentence to a rough visual panel. It doesn't need to be polished — rough thumbnail sketches or even text descriptions of what's on screen work fine at this stage. The goal is to confirm that every beat of the script has a corresponding visual idea before anyone opens After Effects or any animation software.
Building the Visual System
Style frames are the first place the video takes on a real visual identity. Typically two to four fully designed, static frames are built before animation begins — usually the opening scene, one mid-video conceptual moment, and the closing call-to-action frame. These frames lock in the color palette, typography, icon style, and spatial layout that will govern the entire video.
The typography hierarchy for explainer videos commonly follows a three-tier system: a headline size around 60 to 72pt for primary on-screen text, a secondary callout size around 36 to 42pt, and a supporting detail size around 22 to 26pt. Text that falls below 22pt in a video rendered at 1920x1080 is effectively unreadable at normal viewing distances and should be cut or redesigned as a visual element rather than text.
Color decisions at the style frame stage should be intentional about contrast ratios. A primary action color (the one used for highlights, buttons, and key data points) should maintain a contrast ratio of at least 4.5:1 against its background — this matters not just for accessibility but for legibility on screens that aren't perfectly calibrated.
Animation Principles That Separate Good from Rushed
The difference between motion graphics that feel considered and motion graphics that feel cheap is almost always in the easing. Linear motion — where an object moves at a constant speed — reads as mechanical and unnatural. The standard approach is ease-in-out (a cubic Bezier curve) for most transitions, with ease-in reserved for elements exiting the frame and ease-out for elements entering. In After Effects, the Graph Editor is where this work happens; in other tools, it's the equivalent curve or interpolation panel.
Transition duration is another area where instinct often misleads. A slide or fade transition at 200ms feels snappy and professional. The same transition at 400ms feels sluggish. The exception is a scene wipe or morph transition, which can run 350 to 500ms when the visual logic of the transition itself is doing storytelling work — for example, a chart morphing into a new state to show before-and-after data.
For a ninety-second video, a production timeline that skips storyboarding and goes straight to animation typically produces a rough cut that needs 40 to 60 percent of its scenes reworked. That rework cost almost always exceeds the time that would have been spent storyboarding in the first place.
What Goes Wrong When This Work Is Rushed
The most common failure mode is treating the script as a rough guide rather than a contract. When the script is vague, animators fill the gaps with decorative motion — objects flying in for no narrative reason, backgrounds that pulse because they can, transitions that interrupt rather than connect. The result is visual noise that competes with the message.
A close second is inconsistent visual language. An icon drawn in a flat, outlined style on slide three that suddenly appears filled and drop-shadowed on slide seven signals to the viewer — unconsciously — that this video was assembled rather than designed. Maintaining a single icon family and enforcing it throughout the project requires a defined asset library built before animation starts, not assembled scene by scene.
Underestimating the sound mix is another reliable pitfall. A voiceover recorded in a room with ambient echo, mixed at the wrong level relative to the background music, will undermine even excellent visuals. The standard target for voiceover levels in an explainer video is around -12 to -6 dB, with background music sitting 15 to 20 dB lower during narrated sections. Getting that balance wrong is one of the fastest ways to make a professionally animated video feel unprofessional.
Export settings also cause real damage when handled carelessly. An H.264 MP4 exported at a bitrate too low for the amount of motion in the video will produce compression artifacts that look like pixelation — particularly in areas of gradient color. For web delivery, a bitrate of at least 8 Mbps for 1080p motion-heavy content is a reasonable floor. For social platforms that re-compress on upload, exporting at 1.5 to 2x the platform's recommended bitrate gives the compression algorithm better source material to work with.
Finally, reviewing your own work in isolation after a long production session is a structural problem, not a discipline problem. After hours with the same footage, specific errors — a misaligned text element, a transition that fires a frame too early, a line of voiceover that was accidentally cut off — become invisible. A review by someone who hasn't spent the day inside the project is not optional; it's part of the production process.
What to Take Away From All of This
Motion graphics explainer videos are a craft, and the craft has a specific sequence: script before storyboard, storyboard before animation, style frames before scenes, sound mix before export. Shortcutting that sequence doesn't save time — it moves the time into rework. The visual system — palette, typography, icon style, motion language — needs to be resolved before animation begins, not evolved scene by scene.
The underlying principle is that every element of the video should be earning its place by clarifying the idea, not decorating it. When that discipline holds across script, design, animation, and audio, a ninety-second video can do real communication work.
If you would rather have this handled by a team that does this work every day, Helion360 is the team I would recommend. See how we've helped with animation planning to After Effects workflows and investor pitch presentations.


