It’s hard to escape the momentum around Voice UI. From consumer assistants to connected vehicles and industrial systems, the industry’s pacesetters in the AI race are betting that more natural interfaces will transform how people interact with technology, argues Sondre Ager-Wick, Director of User Experience, Qt Group
At the end of 2025, OpenAI began unifying various engineering, product, and research teams to overhaul its voice models as it prepares for a voice-first personal device expected to launch in about a year.
Google DeepMind hired the CEO and several of the top engineers at voice AI startup Hume AI, and Apple just paid $2 billion for Q.ai, an AI company with technology touted to dramatically enhance the Siri voice assistant.
Meanwhile, Meta has brought on former Apple Human Interface design chief Alan Dye to lead a new creative studio within Reality Labs, aimed at defining the next generation of human-computer interaction.
It’s easy to get swept up in the appeal of interactive modes that feel like they are propelling us into an AI era mirroring our sci-fi fantasies.
Beyond the allure of new form factors
As practitioners in this space, we must make sure we don’t mistake shiny new wrappers for true progress. This applies not just to voice-first paradigms, but to the broader category of wearables.
We should be thinking harder about what makes an AI experience actually useful for consumers. Is it the interface’s placement? Is it the look and feel? Or is it something that isn’t talked about with anywhere near the same degree of urgency: the intelligence embedded behind the UI that interprets user intent, context and anticipates needs.
History tells us that while novel form factors are immediately compelling, sustainable product success comes from respecting people’s priorities and delivering outcomes.
The Humane AI Pin and the Rabbit R1 were ambitious attempts at reimagining interaction, but they faced significant real-world hurdles in performance and utility.
While the industry’s incumbents have the resources to iterate on these hardware challenges, these early examples show that a new form factor doesn’t automatically solve a user’s problems. There is no singular mode that functions as a universal interface; the value is in the execution of the task.
Intelligence is the real interface
App design has long revolved around menus, icons, and screens, which often require multiple taps to complete a simple task. Think about adjusting the temperature in your home – even if you have a ‘smart’ thermostat, you currently open an app, navigate to the right room, and manually set a value. A genuinely intelligent system wouldn’t wait to be asked. It would notice you’ve arrived earlier than usual, cross-reference your preferred settings, and have the temperature right before you’ve taken your coat off. The future of AI is about the contextual awareness that means you shouldn’t have to explain yourself twice.
What separates the AI experiences that will earn lasting adoption from those that is the quality of interpretation, rather than the interface itself. The most impactful systems will be those that quickly and accurately read your intent, drawing on multi-modal input, contextual awareness, and a genuine understanding of your preferences, then present you with options to simply choose or confirm. That keeps you in control while removing the friction of having to direct every step. Trust is built not through novelty, but through the system proactively getting it right – consistently, and from early on.
This is where the IoT context becomes particularly instructive. Devices operating at the edge, particularly in vehicles, industrial equipment, smart home systems, can’t always rely on cloud connectivity, but they still need to interpret intent intelligently and respond in real time. The challenge of building AI that works under constrained conditions is the sharpest test of whether your intelligence layer is genuinely robust, or just impressive in a demo.
Reliability is not optional
None of this scales if it only works on flagship devices with always-on, high-speed connectivity. VR has already delivered that cautionary tale with jaw-dropping demos that fell apart the moment real-world constraints kicked in. AI can’t repeat that mistake. For this technology to reach mainstream adoption, it must function reliably across diverse hardware, on legacy networks, and in low-latency environments. At the heart of any context-aware system is data. As we move toward deeper awareness, we must prioritise how information is handled. Our progress as an industry depends on striking a careful balance between usefulness and clear consent. Intelligence only works when people feel in control; without that, it becomes intrusive.
We are already seeing successful implementations of this balance. Apple’s approach of keeping personal data on device while calling on Cloud models like Gemini for heavy-duty processing is a great example. Samsung’s Galaxy AI follows a similar path, handling tasks locally whenever possible.
The path forward
The companies doing this well are doubling down on meaningful, functional progress rather than novelty for its own sake
The next success story in AI UX will come from tools that grasp context and learn what people really need over time, while still leaving users firmly in charge.
The real future of voice UI (and AI interfaces more broadly) is everything we can’t see. The invisible layer of intelligence that understands what you mean, anticipates what you need, and gets out of the way and gets you straight to your desired outcome.
That’s what we should be building.
Author biography:
Sondre Ager-Wick is Director of User Experience, Qt Group.
