Designing For Multimodal Interactions

Explore top LinkedIn content from expert professionals.

  • View profile for Sahar Mor

    I help researchers and builders make sense of AI | ex-Stripe | aitidbits.ai | Angel Investor

    42,602 followers

    Voice is the next frontier for AI Agents, but most builders struggle to navigate this rapidly evolving ecosystem. After seeing the challenges firsthand, I've created a comprehensive guide to building voice agents in 2024. Three key developments are accelerating this revolution: -> Speech-native models - OpenAI's 60% price cut on their Realtime API last week and Google's Gemini 2.0 Realtime release mark a shift from clunky cascading architectures to fluid, natural interactions -> Reduced complexity - small teams are now building specialized voice agents reaching substantial ARR - from restaurant order-taking to sales qualification -> Mature infrastructure - new developer platforms handle the hard parts (latency, error handling, conversation management), letting builders focus on unique experiences For the first time, we have god-like AI systems that truly converse like humans. For builders, this moment is huge. Unlike web or mobile development, voice AI is still being defined—offering fertile ground for those who understand both the technical stack and real-world use cases. With voice agents that can be interrupted and can handle emotional context, we’re leaving behind the era of rule-based, rigid experiences and ushering in a future where AI feels truly conversational. This toolkit breaks down: -> Foundation layers (speech-to-text, text-to-speech) -> Voice AI middleware (speech-to-speech models, agent frameworks) -> End-to-end platforms -> Evaluation tools and best practices Plus, a detailed framework for choosing between full-stack platforms vs. custom builds based on your latency, cost, and control requirements. Post with the full list of packages and tools as well as my framework for choosing your voice agent architecture https://lnkd.in/g9ebbfX3 Also available as a NotebookLM-powered podcast episode. Go build. P.S. I plan to publish concrete guides so follow here and subscribe to my newsletter.

  • View profile for Bertalan Meskó, MD, PhD
    Bertalan Meskó, MD, PhD Bertalan Meskó, MD, PhD is an Influencer

    The Medical Futurist, Global Keynote Speaker, Researcher and Author.

    372,317 followers

    I've been saying for over a year that multimodal large language models will become the ultimate interface between physicians and a range of AI-based solutions. Here is the proof! In this study, the authors developed and evaluated an autonomous clinical AI agent leveraging GPT-4 with multimodal precision oncology tools to support personalized clinical decision-making. They used multiple sources such as histopathology slides, radiological images and search tools like OncoKB, PubMed and Google. "Evaluated on 20 realistic multimodal patient cases, the AI agent autonomously used appropriate tools with 87.5% accuracy, reached correct clinical conclusions in 91.0% of cases and accurately cited relevant oncology guidelines 75.5% of the time. Compared to GPT-4 alone, the integrated AI agent drastically improved decision-making accuracy from 30.3% to 87.2%." Source: https://lnkd.in/dwjGvxcH

  • View profile for SUKIN SHETTY

    Enterprise AI Architect | Building Agentic Systems | Creator of Nemp Memory | Helping Businesses Deploy Real AI | AI Educator

    12,631 followers

    Unlocking Creativity with Agno ‘s Multimodal Capabilities in VibeProto 🤩 I’m excited to share how Agno’s advanced multimodal processing has transformed VibeProto into a powerful tool for non-tech users, enabling them to bring their ideas to life in minutes. 🚀 Today, I want to dive into each modality image, audio, and video Multimodal Power of Agno in VibeProto: 1 Image Input (📷): Upload screenshots, mockups, or wireframes, and Agno’s visual context understanding extracts layouts, color schemes, and design patterns. This feature is perfect for non-tech users who can simply show what they want built. 2 Audio Input (🎵): Record voice descriptions or explanations, and Agno converts natural speech into actionable prototypes. It’s ideal for users who prefer speaking over typing, making the process intuitive and accessible. 3 Video Input (🎥): Upload screen recordings or workflow demonstrations, and Agno analyzes complex interactions to generate code. This modality is a game-changer for replicating existing apps or digitizing manual processes. Demo Highlight: Replicating Airbnb in Minutes In a recent demo, I attached a screen recording of the Airbnb app to VibeProto. Using Agno’s video input capabilities, VibeProto agent analyzed the recording, understood the user flows, and generated a Python script to replicate key functionalities of Airbnb in just minutes. This showcases how VibeProto, with help of Agno can turn visual inputs into working code, democratizing software development for non-tech users. Why This Matters: • Accessibility: Non-tech users can prototype complex applications without coding knowledge, simply by providing visual or auditory inputs. • Efficiency: The process is incredibly fast, as demonstrated by the Airbnb replication, reducing development time significantly. • Innovation: Agno’s multimodal approach pushes the boundaries of what AI-assisted tools can achieve, making VibeProto a standout in the market. Join me in exploring how VibeProto is redefining software development for everyone. Let’s innovate together! 🙌🏼

  • View profile for Eugene L.

    GTM @ ElevenLabs

    21,910 followers

    🔊 Have you ever stayed on a customer‑service call simply because the person on the other end sounded trustworthy? 🎧 Researchers from Beijing University of Technology , the The University of Texas at Austin and the University of Memphis recently tested how different AI voices affect persuasion. Their findings were: • 𝗙𝗹𝗶𝗿𝘁𝘆 𝗱𝗼𝗲𝘀𝗻’𝘁 𝘄𝗼𝗿𝗸. A playful “coquetry” voice actually decreased persuasion, especially for male chatbots. • 𝗦𝘁𝗲𝗿𝗻 𝗶𝗻𝘃𝗶𝘁𝗲𝘀 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀. Stern voices were just as effective as gentle ones and, in male voices, even increased customer questions. • 𝗔𝗴𝗲 𝗶𝘀𝗻’𝘁 𝘁𝗵𝗲 𝗶𝘀𝘀𝘂𝗲. 𝗲𝗻𝗴𝗮𝗴𝗲𝗺𝗲𝗻𝘁 𝗶𝘀. There was no significant difference between “young” and “old” voices. What mattered was that older‑sounding voices kept people talking longer. • 𝗪𝗼𝗿𝗱𝘀 𝗺𝗮𝘁𝘁𝗲𝗿. Using affirmative sentences - particularly in female voices - prompted more customer inquiries, whereas rhetorical questions were less effective. For leaders in banking and finance, this isn’t just academic. Voice is the new front door of your brand. A gentle but confident tone can build trust with high‑net‑worth clients. An affirmative female voice can reassure anxious SME owners. Conversely, a playful chatbot might unintentionally undermine credibility. 𝗦𝗼𝗺𝗲 𝗾𝘂𝗶𝗰𝗸 𝗮𝗰𝘁𝗶𝗼𝗻𝘀 𝘁𝗼 𝗰𝗼𝗻𝘀𝗶𝗱𝗲𝗿: 1. Audit your AI voice scripts. Are you using affirmative statements that invite dialogue? 2. Experiment with different voice personas. Avoid flirty tones and observe how clients react. 3. Treat voice as part of your CX strategy. Integrate data from calls, chatbots and apps so you can personalize the experience for each customer, because customer empathy is your competitive moat. We’ve moved from building “voices” metaphorically to designing them intentionally. The tone of your AI isn’t just a detail, it’s part of the customer experience. Link to research in comments below. #AI #Voice

  • View profile for Bally S Kehal

    ⭐️Top AI Voice | Founder (Multiple Companies) | Teaching & Reviewing Production-Grade AI Tools | Voice + Agentic Systems | AI Architect | Ex-Microsoft

    21,916 followers

    Most voice AI is just a chatbot with a microphone. One company was purpose-built for real phone calls from day one. The architecture lesson most AI builders learn too late: Everyone builds text agents first. Winners build voice first. Text agents get retries. Formatting. Autocorrect. Voice agents get one shot. Real-time. No edits. Then production hits ↓ → Latency: Model takes 3 seconds. Customer hangs up. → Context: "Uh, yeah, so I need to, wait — can you also check my..." → Interruptions: Humans talk over each other. Chat agents break. → Compliance: Every voice interaction is regulated differently. Two traps I see teams fall into: Trap 1: Bolt STT onto a chat agent. Add TTS on output. Call it "voice AI." That's a wrapper. Wrappers break in production. Trap 2: Build your own with Pipecat, LiveKit, Vapi. 6 months later you're managing STT providers, TTS rate limits, LLM deprecations, infrastructure scaling, compliance audits. You wanted a voice assistant. Now you're a voice infrastructure company. PolyAI solved this differently. Full stack built for voice since 2017: → Proprietary ASR + LLM trained on real customer service transactions → 45+ languages. 24/7. Unlimited scale. → Handles surges instantly — storms, outages, promos — zero staffing panic Not just handling calls — generating revenue: → Turning bookings into room upgrades → Enrolling callers into rewards mid-conversation → QA Agents scoring every call automatically → Analyst Agents surfacing patterns no human team catches One healthcare company found fewer complaints from the AI than human reps — on the hardest, most emotional calls. Marriott. FedEx. Caesars. PG&E. 25+ countries. 391% ROI. $10.3M average savings. Payback under 6 months. The companies still running "press 1 for sales, press 2 for support"? Not behind on technology. Behind on architecture. That gap compounds every quarter. My stress test for any voice AI: → Noisy environment → Regional accent → Language switch mid-sentence → Multi-step transaction → Worth 15 minutes if you're evaluating: https://poly.ai/gordon Build or buy — what's your current approach to voice?

  • View profile for Anne Cantera

    ♾️ Director of Conversation / Omnimodal Design @ Stealth | AI Consulting | Agentic AI | Voice & Chat | Model Designer | Multimodal | Adaptive UI | Spec Driven Developer | Founder, Elementyl Intelligence | Teacher

    12,359 followers

    Experimenting please read... I've been living through the exact job market shift everyone else is panicking about. Traditional UX roles are disappearing faster than most people want to admit. My DMs are full of talented UX designers who can't get callbacks. Meanwhile, I'm actively turning down recruiter messages. The difference? I pivoted to conversational AI and voice interface design 3.5 years ago when I saw the IVR-to-AI conversion wave coming. **The Market Reality** Traditional UX job postings dropped 71% from 2022 to 2023. UX research roles fell below 1,000 monthly postings in early 2025. Only 49.5% of designers are finding new roles within three months, down from 67.9% in 2019. Nobody's talking about this: the conversational AI market is projected to hit $32.6 billion by 2030. Companies are desperate for people who can design these experiences. The job titles? Conversational AI Designer, Voice UX Designer, AI Content Designer, Prompt Engineer, LLM Experience Designer. **The Skills That Actually Matter** Most UX designers get stuck thinking it's just learning new tools. It's not. The pivot requires: • Dialog flow architecture (conversations across turns, not screens) • NLP basics (enough to work with engineers) • Voice-first thinking (designing for ears, not eyes) • AI personality design (tone, empathy, error handling) • Multimodal experiences (bridging voice, text, visual) You need hands-on experience with platforms: Cognigy, Dialogflow, Kore.ai, Cresta. If you can't speak about intent recognition, entity extraction, and conversation flows in these systems, you're not ready. Transparency: I'm using AI tools to synthesize patterns faster, but insights come from doing this work. Converting legacy IVR to AI. Designing voice assistants. Building conversational flows that feel human. **Why This Pivot Works** Your user research skills? Critical for conversational context. Wireframing? Translates to dialog mapping. Accessibility knowledge? Essential for inclusive voice design. **The Uncomfortable Truth** Less than 5% of design roles target junior talent. Mass layoffs continue. Traditional UX teams are shrinking. AI automates entry-level tasks. Companies consolidate roles. If you're waiting for the market to bounce back to 2021, I'll be direct: it's not happening. **What Actually Works** The people working right now? AI-adjacent roles. They learned LLMs. Got hands-on with Dialogflow or Cognigy. Repositioned portfolios for conversational thinking. Applied for "AI Experience Designer" not "UX Designer." Stop thinking "learning a specialization." Start thinking survival adaptation. Different energy entirely. Your traditional UX skills aren't worthless. They're the foundation. But they need a new application layer. The question isn't whether to pivot. It's how fast you can move. What are you seeing in the job market?

  • View profile for Jan Beger

    Our conversations must move beyond algorithms.

    91,249 followers

    A voice- and text-enabled conversational agent was primarily used for health information, casual interactions, and clinical data entry — yet over half of users discontinued after a single session, underscoring barriers to sustained digital engagement. 1️⃣ Among 24,537 users of the Albert Health app, 58% engaged in only one session. 2️⃣ The most frequent intents were health information (32%), small talk (20%), and clinical parameter logging (16%). 3️⃣ Voice input dominated casual (64%) and medication-related (53%) interactions; screen-based input was preferred for clinical tasks (61%). 4️⃣ Participants in disease-specific programs exhibited higher sustained engagement than general health users (OR = 0.67). 5️⃣ A higher proportion of voice-based interactions was positively associated with continued use (OR = 1.005); screen-based interaction predicted attrition (OR = 0.994). 6️⃣ Engagement was more likely to be sustained when users employed a balanced mix of clinical and non-clinical intents (OR = 1.56). 7️⃣ Unexpectedly, higher system confidence scores in chatbot responses were associated with reduced user retention (OR = 0.43). 8️⃣ Users aged 15–45 were less likely to sustain engagement compared to pediatric or older adult cohorts. 9️⃣ Fall-back responses (13% of interactions) were frequently due to non-standard speech, slang, or recognition errors, highlighting limitations in natural language processing. 🔟 A modest engagement peak on day 8 aligned with reminder notifications, but overall retention remained low beyond initial use. ✍🏻 Selahattin Colakoglu, Mustafa Durmus, Zeynep Pelin Polat, Asli Yildiz, Emre Sezgin. User Engagement with A Multimodal Conversational Agent for Self-Care and Chronic Disease Management: A Retrospective Analysis. Journal of Medical Systems. 2025. DOI: 10.1007/s10916-025-02202-2

  • View profile for Swapnil Jain

    Co-Founder & CEO at Observe.AI, AI Agents for Customer Experience

    21,553 followers

    We’re seeing a surge in enterprises expanding their Customer Service AI Agents from chat to voice. While many started with Chat AI powered by foundational models + RAG, extending these agents into Voice introduces a whole new set of challenges. Let’s break it down 👇 🎙️ 𝗩𝗼𝗶𝗰𝗲 𝗮𝘀 𝗮 𝗧𝗲𝗰𝗵𝗻𝗼𝗹𝗼𝗴𝘆 - 𝗩𝗼𝗶𝗰𝗲 𝗯𝗿𝗶𝗻𝗴𝘀 𝗶𝘁𝘀 𝗼𝘄𝗻 𝗶𝗻𝗳𝗿𝗮 𝗮𝗻𝗱 𝗺𝗼𝗱𝗲𝗹𝗶𝗻𝗴 𝗰𝗼𝗺𝗽𝗹𝗲𝘅𝗶𝘁𝘆:  • 🔊 Speech-to-text struggles with noise, custom vocab, and number inputs  • 🗣️ Text-to-speech needs fine-tuning for brand tone, emotion, and pronunciation of company-specific terms  • 📞 Real-time performance needs tight integrations with telephony infra: PBXs, SIP trunks, CCaaS platforms, carrier quirks, session management, DTMF fallback, and more 🧠 𝗦𝗽𝗲𝗲𝗰𝗵 𝗮𝘀 𝗛𝘂𝗺𝗮𝗻 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝗰𝗮𝘁𝗶𝗼𝗻 Speech is messy, and people don’t talk like they type:  • “Uh, I mean… yeah” - disfluencies, restarts, imperfect grammar  • Latency matters - voice is synchronous, unlike chat  • Real-world calls often mix multiple intents in the same sentence While pilots and POCs are easy, implementing Voice AI Agents at scale requires solving all these problems. Excited to be in trenches with Enterprises solving these challenges for them at scale.

  • View profile for Nikkitha Shanker

    Co-founder & CEO @SuperBryn | Building Evals & Observability for Voice Agents | 2x Founder

    25,402 followers

    I walked into the Cerebral Valley 𝗩𝗼𝗶𝗰𝗲 𝗦𝘂𝗺𝗺𝗶𝘁 thinking I had a decent read on where 𝗩𝗼𝗶𝗰𝗲 𝗔𝗜 is headed. I walked out realizing how many assumptions I need to revisit. Here's what stood out. It's a long one, but worth your time if you're building in this space. 𝗢𝗻 𝘁𝗵𝗲 𝗶𝗻𝘁𝗲𝗿𝗳𝗮𝗰𝗲 𝗶𝘁𝘀𝗲𝗹𝗳 ↳ Bret Taylor: GUIs are a historical accident. Conversational agents will bypass dashboards and act directly on underlying systems. Are we building for the past? ↳ Eugenia Kuyda: Attention is feed-based and glanceable. You can't scroll or fast-forward audio like text. Voice is a powerful layer, not the primary OS. ↳ Tanay Kothari: Voice won't go mainstream until it hits zero-edit accuracy. Absolute input reliability is the adoption threshold that actually changes behavior. 𝗢𝗻 𝘄𝗵𝗮𝘁 "𝗶𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝗰𝗲" 𝗺𝗲𝗮𝗻𝘀 𝗵𝗲𝗿𝗲 ↳ Justin Uberti: Cascaded STT→LLM→TTS pipelines hit a ceiling. Human conversation is continuous, not turn-based. End-to-end audio-native models catch non-verbal cues without unnatural pauses. We at @Superbryn are building with this insight too. ↳ Scott Stephenson Passing the "voice Turing test" means injecting emotional and conversational context into perception models. Text strips away laughter, sarcasm, tone. AI needs to perceive audio end-to-end to know how to respond, not just what to say. ↳ Russ d'Sa Voice agents are stateful, real-time programs demanding a new infrastructure paradigm. Humans have a 250ms response prior, so shaving every millisecond of latency is a psychological requirement to cross the uncanny valley. 𝗢𝗻 𝗯𝘂𝘀𝗶𝗻𝗲𝘀𝘀 𝗺𝗼𝗱𝗲𝗹𝘀 𝗮𝗻𝗱 𝘁𝗵𝗲 𝘀𝘁𝗮𝗰𝗸 ↳ Jake Saper, Olivia Moore, Grace Isford: The shift is from selling tools to selling outcomes. Voice agents that complete entire tasks (closing a mortgage, settling a claim) move the model from per-minute pricing to charging for successful task completion. ↳ Shiv (Shivdev) Rao: Voice is the "ultimate upstream data source." Encoding professional-client conversations is the foundational wedge to automate all downstream clerical, coding, and billing tasks. ↳ Dylan Fox: Trust in AI requires moving past skeuomorphic design. Consumers will prefer models that present as intelligent machines to be commanded directly, not ones that feign human emotion. ↳ Brandon Yang: Voice is the highest-bandwidth interface for complex, unstructured intention transfer. When delegating to agents, humans naturally "babble" and dump context they'd self-edit while typing. 𝗧𝗵𝗲 𝘁𝗵𝗿𝗼𝘂𝗴𝗵-𝗹𝗶𝗻𝗲: 𝗩𝗼𝗶𝗰𝗲 𝗔𝗜 is no longer a feature. It's a fundamental infrastructure shift that rewrites the stack, the business model, and the interface, all at once. Already building differently this week because of it. What's the voice interaction you're most excited to kill with a conversational agent? #VoiceAI #CerebralValley #ArtificialIntelligence #ConversationalAI

  • View profile for Rifat Bin Alam

    Product Leader | Building 0 to 1 AI Products | ex-Shikho, ShopUp, Unilever | Scaled AI to 3M+ users

    3,345 followers

    Most apps still expect users to learn the app before they can get anything done. But what if the user could simply say what they want? I have been exploring a small proof of concept around Bangla voice-led app automation. In the demo below, a user says: “Amar Airtel number e 30 taka recharge koro.” The system parses the Bangla voice command, understands the intent, identifies the operator and amount, opens the bKash app, and moves through the recharge flow. This is still an early demo, but it points to a bigger product shift. For years, we have designed apps around navigation. Buttons. Menus. Forms. Tabs. Screens. But as LLMs get better at understanding regional languages and converting unstructured text into structured actions, the interface can become much simpler. The user does not need to know where the feature lives. The user only needs to express intent. This is also why I find the Rabbit R1 story interesting. Rabbit tried to create a new hardware category around this idea. But the bigger opportunity may not require new hardware at all. It may come from apps and platforms exposing their services in ways that AI agents can safely act on top of the devices people already use. We are already seeing early signs of this in enterprise software. Salesforce has released its MCP Server in beta, which allows AI assistants and agentic IDEs to interact with Salesforce without the user clicking through the Salesforce UI. Of course, financial workflows need a much higher safety bar. Anything involving payments, PINs, recharge, send money, or banking actions needs strict permission handling, user confirmation, and security controls. But even with those caveats, the direction feels clear to me. The next major UX shift may not be a better menu. It may be removing the menu altogether. What do you think: will voice-led agents become a real interaction layer for most apps, or will trust and security concerns slow this down? #AI #AIAgents #Bangladesh #Fintech #ProductManagement #UXDesign #VoiceAI

Explore categories