Lesson 5: Multimodal AI and Agents (Non-Visual Productivity)
Course: Foundations of Non-Visual AI Productivity (AI Basics)
Lesson content
- Lesson Objective: By the end of this lesson, you will be able to use multimodal AI and agents safely in a non-visual workflow by requesting structured outputs, checking for likely errors, and asking for clear tool-selection rationale.
- What you will learn: two important trends that can boost non-visual productivity: multimodal AI (AI that can work with text + images + audio) and AI agents (AI that can plan tasks and use tools). This lesson keeps one core rule: speed is not enough - use verification-by-design.
-
NVAIP rule for multimodal and agents: always request:
- Structured output (headings, bullets, tables) so it is easy to review non-visually.
- Assumptions + uncertainty (“what might be wrong or missing”).
- A verification checklist (names, dates, numbers, actions taken).
-
Multimodal AI (what it is): a model that can interpret more than text. Depending on the product, it may be able to:
- Describe an image (objects, scene, layout, visible text).
- Explain a screenshot (what is on screen, what a dialog says, what a chart appears to show).
- Extract information (read visible text, list items, identify labels) when the image is clear.
- Work with audio (transcribe or summarize speech) if the tool supports it.
-
Multimodal AI (what it depends on): reliability changes based on:
- Image quality: blur, glare, low contrast, and small text reduce accuracy.
- Handwriting and complex layouts: whiteboards and messy notes can be difficult and often need manual confirmation.
- File access: an AI cannot “read your PDF” unless you paste the text, upload it into a tool that supports file reading, or the text is visible via OCR.
- Hallucination of presence (multimodal warning): sometimes AI may describe objects that are not there. For example, it may claim there is a “Submit” button on a blank page. Correction tip: If a description sounds unlikely, ask: “Describe the coordinates or the relative position of that item” to force the AI to look closer.
-
Multimodal NVAIP workflow (how to use it safely):
- Orient: state your goal (“I need the key message and any numbers in this chart”).
- Plan: request the output format (“give a short summary + a table of values”).
- Execute: ask for a first pass, then ask targeted follow-ups (“read only the labels and numbers”).
- Verify: spot-check critical details (names, numbers, dates). If possible, verify using the original text/data instead of only the image.
- Recover: correct mistakes and ask the AI to restate the corrected final output.
-
Example multimodal prompt (screenshot or chart):
Describe this image for a blind user. Then extract any visible text exactly as written. If there are numbers, put them in a table with columns: label, value, units. List anything you are unsure about. Finally, give me a 5-item checklist of what I should verify before using this information.
- AI agents (what they are): an agent is an AI system designed to plan steps and use tools (for example: search, calendars, email drafting, documents, data tools) to complete a task. Agents are not automatically “all-powerful.” They operate within the tools they are connected to and the permissions they are granted.
-
Agents (what they can do well in NVAIP):
- Planning: break a goal into steps and checkpoints (great for non-visual clarity).
- Drafting: write emails, agendas, checklists, and structured documents quickly.
- Tool-scoped actions: in some environments, they can help create events, update tasks, or organize information - usually with confirmation.
-
Agents (what to be careful about):
- Permissions: always know what the agent can access (documents, email, calendar, files).
- Action transparency: ask for a plan first and a log of actions taken.
- Agent “Reasoning” trace: ask for a clear explanation of why the agent chose a tool and what it planned to do. Prompt phrase: “Show me your step-by-step reasoning for why you chose this tool.” This is highly navigable for screen readers.
- Human-in-the-loop: for important tasks (sending emails, scheduling, deleting/editing records), require confirmation and verify the final outcome.
- Not universal clicking: most agents cannot reliably control any random app on your desktop. Automation depends on the specific tool and environment.
-
Example agent prompt (plan + confirmation):
Task: prepare a 30-minute meeting plan. First, ask me 3 questions if needed. Then give a step-by-step plan. Do not take actions. After I confirm, draft the invite text, agenda, and a checklist of what I should verify (date, time, timezone, attendees).
-
Privacy-first NVAIP (must-follow habits):
- Minimize sensitive data: avoid uploading client data, IDs, medical, financial, or confidential documents unless you have permission and the tool is approved.
- Redact when possible: remove names, email addresses, and identifiers if not needed for the task.
- Use templates: ask AI to generate templates and structure, then fill sensitive details yourself.
- Verify before sharing: treat AI output as a draft and check the final version before sending or publishing.
-
Mini exercise (5 to 7 minutes):
- Multimodal practice: take one non-sensitive screenshot (a simple chart or a page with headings). Ask the multimodal prompt above and review the extracted text and table.
- Verification: spot-check 2 critical details from the image (a number and a label). Correct any mistakes and ask the AI to restate the final output.
- Agent practice: ask an agent (or AI chat) to create a plan for a real task you have today. Require: plan first, then verification checklist, then draft output.
- Key takeaway: multimodal AI and agents can be powerful for non-visual productivity, but success comes from structured prompts + clear constraints + verification-by-design + privacy discipline.
Continue
All lessons in this course
References
- 1. OpenAI, “GPT-4 Technical Report” (2023). https://arxiv.org/abs/2303.08774
- 2. OpenAI, “GPT-4V(ision) System Card” (2023). https://cdn.openai.com/papers/GPTV_System_Card.pdf
- 3. Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models” (2022). https://arxiv.org/abs/2210.03629
- 4. Park et al., “Generative Agents: Interactive Simulacra of Human Behavior” (2023). https://arxiv.org/abs/2304.03442
- 5. Stanford HAI, “AI Index Report” (latest edition). https://aiindex.stanford.edu/report/