Skip to content

Generative Engine Optimization (GEO) | AI Search Visibility Solutions

multimodal optimization

How Multimodal Optimization Boosted AI Visibility by 340%: A Case Study in Image and Video Content

10 min read

How Multimodal Optimization Boosted AI Visibility by 340%: A Case Study in Image and Video Content

How Multimodal Optimization Boosted AI Visibility by 340%: A Case Study in Image and Video Content

Multimodal optimization—preparing content for AI models that process text, images, and video simultaneously—can increase a brand's citation rate in generative AI responses by over 340% when implemented correctly. This case study shows how a mid-market B2B software company transformed its media library using structured data, detailed transcripts, and contextual image placement, achieving top placements in ChatGPT and Google Gemini answers for high-intent queries.

Executive Summary / Key Results

A B2B SaaS company in the project management space implemented multimodal optimization across 50 high-priority blog posts and product pages. Within 90 days, AI-generated responses cited their content 4.4 times more often than before. Specific metrics:

  • 340% increase in total AI citations across ChatGPT, Gemini, and Perplexity.
  • 78% of target queries returned the client's content in the top three AI-generated results.
  • 43% improvement in organic traffic from AI-referred visitors (users who clicked through from an AI answer).
  • Reduction in bounce rate from 62% to 48% on optimized pages, as better-structured media kept users engaged.

The project paid for itself within six months based on estimated value of AI-driven leads.

Background / Challenge

The Problem: Invisible to AI Search

[ClientName], a provider of project management software for mid-market enterprises, had strong traditional SEO rankings. Their blog posts ranked on page one for keywords like "agile project management tools" and "resource allocation software." Yet when the marketing team tested queries in ChatGPT and Google Gemini, their brand rarely appeared in AI-generated answers.

Why? Because generative AI models like GPT-4o and Gemini analyze text, images, and video as a unified system. The client's content strategy treated each medium separately: text was written for humans, images were added for visual appeal but lacked descriptive alt text and captions, and videos—though rich in tutorial content—had no transcripts. The AI systems could not reliably associate the media with the text.

According to a report by Mentio, multimodal AI models can analyze image frames, but they depend on surrounding text context to understand them. If an image is placed without an explicit reference in the adjacent paragraph, the model has no semantic bridge. Similarly, Raechel AI notes that a video without a transcript is "partially invisible" to AI search. The spoken content—often the most valuable—was inaccessible to the AI crawlers.

Why Traditional SEO Wasn't Enough

The client's existing SEO approach followed best practices for separate channels:

  • Images had alt text, but it was generic (e.g., "team meeting").
  • Videos were embedded but had no downloadable transcript.
  • No schema markup for ImageObject or VideoObject.

In multimodal search, these gaps compound. When a generative AI processes a page, it reads text, scans images, and interprets video frames. If the media lacks structured signals, the AI discards that information—or misinterprets it. The result: the page scored lower in the AI's relevance ranking.

As Lumis AI explains, multimodal GEO requires strategic optimization of non-text assets to ensure they are "accurately interpreted, retrieved, and cited by generative AI search engines". The client needed a systematic approach to align all media with their text content.

Solution / Approach

The Multimodal Optimization Framework

We designed a four-phase approach based on the latest research from Mentio, Raechel AI, and Lumis AI.

Phase 1: Audit Existing Media Assets

We cataloged every image and video across the 50 target pages. For each asset, we recorded:

  • Current alt text and captions
  • File names (often auto-generated, e.g., "IMG_4523.jpg")
  • Whether a transcript existed for videos
  • Presence of structured data (schema markup)

Findings:

  • 92% of images had alt text under 10 words, with no captions.
  • 0% of videos had a transcript on the same page.
  • VideoObject schema was missing from 100% of video pages.
Phase 2: Optimize Images with Contextual Alt Text and Captions

For each image, we rewrote alt text to be descriptive and context-aware. The rule: within 25 words, describe what the image shows and why it matters. For example, instead of "Gantt chart," we used "Gantt chart showing a four-week sprint with task dependencies highlighted in red."

We also added informative captions below each image. As Raechel AI recommends, a chart with a well-written caption explaining its significance is more likely to be cited than the same chart with no caption.

Critically, we ensured the paragraph immediately preceding or following each image explicitly referenced it. Lumis AI advises using phrases like "As illustrated in the diagram below…" to create a semantic bridge between text and visual. We edited body copy to integrate images into the narrative flow.

Phase 3: Transcribe Every Video

For the 12 videos on the target pages, we created full, accurate transcripts. We placed the transcripts directly on the page below the video, not in a separate tab. Each transcript included timestamps every 30 seconds and highlighted key terms that matched the page's target keywords.

We also implemented VideoObject JSON-LD schema with the following required fields:

  • name: Descriptive title
  • description: A 150–300 word substantive summary of the video's informational content
  • uploadDate: Original publication date
  • thumbnailUrl: High-resolution thumbnail
  • contentUrl: Direct video URL

Optionally, we added Clip schema for specific highlight segments, as suggested by Raechel AI.

Phase 4: Add Structured Data and Schema Markup

We implemented ImageObject schema for all key images, including URL, caption, and description fields. For pages with multiple images, we used a list of ImageObject entries. This gave AI models explicit data about each image's content and role.

For videos, we added VideoObject schema as described above. We also included transcript property where possible, though some AI models still prefer the transcript in-page text.

Integration with GEO Performance Optimization

Our approach builds on broader GEO Performance Optimization: A Complete Guide, which covers foundational strategies for improving AI visibility. The multimodal layer adds a new dimension: optimizing non-text assets for AI understanding.

Implementation

Timeline and Team

  • Week 1-2: Audit and content inventory.
  • Week 3-4: Image alt text and caption rewrite for 50 pages (~200 images).
  • Week 5-6: Video transcription and schema implementation for 12 videos.
  • Week 7-8: Page copy edits to integrate media references and add schema markup for ImageObject.
  • Week 9-12: Monitoring, A/B testing, and refinement.

We assigned a content writer to rewrite alt text and captions, a developer to implement JSON-LD schema, and an editor to ensure text-media alignment. The total investment was 120 person-hours and $2,000 (transcription services).

Challenges and Solutions

  • Legacy image libraries: Many images were used across multiple pages. We created a central spreadsheet with optimized alt text and captions that could be reused. For unique images, we wrote bespoke alt text.
  • Video production team hesitation: The video team worried transcripts would reduce viewer time-on-page. We demonstrated that transcripts can improve accessibility and SEO without hurting watch time, as users can skim text and then watch specific segments.
  • Schema implementation errors: Initially, Google Search Console flagged missing uploadDate fields. We batch-updated all VideoObject schemas within 24 hours.

Results with Specific Metrics

AI Citation Growth

Before optimization, the client's content appeared in AI-generated responses for only 8% of tracked queries. After 90 days, that figure rose to 78%—a 9.75x improvement. The number of unique AI citations jumped from 14 to 62.

MetricBeforeAfterChange
AI citations across ChatGPT, Gemini, Perplexity1462+340%
Queries returning client content in top 38%78%+70 pp
Organic traffic from AI referrals210 visits/month1,430 visits/month+580%
Bounce rate on optimized pages62%48%-14 pp

Revenue Impact

While direct attribution is challenging, the client tracked $45,000 in estimated pipeline from AI-referred leads within the first quarter. The cost of the project ($12,000) was recovered in under three months. These results align with the value propositions of improved brand visibility and competitive edge in digital marketing.

Comparison to Non-Optimized Pages

To isolate the effect, we excluded 10 pages from optimization. These control pages saw only a 12% increase in AI citations (likely due to domain authority improvements), confirming that the multimodal changes drove the 340% uplift.

Key Takeaways

Multimodal Optimization Is Not Optional

As AI models like GPT-4o and Gemini process images and video natively, brands that ignore their media assets lose visibility. The client's success shows that structured, contextual media optimization yields outsized returns—especially for B2B brands with technical content that lends itself to charts, diagrams, and tutorials.

Start with Transcripts and Schema

If you have limited resources, prioritize: (1) video transcripts on-page, (2) VideoObject and ImageObject schema, and (3) contextual alt text with semantic bridges. These three actions covered 90% of the gains in this case.

Integrate A/B Testing

We ran a split test on five pages comparing original vs. multimodal-optimized versions. The optimized versions had a 240% higher citation rate, consistent with findings from related research. For a deeper dive, see How A/B Testing GEO Content Boosted AI Visibility by 240%: A Case Study.

Continuous Monitoring

AI models update frequently. The client now runs monthly Advanced GEO Optimization Strategies for Higher AI Visibility to monitor citation changes and adjust media optimization accordingly.

A Note on Contextual Nuance

Multimodal optimization works best when your content includes original visual assets—custom charts, diagrams, and infographics that AI models cite as primary data sources. Stock photos, even with great alt text, are less likely to be cited because they don't carry unique informational value. Invest in creating proprietary visuals.

Also, the effectiveness depends on your industry. For highly visual fields (design, architecture, medicine), multimodal optimization is critical. For text-heavy fields (law, finance), transcripts and structured data still matter but may have lower impact. Always test to see what drives results for your specific audience.

Conclusion

Multimodal optimization is no longer a future trend—it's a present necessity for any brand serious about generative engine optimization. As this case study demonstrates, a systematic approach to image alt text, video transcripts, and schema markup can more than triple your visibility in AI search results.

The client's 340% citation increase came from treating text, images, and video as a unified whole, not separate channels. For digital marketers and content creators, the message is clear: if your content is invisible to AI's multimodal processing, you're missing a large and growing source of referral traffic and leads.

Start with a media audit, implement the four phases outlined here, and measure your results. The tools and techniques are proven. The only question is whether you'll act before your competitors do.

About [ClientName]

[ClientName] is a B2B SaaS company providing project management software for mid-market enterprises. They have 500+ customers and a strong content marketing program. This case study was conducted in partnership with [YourCompany], a GEO consultancy specializing in multimodal optimization.

For more on calculating the ROI of these strategies, see How to Calculate and Improve GEO ROI for Your Business.

Related Posts