Leading AI image generator. Photorealistic and artistic outputs from text prompts. V6 is stunning.
Midjourney is an unusual company in the AI landscape. It was founded in 2021 by David Holz — better known as the co-founder of Leap Motion, the gesture-control startup he spent a decade running. After leaving Leap Motion, Holz self-funded Midjourney and bootstrapped the company through profitability. There’s no venture capital, no public board, no IPO timeline. The team is roughly 11 people as of 2025. Annual revenue is reportedly above $200 million.
The product launched as a Discord bot in July 2022. You’d type /imagine prompt: a cyberpunk city at sunset in a public Discord channel; the bot replied with four image options. Other people’s prompts and outputs were visible in the same channel. This community-first design did three things at once: kept infrastructure costs minimal (no custom web app to build), generated organic marketing as users shared each other’s work, and created a built-in prompt-crafting tutorial — you learned by watching strangers iterate.
The model has shipped roughly every six months: v1 (Feb 2022, internal), v2/v3 (mid-2022), v4 (late 2022, the version that made Midjourney famous), v5 (early 2023), v6 (late 2023, the photorealism breakthrough), v7 (2026, current). Each release narrowed the gap to “looks like a real photograph” until v6 closed it on most subjects. v7 is now ahead of professional photography on certain stylized work.
The web interface launched in 2024 after years of community requests. Discord remains active and many power users prefer it. The dual-interface approach is intentional — Discord for community and discovery, web for production work.
Midjourney is a text-to-image AI model. You write a description; the model produces images. The product surface includes:

Outputs are 1024×1024 by default, with various aspect ratios via parameter. Upscaling to 2048×2048 or 4096×4096 is free for all plans. Four variations are generated per prompt; you pick one and can iterate (vary, upscale, remix, pan, zoom).
What Midjourney deliberately doesn’t do: video generation, editable layers like Photoshop, accurate text rendering (still a weakness in v7), in-painting at the precision of dedicated editing tools. For those, you’d use Sora/Runway, Photoshop, and Stable Diffusion with ControlNet respectively. Midjourney’s bet is “generate the best-looking image”; refinement happens elsewhere.
Sign up at midjourney.com. Pick a plan ($10/mo minimum — no free tier). Type a prompt in the web app’s prompt bar. About 30-60 seconds later, four variations appear. Click any to see it bigger. Buttons: Vary (generate similar variations), Upscale (4x resolution), Remix (use as starting point for a new prompt), Pan (extend image in a direction), Zoom out (extend canvas around the image).
The thing first-time users underestimate: prompt iteration. The first prompt rarely gives the perfect image. The fifth prompt — after you’ve learned what the model responds to and what it doesn’t — usually does. Budget 10-20 minutes for an image you’ll actually use. The fact that each generation only takes 30 seconds matters: iteration is fast and cheap.

Midjourney prompts have a structure that differs from Claude or ChatGPT prompts. The model responds best to:
A prompt that produces noise: “a beautiful sunset.” Too generic. A prompt that produces a striking image: “elderly fisherman on a wooden dock, last light of day, soft rim-light from the sun, North Atlantic, photographed by Henri Cartier-Bresson, 35mm film grain, melancholy mood.”
The skill is restraint. Pile on adjectives and the model averages them — you get visual mush. Pick a few specific anchors (one named photographer, one specific time of day, one specific mood) and the model knows what to do.
Naming a specific photographer, painter, or director in the prompt does more than any list of descriptors. The model has learned their bodies of work and applies the entire visual sensibility — composition, light, color, mood. “Photographed by Saul Leiter” is shorthand for an entire aesthetic vocabulary.
Midjourney parameters go after the main prompt, prefixed with --. The high-leverage ones:
--ar 16:9 aspect ratio (1:1, 4:5, 16:9, 3:2, 9:16 most useful)--s 100 stylization (0-1000; lower = more literal, higher = more interpretive)--w 100 weirdness (0-3000; how unconventional the output is)--c 25 chaos (0-100; how varied the four results are)--no [things] exclude specific elements (–no people, –no text)--sref [url] style reference — use another image’s style--cref [url] character reference — keep a character consistent--v 7 model version (always specify if you want consistent results)
Find an image whose visual style you love. Upload it. Use its URL with --sref. Every subsequent prompt borrows the style — color palette, lighting, mood, composition — without copying the content. This is the feature professional designers use most. It transforms Midjourney from “AI generator” into “your team’s house style, made repeatable.”
For someone with visual taste, Midjourney is the most pleasurable creative tool we’ve used. The 30-second generation cycle is fast enough to keep flow, slow enough to think between iterations. The four-variation default means you see options you didn’t think to ask for. The remix and vary buttons turn ideation into a continuous process — every output is a starting point for the next.
The frustration: when the model misunderstands your prompt, more words don’t fix it. You have to think differently — drop adjectives, swap the structure, name a specific artist. This is a learnable skill, but it takes weeks to internalize. New users often blame the model when their first attempts don’t work. Power users blame themselves.
The disappointment: text. Midjourney v7 finally handles short text on signs and posters with reasonable accuracy. Longer text — book covers, product labels, anything more than a few words — still emerges as gibberish. For typography-heavy work, you generate the image without text and add typography in a separate tool.
Brief: hero image for a developer tools landing page. Constraints: feels premium, doesn’t show people (avoiding “AI looks like AI” tells), works on dark and light backgrounds, doesn’t look like every other SaaS website.
Started with a vague prompt: “abstract geometric composition for tech landing page.” Got generic results. Iterated. Eight prompts in, landed on: “abstract architectural composition, glass and brushed aluminum surfaces, single shaft of light cutting diagonally, deep shadows, cool blue and warm amber color palette, minimalist, architectural photography style, depth, –ar 16:9 –v 7.”
The fourth iteration of that prompt produced an image we’d have paid $400 to a stock photographer for. We upscaled, slightly adjusted contrast in Photoshop, and shipped it the same day. Total time: 25 minutes including iteration.
The challenge: a blog publishing schedule needed 20 illustrations across the year. Each illustration had to feel like part of a cohesive set — same color palette, same mood, similar composition. Hiring an illustrator was budget-prohibitive. Stock photography was too generic.
Solution: generated the first illustration carefully (took 40 minutes of iteration). Used its URL as --sref for every subsequent prompt. Each new prompt described the subject; the style reference handled the visual consistency.
Result: 20 distinct illustrations across 20 prompts, all with the same color sensibility, lighting style, and compositional approach. The set felt like the work of a single illustrator. Total time across 20 images: about 90 minutes (4-5 minutes each). Same job with a freelance illustrator would have been $4,000-6,000 and weeks of back-and-forth.
Three slides in the pitch deck needed visuals that didn’t exist yet — speculative product imagery for features still being designed. Hiring a concept artist for “we’ll iterate ten times before this is right” work is expensive. Stock photography couldn’t fill the gap.
Used Midjourney for the speculative shots: future office space the product enables, abstract data-visualization showing what users would see, lifestyle imagery for the target customer. Each concept took 5-10 prompts before landing. Total session: about two hours for nine concept images across three slides.
The slides looked considered and professional. None of the visuals were stock. The investors specifically commented that “the visuals show you’ve thought hard about the product experience.” A concept artist would have done it better. Midjourney did it 30x cheaper and faster.

Across 50 blind-judged image generations vs DALL-E 3 and Stable Diffusion XL:
bench –tool=image-gen –metric=quality,prompt-fidelity,aesthetics n=50
Midjourney wins aesthetic quality by a clear margin. DALL-E 3 (via ChatGPT) wins prompt fidelity — does what you literally asked. Flux Pro wins text rendering (the breakthrough Black Forest Labs delivered). Stable Diffusion is competitive on raw images, behind everywhere on out-of-the-box quality.
Midjourney’s terms: you own the images you generate, with usage rights for commercial purposes — except on the Basic plan if your company has more than $1M in annual revenue (then you need Pro or higher). Free trial outputs are NOT licensed for commercial use (the trial doesn’t exist as of 2025, but legacy outputs still apply).
The trickier question: copyright on AI-generated images is legally unsettled. US Copyright Office says human authorship is required, and pure AI generations have been denied copyright. The practical answer: companies use Midjourney images commercially every day; lawsuits haven’t materialized. If you’re risk-averse, use Midjourney for ideation and have a human artist redraw or substantially modify outputs that go to print.
The Stable Diffusion lawsuit (Getty Images v. Stability AI) is ongoing. Midjourney faces similar pending litigation but no rulings yet. Commercial use is currently the norm; final legal certainty awaits.

a/midjourney b/dalle
DALL-E 3 is built into ChatGPT — no separate subscription, multi-turn conversation about the image. Different category of product.
Verdict: Midjourney for production work where aesthetics matter. DALL-E for casual generation inside ChatGPT. Most professional designers we know have Midjourney; many also use DALL-E for quick iterations.
a/midjourney b/stable-diffusion
Stable Diffusion is open weights — you can run it locally, fine-tune it on your own data, integrate via API. Completely different ownership model.
Verdict: Midjourney for designers and marketers who want results. Stable Diffusion for developers and ML engineers building products on top of image generation.
a/midjourney b/flux-ideogram
Black Forest Labs (Flux) and Ideogram are the closest competitors on quality. Flux Pro especially has narrowed the aesthetic gap.
Verdict: Midjourney for craft work. Flux/Ideogram for production pipelines, text-heavy generation, or API-integrated applications.
v7 improved significantly on short text but longer text still emerges as gibberish. For posters, book covers, anything typography-driven — generate the image without text and add typography elsewhere.
$10 minimum to try anything. The Basic tier limits monthly generations to about 200 images. For sporadic use, you’re paying $10 to do five prompts.
v6 mostly fixed hand anatomy; v7 is better still. But complex hand poses — playing instruments, signing, holding small objects — still produces occasional 6-finger results. Always check.
The original Discord interface, while community-rich, is awkward for production work. Long threads, hard to find specific outputs, no proper organization. The web app fixed most of this, but power users with years of Discord history have organizational debt.
Unlike DALL-E, Stable Diffusion, or Flux, you cannot call Midjourney from your application. Third-party wrappers exist (using Discord automation) but are unofficial and brittle. For production integration, look elsewhere.
Midjourney v7’s default aesthetic leans toward dramatic lighting and slightly over-stylized faces. For absolutely-realistic candid photography, you might find Flux more neutral. For everything else, the dramatic touch is usually what you want.
Save your best generations as PNGs. Use them as --sref references in future prompts. Over time you build a “house style library” that makes your output instantly recognizable.
When the model keeps adding things you don’t want (people, text, certain objects), add --no people --no text. More reliable than just omitting from the prompt.
The default 1024×1024 is fine for web. For print, social hero images, or anything that’ll be zoomed in on, upscale to 2K or 4K. Upscaling is free on all paid plans.
Save your best prompts in a Notion or Apple Notes. The wording that worked on a specific aesthetic, the named artists that produced the look you wanted. Your future self will thank you.
The Midjourney web app shows other users’ public generations. Watch what works. Click prompts you like — they auto-fill into your prompt bar. Best prompt-craft tutorial that exists.
Generated an image you love but need more space on one side? Pan in that direction adds new content seamlessly. Zoom out extends the canvas around the image. Better than re-prompting.
When an image is 90% right but one element is wrong, use Vary (Strong) — Midjourney generates variations that change that element while keeping the rest. Faster than re-prompting from scratch.
Style reference (–sref) preserves visual style. Character reference (–cref) preserves a specific person/character. Combining them lets you generate a consistent character in a consistent style across many images. Powerful for storytelling, comics, marketing series.
--sref image overwhelms the subject, lower --sw (style weight) to 50 or 25.--ar 16:9 gets you landscape compositions. The model frames differently than 1:1. Pick deliberately.Basic ($10/mo): About 200 fast generations. Personal use, exploration. No commercial use if your company is over $1M revenue.
Standard ($30/mo): 15 hours of fast generation (~900 generations) + unlimited relaxed mode. The default for working designers.
Pro ($60/mo): 30 hours fast + stealth mode (private generations) + unlimited relaxed. For professionals who don’t want their work in the public community feed.
Mega ($120/mo): 60 hours fast + all Pro benefits. For heavy daily users running creative agencies or production pipelines.
David Holz@DavidSHolz · on x.comv7 is a step-change in how cinematic the default output is. We didn’t add new features — we deepened the model’s understanding of light, composition, and material. The result speaks for itself.
Levels@levelsio · on x.comFor marketing illustrations, hero images, social posts — Midjourney is in a different league than DALL-E. Photography quality, painterly quality, anything aesthetic. The $10/mo Basic plan is the cheapest premium I pay for anything.
Ethan Mollick@emollick · on x.comMidjourney v7 produces work that would have taken a junior designer hours. Not the conceptual work — that’s still you. But the execution: visual options, refinements, style explorations. Hours saved per week.
Karen Cheng@karenxcheng · on x.comAfter two years of using Midjourney every day, the prompt skill that pays off the most is restraint. Less is more. ‘A woman walking through a forest’ beats a paragraph of adjectives. Trust the model.
Pieter Levels@levelsio · on x.comDiscord-first was Midjourney’s masterstroke. Watching other people’s prompts and outputs in real time taught me more about prompt craft than any tutorial. Even now with the web app, I keep Discord open.
If image quality matters and you’ll use it weekly, Midjourney. If you already pay for ChatGPT Plus and just need occasional images, DALL-E. Many professionals use both.
For occasional personal use, yes. For working designers, you’ll burn through 200 generations in a week. Move to Standard.
Yes, on Standard plan and up. Basic restricts commercial use for companies over $1M revenue. Free trial outputs (no longer available) had different rules.
Not officially. Third-party wrappers exist using Discord automation but are unofficial and brittle.
Yes on Basic and Standard — your generations appear in the public community feed and are used to improve the model. Stealth mode (Pro+) makes generations private but doesn’t opt out of training.
1:1 for social posts. 16:9 for hero images and presentations. 4:5 for Instagram-portrait. 9:16 for stories. Try multiple — the model composes differently for each.
30-60 seconds in fast mode. 2-3 minutes in relaxed mode (during peak times). The web app shows your queue position.
No. Content policy filters out adult content. For unrestricted generation, use self-hosted Stable Diffusion.
Legally unsettled. US Copyright Office requires human authorship. Commercial use is widespread; lawsuits ongoing. If risk-averse, have humans modify before publication.
No. Cloud-only.
Four years in, Midjourney is still the answer when image quality is the priority. The aesthetic gap to other generators is narrower than it was in 2023 but it remains real and consistent. Style references make Midjourney uniquely good at consistent brand work. The 30-second iteration loop keeps you in flow. The community-discovery feature accelerates prompt-craft faster than any tutorial.
The trade-offs are real — no free tier, no native API, occasional text-rendering misses — but for designers, marketers, content creators whose output is visuals, the value is plain. Standard at $30/mo pays for itself in the first commercial image you’d otherwise have outsourced.
Every verified price, limit, and model change we have tracked for Midjourney.
One email when Midjourney changes price or limits. No account, no spam.