Alibaba's Qwen-Image-3.0 is pitched at one thing: making generated images useful enough to actually ship. It renders text as small as 10 pixels, takes prompts up to 4.5k tokens, and laid out a full 3x3 grid of distinct infographics (a math slide, a parasitology explainer, a bank control diagram) in a single pass across 12 languages. The honest caveat: it launched with no open weights, no model card, and no benchmarks, and access is an invite-only API bound for Qwen Chat, so for now you are trusting the cherry-picked example gallery until you can run your own prompts.
fal.ai released an API for the CVPR 2025 Video Depth Anything model. It generates temporally consistent depth maps from video, with three model sizes (Small, Base, Large), five colormaps, side-by-side comparison with the original video, and raw depth export as .npz. At $0.04 per second of video, it's a practical tool for 3D reconstruction, video effects, compositing, and scene understanding.
A tested guide provides 9 copy-paste AI travel photo prompts using a five-part structure that transforms ordinary selfies into travel photos. Only 5% of surveyed travelers could distinguish AI from real photos. The prompts work best in GPT-Image-2 Thinking mode, and the guide emphasizes honest labeling when posting.