Trying world Labs Is Joining AMD: a practical beginner's guide

World Labs is teaming up with AMD to push 3D spatial AI forward. Here is what the partnership means and how to build your first 3D spatial pipeline today.

By Asahluma Tyika

What happened with World Labs and AMD (and why you should care)

Howzit, developers! If you follow the AI space, you probably noticed the massive buzz around Dr. Fei-Fei Li’s startup, World Labs, and its strategic collaboration and backing with AMD.

For the past two years, generative AI has been obsessed with flat screens: text prompts turning into 2D JPEG images or flat MP4 video loops. World Labs is taking a radically different approach known as Spatial Intelligence—building “Large World Models” (LWMs). These models do not just generate pixels; they understand geometry, depth, physics, camera perspective, and 3D structure.

AMD stepping in to power and co-develop with World Labs signals something big for developers: spatial 3D AI is moving off proprietary, locked-down hardware silos and landing straight into open compute platforms like AMD ROCm and Ryzen AI.

Whether you are building indie games in Godot, doing computer vision for robotics, or just tired of regular chatbots, this guide cuts through the venture capital jargon. We will unpack what world models actually are, look at how the AMD compute stack handles them, and walk through a fully runnable Python project that turns a flat 2D image into an interactive 3D point cloud scene on your local machine.


Part 1: 2D Generative AI vs. World Models

To understand why AMD is putting serious weight behind World Labs, you need to understand the structural difference between generating an image and generating a “world.”

When you ask a standard diffusion model for a picture of a kitchen chair, it predicts which colored pixels look plausible next to each other. It has no idea the chair has a back leg hidden behind the front cushion. If you rotate the virtual camera, the illusion falls apart.

A world model estimates the actual coordinates, continuous surfaces, and volume of the scene.

Feature Standard 2D GenAI (Stable Diffusion, Midjourney) Spatial World Models (World Labs, NeRF, 3DGS)
Primary Output 2D pixel grid (.png, .webp) 3D representations (Meshes, Point Clouds, Gaussian Splats)
Physical Awareness None (purely statistical visual patterns) High (understands occlusion, scale, perspective, lighting)
Camera Freedom Fixed single viewpoint Full 6-DoF (Degrees of Freedom: pitch, yaw, roll, X, Y, Z)
Compute Focus Tensor matrix multiplication on 2D latents Volumetric rendering, ray-marching, and spatial geometry

World Labs aims to let developers generate persistent, editable 3D environments from text or single photos. AMD’s goal is to ensure those heavy volumetric calculations run efficiently across AMD Instinct data center GPUs down to local Radeon graphics cards using the ROCm software stack.


Part 2: The AMD ROCm Angle for Python Developers

Historically, running advanced AI meant using Nvidia’s proprietary CUDA framework. AMD’s ROCm (Radeon Open Compute) has matured rapidly to break that lock-in.

For Python developers, the best part about AMD’s current ROCm ecosystem is that your code doesn’t need to change. PyTorch compiled with ROCm uses the exact same torch.cuda API calls. When you see torch.cuda.is_available() in a script, ROCm translates those calls directly to AMD hardware via its HIP (Heterogeneous-Compute Interface for Portability) runtime layer.

# Example: Installing ROCm-enabled PyTorch on Linux
pip install torch torchvision --index-url https://download.pytorch.org/whl/rocm6.1

If you are running on Windows, Apple Silicon, or a CPU, do not stress. The practical implementation below runs cleanly on any hardware by falling back to standard CPU tensors if an AMD ROCm GPU is not active.


Hands-on example: Building your first 2.5D spatial depth scene

Let’s build a working prototype of how spatial models interpret depth and 3D position from a 2D image. We will use a lightweight monocular depth estimation pipeline combined with Open3D to project 2D pixels into true 3D Euclidean coordinates $(X, Y, Z)$ and export an interactive point cloud.

1. Set up your environment

Open your terminal and install the required dependencies:

pip install torch torchvision numpy pillow open3d

2. Create the script: spatial_world_gen.py

Create a new file named spatial_world_gen.py and paste the following complete, runnable code:

import urllib.request
from PIL import Image
import numpy as np
import open3d as o3d
import torch
import torchvision.transforms as transforms

def select_compute_device() -> torch.device:
    """Detects ROCm/CUDA acceleration or falls back cleanly to CPU."""
    if torch.cuda.is_available():
        device_name = torch.cuda.get_device_name(0)
        print(f"[Compute] Hardware acceleration detected: {device_name}")
        return torch.device("cuda")
    print("[Compute] Acceleration not found. Running on CPU.")
    return torch.device("cpu")

def download_sample_image(filename: str = "room.jpg") -> str:
    """Downloads a royalty-free test image of an indoor room."""
    url = "https://images.unsplash.com/photo-1513694203232-719a280e022f?w=640"
    print(f"[Data] Fetching test image from {url}...")
    req = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})
    with urllib.request.urlopen(req) as response, open(filename, "wb") as out_file:
        out_file.write(response.read())
    return filename

def estimate_spatial_depth(image_path: str, device: torch.device):
    """Loads a vision model to calculate estimated depth (Z dimension)."""
    print("[Model] Loading MiDaS depth model from PyTorch Hub...")
    model = torch.hub.load("intel-isl/MiDaS", "MiDaS_small", trust_repo=True)
    model.to(device)
    model.eval()

    # Load appropriate transform
    midas_transforms = torch.hub.load("intel-isl/MiDaS", "transforms", trust_repo=True)
    transform = midas_transforms.small_transform

    raw_image = Image.open(image_path).convert("RGB")
    input_batch = transform(np.array(raw_image)).to(device)

    print("[Model] Estimating spatial geometry...")
    with torch.no_grad():
        prediction = model(input_batch)
        # Resize prediction back to original image size
        prediction = torch.nn.functional.interpolate(
            prediction.unsqueeze(1),
            size=raw_image.size[::-1],
            mode="bicubic",
            align_corners=False,
        ).squeeze()

    depth_map = prediction.cpu().numpy()
    return raw_image, depth_map

def generate_3d_point_cloud(raw_image: Image.Image, depth_map: np.ndarray, output_ply: str = "world_scene.ply"):
    """Projects 2D image pixels + estimated depth into a 3D point cloud."""
    print("[Spatial] Projecting pixels into 3D Cartesian coordinates...")
    width, height = raw_image.size
    rgb_data = np.array(raw_image) / 255.0

    # Normalize depth map between 0.1 and 1.0 to avoid division by zero
    depth_normalized = (depth_map - depth_map.min()) / (depth_map.max() - depth_map.min() + 1e-6)
    depth_metric = 1.0 / (depth_normalized + 0.1)

    # Approximate pinhole camera intrinsics
    focal_length = width
    cx, cy = width / 2.0, height / 2.0

    # Generate pixel grid
    u, v = np.meshgrid(np.arange(width), np.arange(height))

    # Back-project 2D (u, v, depth) to 3D (X, Y, Z)
    z = depth_metric
    x = (u - cx) * z / focal_length
    y = (v - cy) * z / focal_length

    # Flatten arrays to list of vertices
    points = np.stack((x, -y, -z), axis=-1).reshape(-1, 3)
    colors = rgb_data.reshape(-1, 3)

    # Subsample to keep memory lightweight for beginners
    step = 4
    subsampled_points = points[::step]
    subsampled_colors = colors[::step]

    # Create Open3D PointCloud object
    pcd = o3d.geometry.PointCloud()
    pcd.points = o3d.utility.Vector3dVector(subsampled_points)
    pcd.colors = o3d.utility.Vector3dVector(subsampled_colors)

    # Save to disk
    o3d.io.write_point_cloud(output_ply, pcd)
    print(f"[Success] Saved 3D point cloud file to '{output_ply}'.")
    return pcd

if __name__ == "__main__":
    compute_dev = select_compute_device()
    image_file = download_sample_image()
    rgb_img, depth = estimate_spatial_depth(image_file, compute_dev)
    point_cloud = generate_3d_point_cloud(rgb_img, depth)
    
    print("\nDisplaying interactive 3D window (Close window to exit)...")
    o3d.visualization.draw_geometries([point_cloud], window_name="Code easier! | Spatial 3D Preview")

3. Run the script

Execute the file from your terminal:

python spatial_world_gen.py

When the script finishes downloading the model weights and running inference, an interactive window will pop up. Left-click and drag to rotate your camera in real 3D space, right-click and drag to pan, and scroll your mouse wheel to zoom in and out. You are now inspecting geometry reconstructed from a single flat picture.


Common mistakes / troubleshooting

  • GUI Window crashes or fails to open (GLFW Error): If you are running inside WSL2 or an SSH session without an X-server, Open3D’s draw_geometries will fail.
    • Fix: Comment out the o3d.visualization.draw_geometries line and inspect the generated world_scene.ply file using external free desktop tools like MeshLab or Blender.
  • PyTorch Hub SSL Certification Errors: On older macOS or Linux setups, downloading the MiDaS weights may throw an SSL certificate error.
    • Fix: Run python -m pip install --upgrade certifi or load Python’s built-in ssl._create_default_https_context = ssl._create_unverified_context at the very top of your test script.
  • Point cloud looks inverted or like a spiky wall: This happens when depth values are uncalibrated or inverted.
    • Fix: Check the normalization line. Inverted depth produces a hollow “cave” effect rather than an outward room structure; invert depth_normalized (1.0 - depth_normalized) if your model predicts disparity instead of relative distance.
  • ROCm device not recognized on AMD Radeon GPUs: PyTorch defaults back to the CPU even if you have an AMD GPU installed.
    • Fix: Ensure your user belongs to the render and video Linux groups via sudo usermod -aG render,video $USER, and confirm your GPU architecture with the rocminfo terminal command.

Try it yourself

Right now, our pipeline samples every 4th pixel (step = 4). Open spatial_world_gen.py and modify the script to:

  1. Change step = 1 for a dense, high-resolution spatial reconstruction.
  2. Replace the test URL with a photo of your own desk or workstation.
  3. Observe how the reconstructed 3D point cloud handles thin objects (like monitor stands and keyboards) versus broad surfaces (like walls).

What’s next

World Labs joining forces with AMD is only the start of this pipeline consolidation. Here is what we will explore in upcoming posts:

  • Converting raw point clouds into clean, textured polygonal meshes (.gltf / .glb) ready for game engines.
  • How Gaussian Splatting (3DGS) replaces traditional point clouds for real-time 60fps spatial rendering.
  • Benchmarking ROCm vs CUDA on actual 3D reconstruction workloads.

Bookmark this page to keep your spatial AI setup handy when you start building your next 3D environment.


Diagrams

Horizontal pipeline flowchart comparing 2D Generative AI (Prompt -> Diffusion -> Flat 2D PNG) against Spatial World Models (Prompt/Image -> Depth & Geometry Estimation -> 3D Mesh / Point Cloud -> Interactive Physics Engine) running on AMD ROCm compute.

Horizontal pipeline flowchart comparing 2D Generative AI (Prompt -> Diffusion -> Flat 2D PNG) against Spatial World Models (Prompt/Image -> Depth & Geometry Estimation -> 3D Mesh / Point Cloud -> Interactive Physics Engine) running on AMD ROCm compute.

Bar chart comparing coordinate spaces in 3D reconstruction: showing X (width), Y (height), and Z (estimated depth) mapped to vertex density per point cloud cluster.

Bar chart comparing coordinate spaces in 3D reconstruction: showing X (width), Y (height), and Z (estimated depth) mapped to vertex density per point cloud cluster.


This tutorial was drafted with AI assistance and verified with runnable examples. Found a mistake? Email the author — corrections are welcome.

Share: X (Twitter) Facebook LinkedIn