Files
codegirl-games/.github/issues/add-configurable-internal-render-scale.md
T
2026-08-09 08:46:04 +00:00

6.1 KiB
Raw Blame History

Add configurable internal render scale to improve fill-bound FPS

Summary

Complete-frame benchmarks show that the current sprite renderer is primarily pixel/fill bound on Vulkan/Lavapipe, not CPU draw-preparation bound. Rendering the world into a smaller offscreen color target and nearest-blitting it to the native swapchain produced the largest measured FPS improvement.

This should be an opt-in render-quality setting with a native-resolution fallback. For pixel art, a 50% scale is especially useful because it maps to an exact 2× nearest-neighbor upscale.

Evidence

Environment:

  • Commit dc66203259e92ce39291556f7b6c26c8dd999b84
  • Odin dev-2026-05-nightly:ea5175d
  • -debug -o:speed
  • SDL 3.4.12, Vulkan/Lavapipe, immediate present
  • 800×600 swapchain
  • Five 150-frame trials per invocation after 50 warm-up frames
  • Order-balanced baseline/candidate invocations
  • GPU idle wait included before stopping each trial timer

Fill and overdraw scaling

The committed perf-frame harness produced:

Workload Median frame time Median FPS
1 visible sprite 0.592 ms 1,690
24 visible sprites 2.249 ms 445
64 visible sprites 4.017 ms 249
128 visible sprites 6.255 ms 160
128 stacked sprites 7.692 ms 130

Stacking the same 128 sprites increased frame time by approximately 23%, confirming that overdraw matters.

With one centered sprite, render-target scaling produced:

Window size Pixels Median frame time
400×300 120,000 0.268 ms
800×600 480,000 0.545 ms
1600×1200 1,920,000 2.075 ms

The near-linear increase at larger sizes is further evidence of a pixel-bound workload.

Internal render-scale prototype

The temporary candidate kept the physical swapchain at 800×600, rendered sprites into a smaller COLOR_TARGET | SAMPLER texture, and used BlitGPUTexture with nearest filtering to upscale into the swapchain.

Aggregate medians across ten trials per mode:

Workload Native 75% scale 50% scale
128 spread sprites 6.082 ms / 164 FPS 4.082 ms / 245 FPS 3.082 ms / 325 FPS
128 stacked sprites 7.912 ms / 126 FPS 5.116 ms / 195 FPS 3.034 ms / 330 FPS

Compared with native resolution:

  • 75% reduced frame time by 3335% and increased FPS by 4955%.
  • 50% reduced frame time by 4962% and increased FPS by 97161%.

A visual smoke test confirmed that the 50% path rendered the complete scene at the correct orientation and 800×600 output size. It was visibly coarser, as expected. A 75% scale at 800×600 does not produce an integer upscale and can create uneven pixel sizing with nearest filtering.

Reproduction harness

Native spread and stacked workloads:

make perf-frame \
	PERF_FRAME_SCENARIO=0 \
	PERF_FRAME_SPRITES=128 \
	PERF_FRAME_WIDTH=800 \
	PERF_FRAME_HEIGHT=600

make perf-frame \
	PERF_FRAME_SCENARIO=3 \
	PERF_FRAME_SPRITES=128 \
	PERF_FRAME_WIDTH=800 \
	PERF_FRAME_HEIGHT=600

Run the same commands on the candidate branch with internal render scale set to 75% and 50%. Keep every PERF_FRAME_* value unchanged between comparisons.

Suggested fix

Add explicit logical/output and internal-render dimensions to App:

Render_Scale :: enum {
	Native,
	Three_Quarter,
	Half,
}

App :: struct {
	// Existing swapchain fields remain the logical/output dimensions.
	swapchain_texture: ^sdl.GPUTexture,
	swapchain_w:       u32,
	swapchain_h:       u32,

	render_scale:   Render_Scale,
	scene_texture:  ^sdl.GPUTexture,
	scene_w:        u32,
	scene_h:        u32,
}

Create the offscreen target with the swapchain format and both usages required by SDL's blit path:

app.scene_texture = sdl.CreateGPUTexture(
	app.device,
	{
		type = .D2,
		format = sdl.GetGPUSwapchainTextureFormat(app.device, app.window),
		usage = {.COLOR_TARGET, .SAMPLER},
		width = app.scene_w,
		height = app.scene_h,
		layer_count_or_depth = 1,
		num_levels = 1,
		sample_count = ._1,
	},
)

Render the sprite pass into scene_texture. Continue using the logical swapchain dimensions for camera and clip-space calculations so world layout does not change with render scale:

color_info := sdl.GPUColorTargetInfo {
	texture = app.scene_texture,
	clear_color = app.clear_color,
	load_op = .CLEAR,
	store_op = .STORE,
	cycle = true,
}

After ending the render pass, upscale into the full swapchain:

blit := sdl.GPUBlitInfo {
	source = {
		texture = app.scene_texture,
		w = app.scene_w,
		h = app.scene_h,
	},
	destination = {
		texture = app.swapchain_texture,
		w = app.swapchain_w,
		h = app.swapchain_h,
	},
	load_op = .DONT_CARE,
	filter = .NEAREST,
}
sdl.BlitGPUTexture(app.cmd, blit)

Use the existing direct-to-swapchain path at native scale to avoid an unnecessary blit. Recreate the offscreen texture whenever the swapchain size, format, or render-scale setting changes. If native-resolution UI or text is added later, render it after the world blit in a separate swapchain pass.

Rejected experiment: fragment alpha discard

Approximately 49.7% of pixels inside the baked frame rectangles have exact alpha zero. A temporary fragment shader discarded those texels:

vec4 texel = texture(u_tex, v_uv);
if (texel.a == 0.0) {
    discard;
}
out_color = texel;

Despite the high transparent coverage, this regressed frame time by roughly 56% in both spread and stacked workloads. Do not add alpha discard without contradictory hardware-GPU evidence.

Acceptance criteria

  • Native, 75%, and 50% internal render-scale settings are available.
  • Native scale retains the current direct-to-swapchain path.
  • The scene target is recreated safely on resize, format change, or scale change and released during shutdown.
  • Camera/world coordinates remain stable when render scale changes.
  • Nearest filtering is used for pixel-art output.
  • Add screenshot-based checks for output orientation, viewport coverage, and stable sprite placement at every scale.
  • Benchmark spread and stacked scenarios on at least one hardware GPU.
  • Document the quality tradeoff and recommend integer upscale ratios for pixel art.