Skip to main content
Version: 0.31.0

Render WebGPU effects into a Vision Camera feed Mobile

note

This guide is exclusively for Mobile (React Native) applications.

This guide shows how to render each frame of a published VisionCamera feed yourself with WebGPU, using the /webgpu entry point of @fishjam-cloud/react-native-vision-camera-source. For every camera frame, your worklet receives the camera as a GPU texture and an output texture to draw into. The useVisionCameraWebGpuSource hook handles publishing, GPU synchronization with the video encoder, timestamps and frame lifetimes.

Don't need VisionCamera?

The same kind of rendering works on the camera Fishjam manages itself, without VisionCamera. See Render WebGPU effects into the camera.

Every camera frame reaches your onFrame worklet. Calling render(...) gives you the live camera as a GPU texture and an output texture to draw into. Whatever you draw is what peers receive.

Prerequisites​

On top of the Vision Camera setup:

  • react-native-webgpu ≥ 0.10.1
  • unplugin-typegpu in your app's Babel config (the TGSL shaders need its build-time transform)
  • iOS 17+: the camera-import path relies on Metal external-texture features not guaranteed on earlier versions. The base integration has no such requirement, so you can keep your ios.deploymentTarget unchanged and gate WebGPU usage at runtime.
npm install react-native-webgpu typegpu unplugin-typegpu
module.exports = { presets: ["babel-preset-expo"], plugins: ["unplugin-typegpu/babel", "react-native-worklets/plugin"], };

Keep react-native-worklets/plugin as the last plugin.

Get a camera-capable GPU device​

useCameraWebGpuDevice returns an app-wide shared GPUDevice, requested with the features the camera-import path needs:

const { device, error } = useCameraWebGpuDevice();

To use your own device instead, pass it as the device option of useVisionCameraWebGpuSource. It is validated against getRequiredWebGpuCameraFeatures(); a device missing any of the features surfaces a descriptive error instead of failing per frame.

Publish the camera through your own shaders​

First time?

The WebGPU effects tutorial builds the same kind of pipeline step by step on Fishjam's own camera, from a passthrough render pass to a watermark overlay and a color effect. The shaders and render passes carry over; only the setup around them differs.

The example below publishes the camera in grayscale. To apply your own effect, replace the fragment stage. Your component only does the drawing; the hook handles publishing, GPU synchronization with the video encoder, timestamps and frame lifetimes.

note

The shaders are written in TypeGPU (TGSL): typed TypeScript functions compiled to WGSL by unplugin-typegpu. TypeGPU is not required; you can hand-write WGSL and prepend the bindings' bindingDeclarations yourself.

import React, { useCallback, useMemo } from "react"; import tgpu from "typegpu"; import * as d from "typegpu/data"; import { dot } from "typegpu/std"; import { useCamera as useVisionCamera, useCameraPermission, type Frame, } from "react-native-vision-camera"; import { RTCView } from "@fishjam-cloud/react-native-client"; import { useVisionCameraWebGpuSource, useCameraWebGpuDevice, createCameraShaderBindings, getOutputSurfaceFormat, type WebGpuFrameRenderFunction, } from "@fishjam-cloud/react-native-vision-camera-source/webgpu"; // Full-screen triangle; uv spans the visible area. const vertexMain = tgpu.vertexFn({ in: { vertexIndex: d.builtin.vertexIndex }, out: { position: d.builtin.position, uv: d.location(0, d.vec2f) }, })((input) => { const positions = [d.vec2f(-1, -1), d.vec2f(3, -1), d.vec2f(-1, 3)]; const p = positions[input.vertexIndex]; return { position: d.vec4f(p.x, p.y, 0, 1), uv: d.vec2f((p.x + 1) * 0.5, 1 - (p.y + 1) * 0.5), }; }); export function GrayscaleCameraPublisher() { const { hasPermission } = useCameraPermission(); const { device } = useCameraWebGpuDevice(); const effect = useMemo(() => { if (device == null) return null; const cameraBindings = createCameraShaderBindings(device); const fragmentMain = tgpu.fragmentFn({ in: { uv: d.location(0, d.vec2f) }, out: d.vec4f, })((input) => { const color = cameraBindings.sampleCamera(input.uv); const gray = dot(color.xyz, d.vec3f(0.299, 0.587, 0.114)); return d.vec4f(gray, gray, gray, 1); }); // TypeGPU cannot emit the external-texture binding itself, so prepend // cameraBindings.bindingDeclarations to the resolved WGSL. const module = device.createShaderModule({ code: cameraBindings.bindingDeclarations + tgpu.resolve({ externals: { vertexMain, fragmentMain } }), }); const pipeline = device.createRenderPipeline({ layout: device.createPipelineLayout({ bindGroupLayouts: [cameraBindings.bindGroupLayout], }), vertex: { module, entryPoint: "vertexMain" }, fragment: { module, entryPoint: "fragmentMain", targets: [{ format: getOutputSurfaceFormat() }], }, }); return { cameraBindings, pipeline }; }, [device]); const onFrame = useCallback( (frame: Frame, render: WebGpuFrameRenderFunction) => { "worklet"; if (effect == null) return; // drop frames until the pipeline is ready render(({ commandEncoder, outputView, cameraBindGroup }) => { const pass = commandEncoder.beginRenderPass({ colorAttachments: [ { view: outputView, loadOp: "clear", storeOp: "store" }, ], }); pass.setPipeline(effect.pipeline); pass.setBindGroup(0, cameraBindGroup!); pass.draw(3); pass.end(); }); }, [effect], ); const { frameOutput, stream } = useVisionCameraWebGpuSource("my-camera", { width: 720, height: 1280, cameraShaderBindings: effect?.cameraBindings, onFrame, }); useVisionCamera({ device: "front", isActive: hasPermission, outputs: [frameOutput], }); if (!stream) return null; return ( <RTCView mediaStream={stream} style={{ height: 200, width: 200 }} objectFit="cover" /> ); }

How the example works:

  • createCameraShaderBindings(device) gives your shaders sampleCamera(uv), which returns upright RGB on both platforms and handles the YUV decode for you.
  • Everything created from the device (bindings, shaders, pipeline) lives in one useMemo keyed by the device. TypeGPU cannot emit the camera's texture_external binding, so cameraBindings.bindingDeclarations is prepended to the resolved WGSL, and the fragment targets getOutputSurfaceFormat() (rgba8unorm on Android, bgra8unorm on iOS).
  • Passing cameraShaderBindings to the hook makes the render context carry a ready-made cameraBindGroup, rebuilt each frame because the camera's external texture expires with every frame.
  • The worklet encodes one render pass; the hook submits it and synchronizes with the video encoder. frameOutput plugs into VisionCamera's useCamera; stream is the self-view.

Other peers receive the feed among their customVideoTracks.

Rules inside onFrame​

The callback you pass to render(...) receives a WebGpuFrameRenderContext with the device, queue, commandEncoder, the live cameraTexture (a GPUExternalTexture), the output surface (outputTexture, outputView, outputWidth, outputHeight), and camera metadata (cameraWidth, cameraHeight, cameraIsMirrored).

Rules your worklet must follow
  • Always draw into the provided outputView. Calling outputTexture.createView() per frame leaks native wrappers on the frame runtime, because GPUTextureView has no release API.
  • Call render(...) at most once per frame. Skipping it drops the frame; nothing is published for it.
  • Don't call queue.submit() yourself. The hook submits your passes and synchronizes with the video encoder.
  • Finish GPU uploads before frames flow. Helpers like queue.copyExternalImageToTexture submit work internally. Running them from the JS thread while the source is active races the hook's submissions and can crash the app, so upload textures before activating the camera.
  • Camera bind groups cannot be cached across frames, because the external texture changes every frame. The hook rebuilds cameraBindGroup for you; if you build your own, call createCameraBindGroup inside the worklet each frame.
  • Keep onFrame's identity stable (useCallback or module scope).

After render(...) returns you may keep using the frame (for example, to run inference on it), but only until your callback returns; the hook releases it afterwards.

Going further​

  • Overlays: a frame may contain more than one render pass. Encode additional passes into the same outputView (with loadOp: "load") after the camera pass to draw watermarks or other content on top.
  • Cropping helpers: the grayscale example stretches the camera to the output. computeAspectFillCrop and computeSquareCrop compute the crop that fills your output aspect ratio (like objectFit: "cover"); packFrameCropParams packs a crop for your own uniform buffers.
  • Pipelines that cannot sample texture_external can resolve the camera into an owned rgba8unorm texture with createCameraTextureResolver and resolveCameraTexture, at the cost of one extra render pass per frame.

The full toolkit is documented in the Custom Video Source API reference.

Platform notes​

  • sampleCamera(uv) returns upright RGB on both platforms. On Android it performs the BT.709 limited-range YUV→RGB decode in-shader; on iOS the camera already arrives as RGB.
  • getOutputSurfaceFormat() returns the published surface format: rgba8unorm on Android, bgra8unorm on iOS. Use it for your fragment targets instead of hard-coding a format.
  • The context's cameraIsMirrored tells you whether the camera feed is mirrored (typically the front camera).