Render WebGPU effects into a Vision Camera feed Mobile
This guide is exclusively for Mobile (React Native) applications.
This guide shows how to render each frame of a published VisionCamera feed yourself with WebGPU, using the /webgpu entry point of @fishjam-cloud/react-native-vision-camera-source. For every camera frame, your worklet receives the camera as a GPU texture and an output texture to draw into. The useVisionCameraWebGpuSource hook handles publishing, GPU synchronization with the video encoder, timestamps and frame lifetimes.
The same kind of rendering works on the camera Fishjam manages itself, without VisionCamera. See Render WebGPU effects into the camera.
Every camera frame reaches your onFrame worklet. Calling render(...) gives you the live camera as a GPU texture and an output texture to draw into. Whatever you draw is what peers receive.
Prerequisites​
On top of the Vision Camera setup:
react-native-webgpu≥ 0.10.1unplugin-typegpuin your app's Babel config (the TGSL shaders need its build-time transform)- iOS 17+: the camera-import path relies on Metal external-texture features not guaranteed on earlier versions. The base integration has no such requirement, so you can keep your
ios.deploymentTargetunchanged and gate WebGPU usage at runtime.
- npm
- Yarn
- pnpm
- Bun
npm install react-native-webgpu typegpu unplugin-typegpu
yarn add react-native-webgpu typegpu unplugin-typegpu
pnpm add react-native-webgpu typegpu unplugin-typegpu
bun add react-native-webgpu typegpu unplugin-typegpu
module.exports = { presets: ["babel-preset-expo"], plugins: ["unplugin-typegpu/babel", "react-native-worklets/plugin"], };
Keep react-native-worklets/plugin as the last plugin.
Get a camera-capable GPU device​
useCameraWebGpuDevice returns an app-wide shared GPUDevice, requested with the features the camera-import path needs:
const {device ,error } =useCameraWebGpuDevice ();
To use your own device instead, pass it as the device option of useVisionCameraWebGpuSource. It is validated against getRequiredWebGpuCameraFeatures(); a device missing any of the features surfaces a descriptive error instead of failing per frame.
Publish the camera through your own shaders​
The WebGPU effects tutorial builds the same kind of pipeline step by step on Fishjam's own camera, from a passthrough render pass to a watermark overlay and a color effect. The shaders and render passes carry over; only the setup around them differs.
The example below publishes the camera in grayscale. To apply your own effect, replace the fragment stage. Your component only does the drawing; the hook handles publishing, GPU synchronization with the video encoder, timestamps and frame lifetimes.
The shaders are written in TypeGPU (TGSL): typed TypeScript functions compiled to WGSL by unplugin-typegpu. TypeGPU is not required; you can hand-write WGSL and prepend the bindings' bindingDeclarations yourself.
importReact , {useCallback ,useMemo } from "react"; importtgpu from "typegpu"; import * asd from "typegpu/data"; import {dot } from "typegpu/std"; import {useCamera asuseVisionCamera ,useCameraPermission , typeFrame , } from "react-native-vision-camera"; import {RTCView } from "@fishjam-cloud/react-native-client"; import {useVisionCameraWebGpuSource ,useCameraWebGpuDevice ,createCameraShaderBindings ,getOutputSurfaceFormat , typeWebGpuFrameRenderFunction , } from "@fishjam-cloud/react-native-vision-camera-source/webgpu"; // Full-screen triangle; uv spans the visible area. constvertexMain =tgpu .vertexFn ({in : {vertexIndex :d .builtin .vertexIndex },out : {position :d .builtin .position ,uv :d .location (0,d .vec2f ) }, })((input ) => { constpositions = [d .vec2f (-1, -1),d .vec2f (3, -1),d .vec2f (-1, 3)]; constp =positions [input .vertexIndex ]; return {position :d .vec4f (p .x ,p .y , 0, 1),uv :d .vec2f ((p .x + 1) * 0.5, 1 - (p .y + 1) * 0.5), }; }); export functionGrayscaleCameraPublisher () { const {hasPermission } =useCameraPermission (); const {device } =useCameraWebGpuDevice (); consteffect =useMemo (() => { if (device == null) return null; constcameraBindings =createCameraShaderBindings (device ); constfragmentMain =tgpu .fragmentFn ({in : {uv :d .location (0,d .vec2f ) },out :d .vec4f , })((input ) => { constcolor =cameraBindings .sampleCamera (input .uv ); constgray =dot (color .xyz ,d .vec3f (0.299, 0.587, 0.114)); returnd .vec4f (gray ,gray ,gray , 1); }); // TypeGPU cannot emit the external-texture binding itself, so prepend // cameraBindings.bindingDeclarations to the resolved WGSL. constmodule =device .createShaderModule ({code :cameraBindings .bindingDeclarations +tgpu .resolve ({externals : {vertexMain ,fragmentMain } }), }); constpipeline =device .createRenderPipeline ({layout :device .createPipelineLayout ({bindGroupLayouts : [cameraBindings .bindGroupLayout ], }),vertex : {module ,entryPoint : "vertexMain" },fragment : {module ,entryPoint : "fragmentMain",targets : [{format :getOutputSurfaceFormat () }], }, }); return {cameraBindings ,pipeline }; }, [device ]); constonFrame =useCallback ( (frame :Frame ,render :WebGpuFrameRenderFunction ) => { "worklet"; if (effect == null) return; // drop frames until the pipeline is readyrender (({commandEncoder ,outputView ,cameraBindGroup }) => { constpass =commandEncoder .beginRenderPass ({colorAttachments : [ {view :outputView ,loadOp : "clear",storeOp : "store" }, ], });pass .setPipeline (effect .pipeline );pass .setBindGroup (0,cameraBindGroup !);pass .draw (3);pass .end (); }); }, [effect ], ); const {frameOutput ,stream } =useVisionCameraWebGpuSource ("my-camera", {width : 720,height : 1280,cameraShaderBindings :effect ?.cameraBindings ,onFrame , });useVisionCamera ({device : "front",isActive :hasPermission ,outputs : [frameOutput ], }); if (!stream ) return null; return ( <RTCView mediaStream ={stream }style ={{height : 200,width : 200 }}objectFit ="cover" /> ); }
How the example works:
createCameraShaderBindings(device)gives your shaderssampleCamera(uv), which returns upright RGB on both platforms and handles the YUV decode for you.- Everything created from the
device(bindings, shaders, pipeline) lives in oneuseMemokeyed by the device. TypeGPU cannot emit the camera'stexture_externalbinding, socameraBindings.bindingDeclarationsis prepended to the resolved WGSL, and the fragment targetsgetOutputSurfaceFormat()(rgba8unormon Android,bgra8unormon iOS). - Passing
cameraShaderBindingsto the hook makes the render context carry a ready-madecameraBindGroup, rebuilt each frame because the camera's external texture expires with every frame. - The worklet encodes one render pass; the hook submits it and synchronizes with the video encoder.
frameOutputplugs into VisionCamera'suseCamera;streamis the self-view.
Other peers receive the feed among their customVideoTracks.
Rules inside onFrame​
The callback you pass to render(...) receives a WebGpuFrameRenderContext with the device, queue, commandEncoder, the live cameraTexture (a GPUExternalTexture), the output surface (outputTexture, outputView, outputWidth, outputHeight), and camera metadata (cameraWidth, cameraHeight, cameraIsMirrored).
- Always draw into the provided
outputView. CallingoutputTexture.createView()per frame leaks native wrappers on the frame runtime, becauseGPUTextureViewhas no release API. - Call
render(...)at most once per frame. Skipping it drops the frame; nothing is published for it. - Don't call
queue.submit()yourself. The hook submits your passes and synchronizes with the video encoder. - Finish GPU uploads before frames flow. Helpers like
queue.copyExternalImageToTexturesubmit work internally. Running them from the JS thread while the source is active races the hook's submissions and can crash the app, so upload textures before activating the camera. - Camera bind groups cannot be cached across frames, because the external texture changes every frame. The hook rebuilds
cameraBindGroupfor you; if you build your own, callcreateCameraBindGroupinside the worklet each frame. - Keep
onFrame's identity stable (useCallbackor module scope).
After render(...) returns you may keep using the frame (for example, to run inference on it), but only until your callback returns; the hook releases it afterwards.
Going further​
- Overlays: a frame may contain more than one render pass. Encode additional passes into the same
outputView(withloadOp: "load") after the camera pass to draw watermarks or other content on top. - Cropping helpers: the grayscale example stretches the camera to the output.
computeAspectFillCropandcomputeSquareCropcompute the crop that fills your output aspect ratio (likeobjectFit: "cover");packFrameCropParamspacks a crop for your own uniform buffers. - Pipelines that cannot sample
texture_externalcan resolve the camera into an ownedrgba8unormtexture withcreateCameraTextureResolverandresolveCameraTexture, at the cost of one extra render pass per frame.
The full toolkit is documented in the Custom Video Source API reference.
Platform notes​
sampleCamera(uv)returns upright RGB on both platforms. On Android it performs the BT.709 limited-range YUV→RGB decode in-shader; on iOS the camera already arrives as RGB.getOutputSurfaceFormat()returns the published surface format:rgba8unormon Android,bgra8unormon iOS. Use it for your fragment targets instead of hard-coding a format.- The context's
cameraIsMirroredtells you whether the camera feed is mirrored (typically the front camera).
Related guides​
- Publish a Vision Camera feed: publish the camera without custom rendering
- Render WebGPU effects into the camera: the same approach on Fishjam's own camera
- Low-level frame API: the pooled-surface layer the hook builds on
- How custom sources work
- API reference: Vision Camera Source package, Custom Video Source package