AI-Portable
Article image for Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples Articles
Not Applicable

Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples

NVIDIA's DIN Deploy pairs ONNX Runtime with TensorRT RTX to run speech, segmentation, and image generation models locally in native C++.

Condensed by AI-Portable from Editorial queue.

NVIDIA’s DIN Deploy (Do Inference Now) is an open-source set of native C++ samples that shows how to package a model once and run it locally on Windows and Linux, without dragging a Python runtime into the final application. According to NVIDIA’s developer blog, the repository bridges the gap between a raw Hugging Face checkpoint and a hardware-accelerated desktop application by pairing ONNX Runtime with the TensorRT RTX execution provider.

From checkpoint to compiled app

Each sample follows the same architecture: a Python exporter downloads a model from Hugging Face and converts it to an ONNX artifact, while the application half is a C++ command-line tool built on ONNX Runtime session and tensor APIs. That separation keeps export logic out of deployment code. Most shared code uses only standard ORT tensor APIs, so it can run on any execution provider that supports those interfaces. Optional CUDA paths handle vendor-specific acceleration, but they are not required for the pipeline to function. The same ONNX Runtime API surface is also available through WinML 2.0, which gives Windows developers another route into the same models.

The repository ships CMake presets for Windows, Linux, x86-64, and Arm64. DirectX interop is Windows-only. By default, CMake downloads ONNX Runtime and TensorRT RTX automatically, making the initial build less dependent on manual environment setup.

What the samples actually run

DIN Deploy covers three distinct local AI workflows:

  • Automatic speech recognition (ASR): OpenAI Whisper for offline transcription, plus NVIDIA Parakeet TDT and NVIDIA Nemotron ASR Streaming for streaming pipelines.
  • Interactive segmentation: Meta SAM 2.1 samples for image and video masking, with model outputs converted to usable segmentation masks for selection and tracking.
  • Prompt-driven image generation: FLUX.2-klein-4B shows how to integrate GPU-resident resources with a cross-vendor shader interface.

Performance numbers measured on DGX Spark are substantial. GPU acceleration ranges from 39× real-time for Nemotron ASR streaming to 206× for Parakeet TDT. SAM 2.1 jumps from 0.5 FPS on CPU to 38.3 FPS on GPU. These figures highlight why a compiled C++ path with TensorRT RTX matters for interactive local applications: the latency difference is not incremental, it is the difference between unusable and real-time.

Graphics interop and drop-in quantization

The FLUX.2 sample also demonstrates ONNX Runtime 1.25 graphics interop with Vulkan and DirectX for sampling. That allows applications to keep image data on the GPU and avoid expensive round-trips through system memory. The sample further shows how post-training quantization using NVIDIA Model Optimizer produces a quantized ONNX model. Because the ONNX interface remains unchanged, the quantized model becomes a drop-in replacement that requires no application-code changes—though the quantization itself is hardware-dependent and assumes an RTX-class GPU for acceleration.

For developers who want to ship local inference as a native binary rather than a Python script, DIN Deploy lowers the practical barrier. The repository’s CMake presets, shared ONNX Runtime API, and explicit split between model export and deployment logic make it straightforward to copy a pipeline into an existing application. The main caveat is hardware: the samples shine on RTX GPUs, and the CPU path is dramatically slower, as the SAM 2.1 comparison shows.

Original source ↗