I Built a Self-Hosted Memory System for AI Agents
A deep walkthrough of a self-hosted memory architecture for AI agents — covering persistence, retrieval, and production-grade infrastructure design.
Technical tutorials, architecture breakdowns, app demos, and engineering devlogs. Each video opens on a dedicated watch page.
A deep walkthrough of a self-hosted memory architecture for AI agents — covering persistence, retrieval, and production-grade infrastructure design.
28 videos
Building an autonomous AI DevOps agent that handles code writing, review, and merge operations across a real development workflow.
Integrating MCP tool servers into Open WebUI with Kubernetes orchestration and MCPO for production-grade agent tooling.
Practical strategies to reduce AI token consumption and lower inference costs without sacrificing capability.
A full walkthrough of a homelab Kubernetes cluster built on Orange Pi 6 Plus and HP ProBook hardware for AI workload testing.
Introducing an open source memory control plane for AI agents — managing context, persistence, and retrieval across agent architectures.
Comparing MCP authentication and registry approaches — why Context Forge provides a better enterprise workflow for Azure-based MCP setups.
Setting up role-based access control for LLM models and MCP tool servers in Open WebUI — essential for multi-user enterprise deployments.
A practical comparison of the five best free LLM API endpoints for developers — covering rate limits, model quality, and production viability.
An architectural breakdown of why OpenAPI specifications fall short for complex agent interactions and how MCP solves the gap.
End-to-end build of a custom AI agent connected to GroupMe messaging using MCP tools and Open WebUI as the orchestration layer.
Building a Grafana monitoring dashboard for Open WebUI and MCP servers — observability for self-hosted AI infrastructure.
When existing tools fail, you build your own. A custom NPU inference backend for local LLM workloads where LM Studio fell short.
An honest look at how AI is reshaping DevOps roles — what changes, what stays, and how to stay ahead of the curve.
Using Claude Code with MCP to control video editing software — a real example of AI workflow automation beyond coding.
Setting up local LLM inference on Intel NPU hardware using the OpenVINO runtime and a new Windows AI wrapper.
Replacing manual project management with an MCP-powered AI agent that controls ClickUp via natural language.
How to connect free LLM API providers to Cline for AI-assisted coding without burning through paid credits.
A practical guide to getting the most out of free AI coding tools by combining the right models, providers, and workflows.
Learn how to run Large Language Models locally on your machine using just your CPU's Neural Processing Unit (NPU). No expensive GPU required!
A complete guide to setting up Open WebUI and connecting it to the Hugging Face API to run powerful LLMs completely for free.
A quick demonstration of how to set up, track, and log medication reminders directly on your Wear OS smartwatch with MedTick.
A quick look at WristLux — a standalone Wear OS light meter that reads ambient lux from the watch sensor for desk, plant, photography, and outdoor checks.
A walkthrough of FlexLog — a standalone Wear OS workout logger for set and rep tracking, rest timers, templates, and history directly from the wrist.
A full demo of FidgetDrop — a privacy-first Android haptic fidget app with satisfying tactile interactions, local records, and the optional FidgetDrop Pro upgrade.
A look at Tally Counter for Wear OS — up to eight named counters, step sizes, targets, rotary input, and a Tile and complication for counting straight from the wrist.
A demo of FidgetDrop for Wear OS — three haptic-first fidget modes, optional locally generated sound, Discreet Mode, and a Tile and complication for quick access.
A real-world experiment: give Claude Fable 5 one prompt per app, build it, and ship it to the Play Store — covering FlickDeck, WatchCheck, FidgetDrop Wear OS, and more.