### Building a notification routing platform that’s both reliable and performant requires careful technology choices. Here’s a transparent look at what powers NotifyGate and why we chose each piece of our stack.

## The Core: Elixir & Phoenix

NotifyGate is built on **Elixir** and the **Phoenix Framework**. If you’re wondering why we didn’t choose Node.js, Python, or Go, here’s our reasoning:

### Why Elixir?

**1\. Built for Concurrency**

Notifications are inherently concurrent. At any moment, we’re:
- Receiving thousands of events
- Evaluating routing rules in parallel
- Sending to multiple destinations simultaneously
- Processing webhooks from external services

Elixir runs on the BEAM VM (Erlang’s virtual machine), designed from the ground up for exactly this kind of workload. Handling 100,000 concurrent processes? No problem. The BEAM handles this effortlessly.

```elixir
# Process millions of events concurrently
events
|> Task.async_stream(&route_notification/1, max_concurrency: 10_000)
|> Stream.run()
```

**2\. Fault Tolerance by Design**

When a Slack webhook fails, we don’t want it to crash the entire system. Elixir’s “let it crash” philosophy with supervision trees means failures are isolated and recovered automatically.

```elixir
# If a worker crashes, the supervisor restarts it
children = [\
  {Task.Supervisor, name: NotifyGate.TaskSupervisor},\
  {NotifyGate.NotificationWorker, []}\
]

Supervisor.start_link(children, strategy: :one_for_one)
```

**3\. Soft Real-Time Capabilities**

The BEAM’s preemptive scheduler ensures low-latency responses even under heavy load. Your critical alerts get delivered quickly, even when we’re processing thousands of events per second.

### Why Phoenix?

Phoenix gives us:
- **LiveView** for real-time UI updates without writing JavaScript
- **Channels** for WebSocket connections (real-time dashboard updates)
- **Fast routing** and request handling
- **Built-in instrumentation** with Telemetry

## Data Layer: PostgreSQL with Ash Framework

### PostgreSQL: The Reliable Workhorse

We use **PostgreSQL** for persistent storage because:
- **ACID compliance** ensures your events are never lost
- **JSONB support** lets us store flexible event data without schema migrations
- **Advanced indexing** (GIN, BRIN) keeps queries fast even with millions of events
- **Battle-tested reliability** from decades of production use

### Ash Framework: Declarative Domain Logic

**Ash** is our secret weapon for rapid, maintainable development:

```elixir
defmodule NotifyGate.Events.Event do
  use Ash.Resource,
    domain: NotifyGate.Events,
    data_layer: AshPostgres.DataLayer

attributes do
    uuid_primary_key :id
    attribute :type, :string, allow_nil?: false
    attribute :priority, :atom, constraints: [one_of: [:low, :medium, :high, :critical]]
    attribute :metadata, :map
    attribute :received_at, :utc_datetime_usec
  end

actions do
    defaults [:read]

create :create do
      accept [:type, :priority, :metadata]
      change set_attribute(:received_at, &DateTime.utc_now/0)
    end
  end
end
```

Ash gives us:
- **Declarative resources** that auto-generate APIs, changesets, and queries
- **Policy-based authorization** built into the framework
- **Automatic GraphQL and JSON:API** endpoints
- **Pub/Sub integration** for real-time features

## Background Jobs: Oban

**Oban** handles all our asynchronous work:
- Scheduled notifications
- Retry logic for failed deliveries
- Rate limiting enforcement
- Event cleanup and archival

```elixir
# Schedule a notification with automatic retries
%{
  event_id: event.id,
  destination: "slack",
  channel_id: "C123456"
}
|> NotifyGate.Workers.NotificationWorker.new()
|> Oban.insert()
```

Why Oban over Sidekiq, Bull, or other job queues?
- **No external dependencies** \- uses PostgreSQL, not Redis
- **Reliable** \- jobs are stored in the database with ACID guarantees
- **Observable** \- built-in web UI for monitoring
- **Smart queuing** \- rate limiting, prioritization, and unique jobs out of the box

## Caching & Performance: Cachex + ETS

For hot path performance, we use **Cachex** backed by **ETS** (Erlang Term Storage):

```elixir
# Cache routing rules in memory
Cachex.fetch(:rules_cache, workspace_id, fn ->
  NotifyGate.Rules.list_active_rules(workspace_id)
end)
```

- **Sub-millisecond lookups** for routing rules and templates
- **Automatic expiration** to keep data fresh
- **Distributed caching** across nodes (when we scale horizontally)
- **No external cache server** needed

## Rate Limiting: Hammer

**Hammer** provides flexible, distributed rate limiting:

```elixir
# Limit to 100 notifications per hour per channel
case Hammer.check_rate("slack:#{channel_id}", 60_000 * 60, 100) do
  {:allow, _count} -> send_notification()
  {:deny, _limit} -> {:error, :rate_limited}
end
```

Uses Mnesia for distributed rate limiting across multiple servers without Redis or external coordination.

## API & Integrations

### HTTP Client: Req

**Req** is our HTTP client for sending notifications to external services:

```elixir
# Send to Slack with automatic retries
Req.post!("https://slack.com/api/chat.postMessage",
  json: %{channel: "#alerts", text: message},
  headers: [{"Authorization", "Bearer #{token}"}],
  retry: :transient
)
```

Clean API, built-in retry logic, and excellent error handling.

### Email: Swoosh + Gen\_SMTP

For email notifications:
- **Swoosh** provides a unified interface for multiple email providers
- **Gen\_SMTP** for direct SMTP delivery
- Support for SendGrid, Mailgun, AWS SES, and custom SMTP

## Frontend: Phoenix LiveView + TailwindCSS

### LiveView: Real-Time Without JavaScript

Our entire dashboard is built with **Phoenix LiveView**:

```elixir
defmodule NotifyGateWeb.DashboardLive do
  use NotifyGateWeb, :live_view

def mount(_params, _session, socket) do
    if connected?(socket) do
      Phoenix.PubSub.subscribe(NotifyGate.PubSub, "events")
    end

{:ok, assign(socket, events: load_recent_events())}
  end

def handle_info({:new_event, event}, socket) do
    {:noreply, update(socket, :events, &[event | &1])}
  end
end
```

Real-time updates, server-side rendering, and minimal JavaScript. Your dashboard updates live as events flow through the system.

### TailwindCSS v4: Modern Styling

- **Utility-first CSS** for rapid UI development
- **Dark mode** support out of the box
- **Custom design system** with CSS variables
- **Optimized builds** \- only ships CSS you actually use

## Monitoring & Observability

### Telemetry: Built-in Instrumentation

Phoenix and Oban emit telemetry events we aggregate for monitoring:

```elixir
:telemetry.attach(
  "notify-gate-events",
  [:notify_gate, :notification, :sent],
  &NotifyGate.Metrics.handle_event/4,
  nil
)
```

### Metrics We Track
- Event ingestion rate
- Notification delivery time (p50, p95, p99)
- Rule evaluation latency
- Destination health (success/failure rates)
- Queue depth and job processing times

## Why This Stack Works for Us

### Operational Simplicity

Our entire production infrastructure:
- **One language** (Elixir) for everything
- **One database** (PostgreSQL) for both data and jobs
- **No external queues** (no Redis, RabbitMQ, or Kafka needed)
- **Minimal dependencies** to maintain and upgrade

### Performance at Scale

With this stack, a single server handles:
- **10,000+ events per second**
- **100,000+ concurrent WebSocket connections** (LiveView)
- **Millions of database records** with fast queries
- **Sub-100ms p95 latency** for most operations

### Developer Productivity
- **Fast iteration** with LiveView (no API layer needed)
- **Type safety** without the boilerplate (Elixir’s pattern matching)
- **Built-in testing** tools (ExUnit, Phoenix.ConnTest, Phoenix.LiveViewTest)
- **Hot code reloading** in development

## What We’d Change

No stack is perfect. If we were starting today, we might consider:
- **Alternative:** Go + NATS for certain high-throughput workloads
- **Addition:** More aggressive use of read replicas for analytics queries
- **Exploration:** Clickhouse for long-term event analytics

But honestly? We’re thrilled with our choices. The stack is productive, performant, and proven at scale.

## The Big Picture

Here’s our architecture in a nutshell:

```
┌─────────────┐
│   Client    │
│  (Your App) │
└──────┬──────┘
       │ POST /api/v1/events
       ▼
┌─────────────────────┐
│   Phoenix API       │
│  (Event Ingestion)  │
└──────┬──────────────┘
       │
       ▼
┌─────────────────────┐
│   Rule Engine       │
│ (Elixir processes)  │
└──────┬──────────────┘
       │
       ▼
┌─────────────────────┐
│   Oban Workers      │
│ (Async delivery)    │
└──────┬──────────────┘
       │
       ├─► Slack
       ├─► Discord
       ├─► Email
       └─► Webhooks

All backed by PostgreSQL
All monitored via Telemetry
All running on the BEAM VM
```

## Open Source Credits

NotifyGate stands on the shoulders of giants. Huge thanks to the maintainers of:
- **Elixir & Phoenix** \- José Valim and the core team
- **Ash Framework** \- Zach Daniel and contributors
- **Oban** \- Parker Selbert and Sorentwo
- **And dozens more** open source libraries we depend on

## Want to Learn More?

Interested in the nitty-gritty details of how we built specific features?
- [How We Handle 10k Events/Second](/content/blog/scaling-event-processing/index.html) (coming soon)
- [Rule Engine Internals](/content/blog/rule-engine-architecture/index.html) (coming soon)
- [Zero-Downtime Deployments with Elixir](/content/blog/zero-downtime-deploys/index.html) (coming soon)

Got technical questions? [Let’s talk](mailto:engineering@notifygate.io) \- we love chatting about architecture.
