How Async Rust Works Under the Hood: Futures, Executors, and Pinning Explained
Learn how async Rust really works: Futures, poll, Wakers, executors, and Pin explained in plain English with code you can run.
Introduction
You write async fn, sprinkle in a few .await calls, and everything just works. Until it doesn't.
Then one day you hit an error like future cannot be sent between threads safely. Or you see Pin<Box<dyn Future<Output = ()>>> in a function signature and feel your confidence drain away. Or your whole server freezes because of one innocent-looking function call.
Most of the confusion around async Rust comes from one thing: it works very differently from async in languages like JavaScript, Go, or C#. There is no built-in runtime and no hidden scheduler. The language gives you a small set of building blocks, and the ecosystem builds everything else on top.
In this guide, we'll open the hood and look at those building blocks one by one: the Future trait, poll, wakers, executors, and the famously confusing Pin. We'll also write a tiny working executor from scratch. You can follow along with a plain cargo new project and no external crates.
By the end, you'll know what the compiler does with your async fn, why Pin exists, and how to avoid the most common production mistakes.
The Big Idea: Async Rust Is Lazy and Cooperative
Before the details, two ideas explain almost everything else.
- Futures are lazy. Calling an
async fndoes not run any of its body. It returns a value (a future) that describes the work. Nothing happens until something polls that value. - Scheduling is cooperative. A task runs until it hits an
.awaitthat can't make progress, and only then does it give control back. Nobody forcibly interrupts it.
Compare this with OS threads, where the kernel can pause your code at any instruction. In async Rust, you decide where the pause points are, and they are your .awaits.
Note: If a task never reaches an
.await, it never gives the thread back. That's the root cause of most "my async server froze" bugs. We'll come back to this.
What Is a Future, Really?
A future is any type that implements this trait from the standard library:
pub trait Future {
type Output;
fn poll(self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<Self::Output>;
}
pub enum Poll<T> {
Ready(T),
Pending,
}
That's the whole thing. A future is something you can ask, "Are you done yet?" by calling poll. It answers in one of two ways:
Poll::Ready(value): finished, here's the result.Poll::Pending: not done yet, ask me again later.
Notice that poll takes a Context. That's how a future says, "Don't keep asking me in a loop. I'll tell you when I'm ready." We'll get to that in the waker section.
Polling vs. Callbacks
Many languages use callbacks or promises that push results to you. Rust uses a poll-based model, where the executor pulls progress out of futures.
This has real benefits:
- Zero-cost abstractions. Futures are plain structs. No heap allocation is required just to have a future.
- No hidden runtime. You choose your executor, or write your own.
- Easy cancellation. Dropping a future cancels it. There is no separate cancel token to wire up.
- Backpressure. The consumer controls when work progresses.
What the Compiler Does With async fn
Here is a simple async function:
async fn greet() -> String {
let name = fetch_name().await;
format!("Hello, {name}")
}
The compiler turns this into an anonymous type that implements Future<Output = String>. Conceptually, it is a state machine, an enum with one variant per suspension point:
// Simplified. The real generated code is more involved.
enum GreetFuture {
Start,
WaitingForName { fetch: FetchNameFuture },
Done,
}
Each time the executor calls poll:
- The future resumes from its current state.
- It runs until the next
.awaitthat returnsPending, or until it finishes. - It saves any local variables it still needs inside the state, then returns.
This is why async Rust is so efficient. All the locals that live across an .await are stored inside the future itself, in one struct. No separate stack is needed per task, as with threads or green threads.
A Practical Consequence: Future Size
Because the future stores everything that lives across an .await, large local values make large futures:
async fn heavy() {
let buffer = [0u8; 64 * 1024]; // 64 KB lives inside the future
do_something().await;
println!("{}", buffer.len());
}
If you spawn thousands of tasks like this, memory adds up. Large buffers are better placed on the heap with Vec or Box.
Wakers: How a Future Says "Try Me Again"
If the executor just polled every future in a tight loop, you'd burn 100% CPU doing nothing. Wakers solve this.
When a future returns Pending, it must first arrange for someone to call wake() when progress becomes possible. The Waker lives inside the Context passed to poll.
The flow looks like this:
- The executor polls your future with a
Contextcontaining aWaker. - The future tries to do work (say, read from a socket) and finds no data yet.
- The future clones and stores the waker, then returns
Pending. - Later, something external happens, such as the OS reporting that the socket is readable.
- That event source calls
waker.wake(). - The executor puts the task back in its queue and polls it again.
The golden rule: Returning
Pendingwithout arranging a wake-up means your task may sleep forever. The compiler can't catch this, so it's one of the easiest bugs to write when implementing futures by hand.
Executors and Reactors: Who Actually Drives Everything?
People often use "runtime" loosely, but separate jobs are being done:
| Component | Job | Real-world example |
|---|---|---|
| Executor | Holds tasks, polls them when woken, runs them on threads | Tokio's scheduler, smol's executor |
| Reactor | Watches for I/O and timer events, calls wake() |
Tokio's I/O driver built on mio (epoll, kqueue, IOCP) |
| Futures | Describe the work and register wakers | Your async fns, TcpStream::read, sleep |
An executor is surprisingly simple at its core: a queue of tasks plus a loop. A reactor is the part that talks to the operating system.
Build a Tiny Executor From Scratch
Let's make this concrete. We'll build two things using only the standard library:
- A future that completes after a delay.
- A
block_onfunction that runs a future to completion.
Step 1: A Hand-Written Future
use std::future::Future;
use std::pin::Pin;
use std::sync::{Arc, Mutex};
use std::task::{Context, Poll, Waker};
use std::thread;
use std::time::Duration;
struct Shared {
done: bool,
waker: Option<Waker>,
}
pub struct Delay {
shared: Arc<Mutex<Shared>>,
}
impl Delay {
pub fn new(dur: Duration) -> Self {
let shared = Arc::new(Mutex::new(Shared { done: false, waker: None }));
let thread_shared = shared.clone();
// Stand-in for a reactor: a helper thread that "fires an event".
thread::spawn(move || {
thread::sleep(dur);
let mut s = thread_shared.lock().unwrap();
s.done = true;
if let Some(w) = s.waker.take() {
w.wake();
}
});
Delay { shared }
}
}
impl Future for Delay {
type Output = ();
fn poll(self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<()> {
let mut s = self.shared.lock().unwrap();
if s.done {
Poll::Ready(())
} else {
// Store the latest waker so the helper thread can wake us.
s.waker = Some(cx.waker().clone());
Poll::Pending
}
}
}
Two details are worth noticing:
- We overwrite the stored waker on every poll. The executor may hand us a different waker each time, and the most recent one is the one that counts.
- The waker is called after we set
done = true. Setting the flag first avoids a race where the future is polled and sees stale state.
Step 2: A Minimal block_on
use std::pin::pin;
use std::sync::Arc;
use std::task::Wake;
use std::thread::Thread;
struct ThreadWaker(Thread);
impl Wake for ThreadWaker {
fn wake(self: Arc<Self>) {
self.0.unpark();
}
}
fn block_on<F: Future>(fut: F) -> F::Output {
let mut fut = pin!(fut);
let waker: Waker = Arc::new(ThreadWaker(thread::current())).into();
let mut cx = Context::from_waker(&waker);
loop {
match fut.as_mut().poll(&mut cx) {
Poll::Ready(value) => return value,
Poll::Pending => thread::park(), // sleep until woken
}
}
}
Step 3: Run It
fn main() {
block_on(async {
println!("start");
Delay::new(Duration::from_secs(1)).await;
println!("one second later");
});
}
Run cargo run and you'll see the two messages one second apart. The main thread sleeps in park() during the wait and uses almost no CPU.
This is the same pattern real runtimes follow. They just add task queues, many threads, and a proper reactor. A runtime like Tokio is this idea, scaled up and polished.
Why Does Pin Exist?
Now the part everyone dreads. Let's make it simple.
The Problem: Self-Referential Futures
Look at this function:
async fn example() {
let data = [1, 2, 3];
let r = &data; // a reference to a local
something().await; // suspension point
println!("{:?}", r);
}
The reference r points to data. Both live inside the generated future, across the .await. So the future contains a pointer to itself.
Now imagine the future is moved to a different memory address, perhaps by passing it into a function or pushing it into a Vec. The bytes get copied, but r still holds the old address. That is a dangling pointer, which means undefined behavior.
Normally Rust's borrow checker prevents this kind of thing. But for compiler-generated state machines, it can't express "this struct borrows from itself." The solution is to make sure the future never moves once it has started being polled.
The Solution: Pin
Pin<P> wraps a pointer type like &mut T or Box<T>. It promises that the value pointed to will not be moved until it's dropped, unless the type opts out.
That's why poll takes self: Pin<&mut Self> and not &mut self. A plain &mut would let you do std::mem::swap and move the future. A pinned reference won't.
Unpin: The Escape Hatch
Most types don't care about being moved. u32, String, and Vec<T> are all fine. These types automatically implement the marker trait Unpin, and for them pinning is a no-op.
Futures generated by async blocks are typically !Unpin. That is what makes pinning matter for them.
Here is a summary of the common ways to pin something:
| Method | Where it lives | Notes |
|---|---|---|
Box::pin(fut) |
Heap | Safest and simplest. Costs one allocation. |
std::pin::pin!(fut) |
Stack (current scope) | No allocation. Stable since Rust 1.68. |
Pin::new(&mut fut) |
Wherever it is | Only works if fut: Unpin. |
Pin::new_unchecked |
Wherever it is | unsafe. You take responsibility for never moving it. |
When You'll Actually Meet Pin in Daily Code
- Storing futures in collections. You can't put different async blocks in a
Vecdirectly, since each has a unique type. UseVec<Pin<Box<dyn Future<Output = T>>>>. - Using
select!. Futures reused across loop iterations often needpin!. - Recursion. An async function that calls itself needs boxing, because the future would otherwise have infinite size:
use std::future::Future;
use std::pin::Pin;
fn countdown(n: u32) -> Pin<Box<dyn Future<Output = ()>>> {
Box::pin(async move {
if n > 0 {
println!("{n}");
countdown(n - 1).await;
}
})
}
Tip: If
Pinerrors confuse you, the fix is usuallyBox::pin(...)orpin!(...). Reach for those first. You rarely needunsafe.
How Tokio Schedules Your Tasks
Let's connect all of this to the runtime most people use. When you write:
#[tokio::main]
async fn main() {
let handle = tokio::spawn(async {
// some work
42
});
let result = handle.await.unwrap();
println!("{result}");
}
tokio::spawn wraps your future in a task and hands it to the scheduler. By default, Tokio uses a multi-threaded work-stealing scheduler: one worker thread per CPU core, each with its own queue. If a worker runs out of tasks, it steals from another.
Because your task might be resumed on a different thread after an .await, the future must be Send. That's the origin of the dreaded error:
error: future cannot be sent between threads safely
It usually means something non-Send (like an Rc, or a std::sync::MutexGuard) is held across an .await.
If you don't need multiple threads, Tokio also offers a single-threaded flavor:
#[tokio::main(flavor = "current_thread")]
async fn main() { /* ... */ }
Choosing a Runtime
| Runtime | Best for | Character |
|---|---|---|
| Tokio | Servers, networking, most production apps | Largest ecosystem, work-stealing, feature-rich |
| smol | Smaller apps and libraries | Lightweight and minimal |
| Embassy | Embedded, no_std |
Async on microcontrollers, no heap needed |
Common Async Rust Pitfalls (and How to Fix Them)
1. Blocking the Executor
This is the number one mistake.
async fn bad() {
std::thread::sleep(std::time::Duration::from_secs(5)); // blocks the worker thread!
}
While that thread sleeps, every task scheduled on it is stuck. Use the async version:
async fn good() {
tokio::time::sleep(std::time::Duration::from_secs(5)).await;
}
For CPU-heavy or genuinely blocking work (file parsing, hashing, a sync library), move it off the async workers:
let result = tokio::task::spawn_blocking(|| expensive_computation()).await?;
2. Holding a Lock Across .await
// Risky: std::sync::MutexGuard held over an await point
let guard = data.lock().unwrap();
do_io().await;
println!("{}", *guard);
This can cause deadlocks and Send errors. Prefer to drop the guard before awaiting:
let value = {
let guard = data.lock().unwrap();
guard.clone()
}; // guard dropped here
do_io().await;
If you truly need to hold a lock across an await, use tokio::sync::Mutex.
3. Forgetting to .await
Because futures are lazy, this does nothing:
async fn run() {
save_to_db(); // never runs; just creates a future
}
The compiler warns with unused_must_use. Treat that warning as an error.
4. Cancellation Surprises
Dropping a future cancels it at its last .await. That's powerful, but with tokio::select! a losing branch is simply dropped, and any partial work in it is lost. If an operation isn't cancel-safe, such as a read that has consumed some bytes into a temporary buffer, you can lose data. Always check the docs for cancel-safety notes on the functions you use inside select!.
5. Accidentally Sequential Code
let a = fetch_a().await;
let b = fetch_b().await; // starts only after a finishes
If the two are independent, run them concurrently:
let (a, b) = tokio::join!(fetch_a(), fetch_b());
Best Practices Checklist
- Keep the time between
.awaits short. Long synchronous stretches starve other tasks. - Use
spawn_blockingfor blocking or CPU-heavy work. - Don't hold
stdmutex guards across.await. - Prefer
join!for independent operations. - Use bounded channels so slow consumers apply backpressure.
- Avoid large stack arrays in async functions. Box them.
- Check cancel-safety before using futures inside
select!. - If you're writing a library, avoid tying users to one runtime when you can.
Conclusion
Async Rust looks mysterious mainly because the machinery is exposed rather than hidden. Once you see the pieces, they fit together neatly:
- An
async fnbecomes a state machine that implementsFuture. - Executors poll futures, and futures return
ReadyorPending. - A
Pendingfuture stores a waker, and the reactor calls it when progress is possible. - Pin exists so self-referential state machines never move in memory.
- Cooperative scheduling means you must never block a worker thread.
A good next step is to take the Delay and block_on code above and extend it. Add a task queue so block_on can run several futures, then compare your version with Tokio's source. Building even a toy executor will teach you more than reading a dozen articles.