Unique identifiers are a foundational primitive in software systems. Every database record, every API resource, every distributed event needs a way to be referenced uniquely across systems and time. UUID and CUID are two of the most widely used approaches, and choosing between them has real implications for your database performance, URL readability, and system architecture.

UUID: The Standard

UUID (Universally Unique Identifier) is defined by RFC 4122 and has four versions in common use. UUID v4 is the most prevalent: a randomly generated 128-bit number represented as a 36-character hexadecimal string in the format xxxxxxxx-xxxx-4xxx-yxxx-xxxxxxxxxxxx.

UUID v4 generates identifiers with approximately 5.3 × 1036 possible values, making collision probability effectively zero for practical applications. Two independently generated UUID v4s have roughly a 1 in 5.3 undecillion chance of matching.

UUID v4 characteristics:

  • 36 characters (32 hex + 4 hyphens)
  • Randomly generated — no correlation between sequential IDs
  • Widely supported across every database and programming language
  • No information leaked about generation time or sequence

UUID and Database Performance

The random nature of UUID v4 creates a significant problem for database indexes. B-tree indexes — the structure most relational databases use for primary keys — perform best with sequential values. Random UUIDs cause index fragmentation: each new insert lands at a random position in the index rather than appending to the end. At scale, this creates excessive page splits and cache misses.

This is why UUID v7 (timestamp-first, sequential within a millisecond) was introduced. UUID v7 retains global uniqueness while ordering identifiers chronologically, which dramatically improves B-tree index efficiency. For high-volume OLTP workloads, the difference between v4 and v7 can be a 2–5× improvement in insert throughput.

CUID: The Alternative

CUID (Collision-resistant Unique Identifier) was designed with URL-safety, sequential ordering, and horizontal scalability in mind. A typical CUID2 looks like clh3vq4bt0000356sg4c3khlp — 24 lowercase alphanumeric characters with no hyphens.

CUID2 characteristics:

  • 24 characters (URL-safe, no hyphens needed)
  • Timestamp prefix makes IDs roughly time-sortable
  • Includes a fingerprint component to reduce collisions in distributed systems
  • Designed specifically for web applications with multiple concurrent generators

The Practical Decision

Use UUID v4 when you need RFC-standard identifiers, maximum ecosystem support, or when the randomness is a security feature (preventing enumeration of resources). Most existing systems, ORMs, and database tools have native UUID support.

Use UUID v7 when you are building a new high-throughput system and database index performance matters. The time-sequential nature eliminates index fragmentation while maintaining global uniqueness.

Use CUID2 when you want shorter, URL-safe identifiers without hyphens and do not need RFC compliance. It is a good choice for new web applications where the identifier will appear in URLs.

UUIDv7 Changes the Performance Argument

Most of the case against UUID primary keys is really a case against version 4 specifically. A v4 UUID is random across its whole length, so consecutive inserts land in scattered positions within a B-tree index, causing page splits and poor cache locality.

Version 7, standardised in 2024, puts a 48-bit millisecond timestamp at the front and fills the rest with randomness. Values generated in sequence sort in roughly the order they were created, so inserts append to the end of the index the way an auto-increment integer does. It keeps the distributed-generation property that made UUIDs attractive and removes the index behaviour that made database administrators dislike them.

Storage Costs More Than People Expect

Stored as text, a UUID is 36 characters — 32 hex digits plus four hyphens. Stored as binary it is 16 bytes. On a table with 50 million rows and three indexes referencing the key, that difference is measured in gigabytes, and it is repeated in every foreign key that points at the table.

If you use UUIDs in PostgreSQL, use the native uuid type rather than text. In MySQL, BINARY(16) with conversion at the boundary costs more application code but less than half the storage of CHAR(36).

Neither Is a Security Token

A v4 UUID has 122 random bits, which is plenty of entropy, but the specification does not require a cryptographically secure source and some implementations have historically used weak ones. More importantly, identifiers leak: they appear in URLs, logs, referrer headers, and support tickets.

Use a purpose-built token for anything that grants access — a password reset link, a session, an API key — generated from a cryptographically secure random source and stored hashed. Identifiers name things; tokens authorise them, and the two jobs should not share a value.

Frequently Asked Questions

Can two UUIDs collide?

In practice, no. Version 4 draws from 2¹²² possible values, around 5.3 undecillion. Generating a billion per second for a century leaves the probability of a single collision negligible. Collisions that do occur in the wild almost always trace to a broken random source rather than to the odds.

Should the primary key itself be a UUID?

A common pattern keeps an internal auto-increment integer as the physical primary key and adds an indexed UUID as the public identifier. You get compact joins internally and non-guessable identifiers externally. With UUIDv7 the case for a single UUID key is much stronger than it was.

What about ULID?

ULID solves the same problem as UUIDv7 — a time-ordered, distributed-friendly identifier — and predates it. It encodes to 26 characters in Crockford base32, which is shorter and case-insensitive. UUIDv7 has the advantage of being a formal standard with native database support arriving, so new projects have less reason to reach for ULID than they did.

Are sequential integer IDs a security problem?

They allow enumeration: if your order is /orders/1042, someone will try 1041. That is an access-control failure rather than an identifier failure — the fix is checking authorisation on every request. Non-guessable identifiers reduce the damage from a missing check but are not a substitute for having the check.

Which should I pick for a distributed system?

UUIDv7 for most cases. It generates without coordination, sorts by creation time, and indexes well. Reach for v4 only when you specifically want no time information embedded in the identifier, since the timestamp in v7 is readable by anyone who receives it.

Related tools: the hash generator for the fingerprinting side of identity, and the password entropy calculator for reasoning about how much randomness a token actually needs. On the difference between encoding and protecting a value, see why Base64 is not encryption, or browse the developer tools.

Generate UUIDs instantly using the UUID Generator — supporting both v4 and v7 formats for direct use in development workflows.