Engineering Guides • Published August 21, 2026 • 25 min read

Regex Cheat Sheet for Developers: Character Classes, Anchors & Quantifiers

Master Regular Expressions with this complete developer cheat sheet. Syntax tables for character classes, anchors, quantifiers, lookarounds, JavaScript &.

Regex Cheat Sheet for Developers: Character Classes, Anchors & Quantifiers
An exhaustive reference guide and cheat sheet for Regular Expressions (Regex), covering character classes, anchors, quantifiers, lookarounds, capturing groups, lazy vs greedy matching, ReDoS vulnerability prevention, and real-world code recipes.
Regular expression pattern matching visualizer and execution engine
Figure 1: Pattern matching execution tree and evaluation path in modern NFA regex engines

Regular Expressions (commonly known as Regex) are one of the most powerful and flexible techniques in computer science. Whether you are validating user inputs in web forms, filtering log streams in Linux terminals, extracting structured entities from unstructured documents, or executing bulk search-and-replace refactoring in IDEs, mastering Regex turns hours of complex string manipulation code into a single, elegant line of pattern-matching rules.

However, because Regex syntax relies on dense meta-character tokens, developers often find regex patterns difficult to read, write, or debug.

This comprehensive technical guide serves as an exhaustive Regex manual and cheat sheet. Bookmark this page for quick reference during daily engineering work!

Test your regex patterns interactively with real-time match highlighting using our free Regex Tester & Debugger.


1. Character Classes & Metacharacters

Character classes represent specific categories of single characters.

| Token | Description | Equivalent Character Set | Matching Example |

| :--- | :--- | :--- | :--- |

| . | Any character except newline (unless s flag set) | Any printable or whitespace | a.c matches "abc", "a1c", "a#c" |

| \d | Any digit character | [0-9] | \d{3} matches "404" |

| \D | Any non-digit character | [^0-9] | \D+ matches "HTTP" |

| \w | Any word character (alphanumeric + underscore) | [a-zA-Z0-9_] | \w+ matches "dev_user42" |

| \W | Any non-word character | [^a-zA-Z0-9_] | \W matches "@", "!", " " |

| \s | Any whitespace character (space, tab, newline) | [ \t\r\n\f\v] | \s+ matches indentations |

| \S | Any non-whitespace character | [^ \t\r\n\f\v] | \S+ matches "DevToolAdda" |

Custom Bracket Sets & Character Ranges

Custom character classes defined inside square brackets [ ] match any single character contained within the brackets:

| Syntax | Description | Example Pattern | Matching Sample |

| :--- | :--- | :--- | :--- |

| [abc] | Match character a, b, or c | [aeiou] | Matches any lowercase vowel |

| [^abc] | Negated set: Match any character EXCEPT a, b, c | [^0-9] | Matches non-numeric characters |

| [a-z] | Range: Match ASCII character between a and z | [a-fA-F0-9] | Matches Hexadecimal digits |

| [0-9] | Numeric range: Match digit from 0 to 9 | [1-5] | Matches numbers 1, 2, 3, 4, 5 |

Unicode Property Escapes (\p{...})

Modern regular expression engines (with the u or v flag enabled) support Unicode Property Escapes, allowing matches based on Unicode character properties, scripts, and categories:

  • \p{L} or \p{Letter}: Matches any letter character across all written human scripts (Kanji, Cyrillic, Arabic, Devanagari, Latin).
  • \p{N} or \p{Number}: Matches any numeric symbol including superscripts and fractions.
  • \p{Emoji}: Matches Unicode emoji symbols.
  • \p{Script=Greek}: Matches characters from the Greek script.

2. Anchors & Boundary Matching

Anchors do not match physical characters; instead, they assert that the current search index matches a specific structural boundary.

| Anchor | Description | Example Pattern | Matching Result |

| :--- | :--- | :--- | :--- |

| ^ | Start of string (or line in multi-line m mode) | ^https | Matches "https://site.com", fails "site.com" |

| $ | End of string (or line in multi-line m mode) | \.png$ | Matches "image.png", fails "image.png.bak" |

| \b | Word boundary (position between \w and \W) | \bcat\b | Matches "cat", fails "category" or "bobcat" |

| \B | Non-word boundary (position not at word edge) | \Bcat | Matches "bobcat", fails "cat" |

| \A | Start of the entire string (ignores multiline m flag) | \AHELLO | Matches string start exclusively |

| \z | End of the entire string (ignores multiline m flag) | END\z | Matches string termination exclusively |


3. Quantifiers: Repeats & Counts

Quantifiers specify how many times the preceding token or group must repeat.

| Quantifier | Description | Repeat Range | Example |

| :--- | :--- | :--- | :--- |

| ** | Zero or more times | $0 \le n < \infty$ | gol matches "gl", "gol", "goool" |

| + | One or more times | $1 \le n < \infty$ | go+l matches "gol", "goool", fails "gl" |

| ? | Zero or one time (optional) | $0 \text{ or } 1$ | colou?r matches "color" and "colour" |

| {n} | Exactly n times | $n$ | \d{4} matches "2026" |

| {n,} | At least n times | $n \le count$ | \d{2,} matches "12", "12345" |

| {n,m} | Between n and m times inclusive | $n \le count \le m$ | [a-z]{3,5} matches "code", "dev" |

Greedy, Lazy (Non-Greedy), and Possessive Matching

By default, quantifiers in NFA regex engines are greedy: they consume as much input text as possible before evaluating subsequent tokens.

Adding a ? turns a quantifier into lazy mode, making it consume as little text as necessary to satisfy the match.

Adding a + turns a quantifier into possessive mode (supported in Java, PCRE, and Python regex module), which consumes text greedily but never backtracks.

Consider target text: <div>Header</div><div>Body</div>

  • *Greedy (<div>.</div>):** Matches from the first <div> all the way to the last </div>:

<div>Header</div><div>Body</div>

  • *Lazy (<div>.?</div>):** Stop at the very first closing </div>:

<div>Header</div>

  • *Possessive (<div>.+</div>):** Consumes the entire line to the end and fails immediately without backtracking to find the closing tag.

4. Groups, Backreferences & Advanced Lookarounds

Groups & Alternation Syntax

| Syntax | Group Type | Description |

| :--- | :--- | :--- |

| (abc) | Capturing Group | Groups tokens together and stores matched text in memory index ($1, $2). |

| (?:abc) | Non-Capturing Group | Groups tokens for quantifiers or alternation WITHOUT memory allocation. |

| (?<name>abc) | Named Capturing Group | Captures matched substring accessible by name (groups.name). |

| a|b | Alternation | Matches either expression a OR expression b (e.g., cat|dog). |

| \1, \2 | Backreference | Matches exact text captured by numbered group 1 or group 2 earlier. |

| \k<name> | Named Backreference | Matches exact text captured by named group name earlier. |

Advanced Lookaround Assertions

Lookarounds test assertions without consuming characters in the match output (zero-width assertions).

| Lookaround Type | Syntax | Meaning | Production Example |

| :--- | :--- | :--- | :--- |

| Positive Lookahead | (?=...) | Match if followed by ... | \d+(?=\s?USD) matches "100" in "100 USD" |

| Negative Lookahead | (?!...) | Match if NOT followed by ... | \d+(?!\s?USD) matches "100" in "100 EUR" |

| Positive Lookbehind | (?<=...) | Match if preceded by ... | (?<=\$)\d+ matches "50" in "$50" |

| Negative Lookbehind | (?<!...) | Match if NOT preceded by ... | (?<!\$)\d+ matches "50" in "£50" |


5. Regex Engine Flags

Execution flags modify pattern evaluation behavior across programming languages:

| Flag | Name | Function |

| :--- | :--- | :--- |

| g | Global | Find all occurrences instead of stopping after the first match. |

| i | Ignore Case | Case-insensitive matching (a matches A and a). |

| m | Multiline | Makes ^ and $ match start/end of each line (not just entire string). |

| s | dotAll | Allows dot (.) to match newline characters (\n). |

| u | Unicode | Full Unicode handling and 32-bit surrogate pair matching. |

| v | Unicode Sets | Enhanced set notation syntax and multi-character character classes (ES2024). |

| y | Sticky | Matches only at the exact lastIndex position in target text. |


6. Real-World Production Regex Recipes

1. RFC 5322 Compliant Email Address

^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$

2. URL Validation (HTTP/HTTPS)

^https?:\/\/(?:www\.)?[-a-zA-Z0-9@:%._\+~#=]{1,256}\.[a-zA-Z0-9()]{1,6}\b(?:[-a-zA-Z0-9()@:%_\+.~#?&\/=]*)$

3. Strong Password Enforcement

(Requires at least 8 characters, 1 uppercase, 1 lowercase, 1 number, and 1 special symbol)

^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[@$!%*?&])[A-Za-z\d@$!%*?&]{8,}$

4. IPv4 Address Validation

^(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)$

5. IPv6 Address Validation

^(?:[0-9a-fA-F]{1,4}:){7}[0-9a-fA-F]{1,4}$

6. ISO 8601 Date Format (YYYY-MM-DD)

^\d{4}-(?:0[1-9]|1[0-2])-(?:0[1-9]|[12]\d|3[01])$

7. Credit Card Number Validation (Visa, MasterCard, Amex)

^(?:4[0-9]{12}(?:[0-9]{3})?|5[1-5][0-9]{14}|3[47][0-9]{13})$

8. Semantic Versioning (SemVer 2.0)

^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)(?:-((?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*)(?:\.(?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*))*))?(?:\+([0-9a-zA-Z-]+(?:\.[0-9a-zA-Z-]+)*))?$

9. Hexadecimal Color Code (#RGB, #RRGGBB, #RRGGBBAA)

^#?([a-fA-F0-9]{3}|[a-fA-F0-9]{6}|[a-fA-F0-9]{8})$

10. URL Slug Identifier

^[a-z0-9]+(?:-[a-z0-9]+)*$

7. Multi-Language Code Snippets

JavaScript (ES6+ / ES2024)

// 1. Pattern Matching with RegExp
const emailPattern = /^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+.[a-zA-Z]{2,}$/;
console.log(emailPattern.test("dev.com")); // true

// 2. Named Capture Groups
const logPattern = /[(?<level>INFO|WARN|ERROR)]s(?<message>.+)/;
const logLine = "[ERROR] Database connection timeout";
const match = logPattern.exec(logLine);

if (match && match.groups) {
  console.log("Log Level:", match.groups.level);     // "ERROR"
  console.log("Message:", match.groups.message); // "Database connection timeout"
}

// 3. Search and Replace with Replacement Callback Function
const text = "Item price: $50 and tax: $5";
const converted = text.replace(/$(d+)/g, (match, amount) => {
  return "€" + (parseInt(amount) * 0.92).toFixed(2);
});
console.log(converted); // "Item price: €46.00 and tax: €4.60"

Python 3

import re

# 1. Searching & Capturing with Named Groups
log_pattern = re.compile(r'[(?P<level>INFO|WARN|ERROR)]s(?P<msg>.+)')
match = log_pattern.match("[WARN] High RAM utilization detected")

if match:
    print(match.groupdict())
    # Output: {'level': 'WARN', 'msg': 'High RAM utilization detected'}

# 2. Extracting All Matches
text = "Support contacts: support@site.com and sales@site.org"
emails = re.findall(r'[w.-]+@[w.-]+.w+', text)
print(emails) # ['support@site.com', 'sales@site.org']

# 3. Subbing with Callable
def uppercase_words(m):
    return m.group(0).upper()

s = "hello world from regex"
print(re.sub(r'w+', uppercase_words, s)) # "HELLO WORLD FROM REGEX"

Go (Golang)

package main

import (
	"fmt"
	"regexp"
)

func main() {
	re := regexp.MustCompile("(?P<first>\w+)\s(?P<last>\w+)")
	fmt.Println(re.MatchString("DevTool Adda")) // true

	match := re.FindStringSubmatch("John Doe")
	for i, name := range re.SubexpNames() {
		if i != 0 && name != "" {
			fmt.Printf("%s: %s
", name, match[i])
		}
	}
}

8. Preventing ReDoS (Catastrophic Backtracking)

ReDoS (Regular Expression Denial of Service) occurs when an inefficient regex pattern contains nested quantifiers applied to overlapping character sets (such as (a+)+$).

When an attacker passes a non-matching string like aaaaaaaaaaaaaaaaaaaaaaaaX, the engine explores millions of backtracking branches, causing CPU utilization to spike to 100% and locking up node.js or python server threads.

Best Practices to Guard Against ReDoS:

  1. Avoid nested quantifiers like (a+) or (x)*.
  2. Keep character sets mutually exclusive.
  3. Configure execution timeout limits on server regex evaluation engines.
  4. Test regex patterns against long non-matching strings before pushing to production.

Try out your regex patterns instantly in our interactive Regex Tester & Debugger.

Code editor displaying complex regular expression patterns and unit tests
Figure 2: Production unit tests validating complex regex patterns for email and URL parsing

Frequently Asked Questions

Q1. What is the difference between greedy and lazy quantifiers in regex?

Greedy quantifiers (, +, {n,m}) match as much text as possible before evaluating the remainder of the pattern. Lazy quantifiers (?, +?, {n,m}?) match as little text as necessary to satisfy the match.

Q2. How do lookahead and lookbehind assertions work in regex?

Lookarounds are non-capturing zero-width assertions. They verify whether a pattern precedes or follows the current position without consuming characters or adding them to the captured match result.

Q3. What is Regular Expression Denial of Service (ReDoS)?

ReDoS occurs when a poorly constructed regex with nested quantifiers (like (a+)+$) evaluates non-matching inputs, causing exponential back-tracking steps that freeze the CPU thread on server processes.

Q4. What do regex flags like g, i, m, s, u, v, and y mean in JavaScript?

g = global match all occurrences; i = case-insensitive; m = multi-line ^ and $ anchors; s = dotAll (. matches newlines); u = full Unicode support; v = set notation Unicode mode; y = sticky search at target index.

Test & Debug Regular Expressions

Struggling with complex regex patterns? Use our interactive tester with real-time matching, group highlights, and multi-language code generation.

Open Regex Tester