Course topics

By WebNest Studio

Python Tutorial

Regular Expressions in Python

A regular expression (regex) is a pattern that describes text: "a 6-digit PIN code", "an email address", "a date like 27-09-2026". Python's re module uses these patterns to search, validate, extract and replace text — tasks that would need many lines of string methods.

This lesson covers pattern syntax, the main re functions (match, search, fullmatch, findall, finditer, sub, split), groups and named groups, flags, compiling patterns, and practical validation and extraction examples.

Pattern Syntax

The essentials:

  • Characters: . any character, \d digit, \w word character, \s whitespace (capitals negate: \D, \W, \S).
  • Sets: [aeiou], ranges [a-z0-9], negation [^0-9].
  • Quantifiers: * (0+), + (1+), ? (0 or 1), {3}, {2,4}; add ? for non-greedy (.*?).
  • Anchors: ^ start, $ end, \b word boundary.
  • Groups: ( ) capture, (?P<name> ) named, (?: ) non-capturing, | alternation; lookarounds (?= ), (?! ).

Raw Strings

Always write patterns as raw strings — r"\d+" — so Python does not interpret backslashes before the regex engine sees them.

re Functions

re.search finds the first match anywhere; re.match only at the start; re.fullmatch requires the whole string to match (best for validation). re.findall returns all matches (or groups) as a list; re.finditer yields match objects; re.sub replaces (the replacement can reference groups or be a function); re.split splits on a pattern. Match objects provide group(), groups(), groupdict(), start() and end().

Flags and Compiling

re.IGNORECASE, re.MULTILINE (^/$ per line), re.DOTALL (. matches newlines) and re.VERBOSE (allow whitespace and comments in patterns). re.compile(pattern) creates a reusable pattern object — clearer when a pattern is used many times.

Examples

search, match, fullmatch, findall and finditer

Python
import re

text = "Order WN-1042 shipped on 2026-09-27; order WN-1043 on 2026-09-30."
print(re.search(r"WN-\d+", text).group())
print(re.match(r"WN-\d+", text), re.match(r"Order", text).group())
print(bool(re.fullmatch(r"\d{6}", "411001")), bool(re.fullmatch(r"\d{6}", "41100")))
print(re.findall(r"WN-(\d+)", text))
print(re.findall(r"(\d{4})-(\d{2})-(\d{2})", text))
for m in re.finditer(r"\d{4}-\d{2}-\d{2}", text):
    print(m.group(), m.start(), m.end())
Output
WN-1042
None Order
True False
['1042', '1043']
[('2026', '09', '27'), ('2026', '09', '30')]
2026-09-27 25 35
2026-09-30 54 64

Named groups, sub with backreferences and functions, split

Python
import re

log = "2026-09-27 10:15:02 ERROR payment failed user=asha amount=2999"
pattern = re.compile(r"(?P<date>\S+) (?P<time>\S+) (?P<level>[A-Z]+) (?P<msg>.*)")
m = pattern.match(log)
print(m.group("level"), "|", m.groupdict()["msg"])

print(re.sub(r"(\d{4})-(\d{2})-(\d{2})", r"\3/\2/\1", "Due 2026-09-27"))
print(re.sub(r"\d{4}(?=\d{4}$)", "****", "Card 1234567812345678"))
print(re.sub(r"\d+", lambda m: str(int(m.group()) * 2), "3 pens and 10 books"))
print(re.split(r"[,;\s]+", "python, java;  sql go"))
print(re.sub(r"\s+", " ", "  too    many   spaces ").strip())
Output
ERROR | payment failed user=asha amount=2999
Due 27/09/2026
Card 12345678****5678
6 pens and 20 books
['python', 'java', 'sql', 'go']
too many spaces

Validating common Indian and web formats

Python
import re

validators = {
    "email": re.compile(r"[\w.+-]+@[\w-]+(\.[\w-]+)+"),
    "mobile": re.compile(r"(\+91[\s-]?)?[6-9]\d{9}"),
    "pincode": re.compile(r"[1-9]\d{5}"),
    "pan": re.compile(r"[A-Z]{5}\d{4}[A-Z]"),
    "strong_password": re.compile(r"(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[^\w\s]).{8,}"),
}
samples = {
    "email": ["asha@webnest.in", "bad@", "a.b+tag@mail.co.uk"],
    "mobile": ["9876543210", "+91 9876543210", "1234567890"],
    "pincode": ["411001", "011001"],
    "pan": ["ABCDE1234F", "ABCD1234F"],
    "strong_password": ["Passw0rd!", "password"],
}
for kind, values in samples.items():
    results = [f"{v}={'ok' if validators[kind].fullmatch(v) else 'no'}" for v in values]
    print(f"{kind:<16}", ", ".join(results))
Output
email            asha@webnest.in=ok, bad@=no, a.b+tag@mail.co.uk=ok
mobile           9876543210=ok, +91 9876543210=ok, 1234567890=no
pincode          411001=ok, 011001=no
pan              ABCDE1234F=ok, ABCD1234F=no
strong_password  Passw0rd!=ok, password=no

Flags: IGNORECASE, MULTILINE, DOTALL and VERBOSE

Python
import re

print(re.findall(r"python", "Python PYTHON python", flags=re.IGNORECASE))
notes = "TODO: write tests\nnote: fine\nTODO: deploy"
print(re.findall(r"^TODO: (.*)$", notes, flags=re.MULTILINE))
html = "<p>line one\nline two</p>"
print(re.search(r"<p>(.*)</p>", html), re.search(r"<p>(.*)</p>", html, re.DOTALL).group(1).split("\n"))
print(re.findall(r"<.+?>", "<b>bold</b>"), re.findall(r"<.+>", "<b>bold</b>"))

price = re.compile(r"""
    (?P<currency>Rs\.|₹)\s*   # currency symbol
    (?P<amount>[\d,]+)        # digits with commas
""", re.VERBOSE)
m = price.search("Total: ₹ 1,299 only")
print(m.group("currency"), int(m.group("amount").replace(",", "")))
Output
['Python', 'PYTHON', 'python']
['write tests', 'deploy']
None ['line one', 'line two']
['<b>', '</b>'] ['<b>bold</b>']
₹ 1299

Common Mistakes

  • Writing patterns without raw strings, so "\b" becomes a backspace character.
  • Using re.match when you meant re.search (match only checks the start).
  • Validating with search instead of fullmatch, accepting strings that merely contain a valid part.
  • Greedy .* swallowing too much — use non-greedy .*? or a more specific class.
  • Parsing HTML or complex formats with regex instead of a proper parser.

Key Points to Remember

  • Write patterns as raw strings: r"\d+".
  • search (anywhere), match (start), fullmatch (whole string), findall/finditer (all), sub, split.
  • Groups, named groups and backreferences extract and rearrange text.
  • Flags: IGNORECASE, MULTILINE, DOTALL, VERBOSE; compile reused patterns.

Practice the examples

Change an input, predict the result, then compare it with the output. Explain why the result changes.