Python Tutorial
Regular Expressions in Python
A regular expression (regex) is a pattern that describes text: "a 6-digit PIN code", "an email address", "a date like 27-09-2026". Python's re module uses these patterns to search, validate, extract and replace text — tasks that would need many lines of string methods.
This lesson covers pattern syntax, the main re functions (match, search, fullmatch, findall, finditer, sub, split), groups and named groups, flags, compiling patterns, and practical validation and extraction examples.
Pattern Syntax
The essentials:
- Characters:
.any character,\ddigit,\wword character,\swhitespace (capitals negate:\D,\W,\S). - Sets:
[aeiou], ranges[a-z0-9], negation[^0-9]. - Quantifiers:
*(0+),+(1+),?(0 or 1),{3},{2,4}; add?for non-greedy (.*?). - Anchors:
^start,$end,\bword boundary. - Groups:
( )capture,(?P<name> )named,(?: )non-capturing,|alternation; lookarounds(?= ),(?! ).
Raw Strings
Always write patterns as raw strings — r"\d+" — so Python does not interpret backslashes before the regex engine sees them.
re Functions
re.search finds the first match anywhere; re.match only at the start; re.fullmatch requires the whole string to match (best for validation). re.findall returns all matches (or groups) as a list; re.finditer yields match objects; re.sub replaces (the replacement can reference groups or be a function); re.split splits on a pattern. Match objects provide group(), groups(), groupdict(), start() and end().
Flags and Compiling
re.IGNORECASE, re.MULTILINE (^/$ per line), re.DOTALL (. matches newlines) and re.VERBOSE (allow whitespace and comments in patterns). re.compile(pattern) creates a reusable pattern object — clearer when a pattern is used many times.
Examples
search, match, fullmatch, findall and finditer
import re
text = "Order WN-1042 shipped on 2026-09-27; order WN-1043 on 2026-09-30."
print(re.search(r"WN-\d+", text).group())
print(re.match(r"WN-\d+", text), re.match(r"Order", text).group())
print(bool(re.fullmatch(r"\d{6}", "411001")), bool(re.fullmatch(r"\d{6}", "41100")))
print(re.findall(r"WN-(\d+)", text))
print(re.findall(r"(\d{4})-(\d{2})-(\d{2})", text))
for m in re.finditer(r"\d{4}-\d{2}-\d{2}", text):
print(m.group(), m.start(), m.end())
WN-1042
None Order
True False
['1042', '1043']
[('2026', '09', '27'), ('2026', '09', '30')]
2026-09-27 25 35
2026-09-30 54 64
Named groups, sub with backreferences and functions, split
import re
log = "2026-09-27 10:15:02 ERROR payment failed user=asha amount=2999"
pattern = re.compile(r"(?P<date>\S+) (?P<time>\S+) (?P<level>[A-Z]+) (?P<msg>.*)")
m = pattern.match(log)
print(m.group("level"), "|", m.groupdict()["msg"])
print(re.sub(r"(\d{4})-(\d{2})-(\d{2})", r"\3/\2/\1", "Due 2026-09-27"))
print(re.sub(r"\d{4}(?=\d{4}$)", "****", "Card 1234567812345678"))
print(re.sub(r"\d+", lambda m: str(int(m.group()) * 2), "3 pens and 10 books"))
print(re.split(r"[,;\s]+", "python, java; sql go"))
print(re.sub(r"\s+", " ", " too many spaces ").strip())
ERROR | payment failed user=asha amount=2999
Due 27/09/2026
Card 12345678****5678
6 pens and 20 books
['python', 'java', 'sql', 'go']
too many spaces
Validating common Indian and web formats
import re
validators = {
"email": re.compile(r"[\w.+-]+@[\w-]+(\.[\w-]+)+"),
"mobile": re.compile(r"(\+91[\s-]?)?[6-9]\d{9}"),
"pincode": re.compile(r"[1-9]\d{5}"),
"pan": re.compile(r"[A-Z]{5}\d{4}[A-Z]"),
"strong_password": re.compile(r"(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[^\w\s]).{8,}"),
}
samples = {
"email": ["asha@webnest.in", "bad@", "a.b+tag@mail.co.uk"],
"mobile": ["9876543210", "+91 9876543210", "1234567890"],
"pincode": ["411001", "011001"],
"pan": ["ABCDE1234F", "ABCD1234F"],
"strong_password": ["Passw0rd!", "password"],
}
for kind, values in samples.items():
results = [f"{v}={'ok' if validators[kind].fullmatch(v) else 'no'}" for v in values]
print(f"{kind:<16}", ", ".join(results))
email asha@webnest.in=ok, bad@=no, a.b+tag@mail.co.uk=ok
mobile 9876543210=ok, +91 9876543210=ok, 1234567890=no
pincode 411001=ok, 011001=no
pan ABCDE1234F=ok, ABCD1234F=no
strong_password Passw0rd!=ok, password=no
Flags: IGNORECASE, MULTILINE, DOTALL and VERBOSE
import re
print(re.findall(r"python", "Python PYTHON python", flags=re.IGNORECASE))
notes = "TODO: write tests\nnote: fine\nTODO: deploy"
print(re.findall(r"^TODO: (.*)$", notes, flags=re.MULTILINE))
html = "<p>line one\nline two</p>"
print(re.search(r"<p>(.*)</p>", html), re.search(r"<p>(.*)</p>", html, re.DOTALL).group(1).split("\n"))
print(re.findall(r"<.+?>", "<b>bold</b>"), re.findall(r"<.+>", "<b>bold</b>"))
price = re.compile(r"""
(?P<currency>Rs\.|₹)\s* # currency symbol
(?P<amount>[\d,]+) # digits with commas
""", re.VERBOSE)
m = price.search("Total: ₹ 1,299 only")
print(m.group("currency"), int(m.group("amount").replace(",", "")))
['Python', 'PYTHON', 'python']
['write tests', 'deploy']
None ['line one', 'line two']
['<b>', '</b>'] ['<b>bold</b>']
₹ 1299
Common Mistakes
- Writing patterns without raw strings, so "\b" becomes a backspace character.
- Using re.match when you meant re.search (match only checks the start).
- Validating with search instead of fullmatch, accepting strings that merely contain a valid part.
- Greedy .* swallowing too much — use non-greedy .*? or a more specific class.
- Parsing HTML or complex formats with regex instead of a proper parser.
Key Points to Remember
- Write patterns as raw strings: r"\d+".
- search (anywhere), match (start), fullmatch (whole string), findall/finditer (all), sub, split.
- Groups, named groups and backreferences extract and rearrange text.
- Flags: IGNORECASE, MULTILINE, DOTALL, VERBOSE; compile reused patterns.
Practice the examples
Change an input, predict the result, then compare it with the output. Explain why the result changes.