History Archaeology
34 pages · ~51 min✓ Reviewed
Builds on Rebase vs Merge. Next up: Git Hooks & Automation.
Part 1 · History Archaeology
Answering "When and Why Did This Change Happen?"
Every codebase eventually produces the same question: a line looks wrong, a feature stopped working, or a setting has a strange value, and nobody remembers how it got that way. Git has recorded the answer all along. The skill is knowing which part of the record to read. Searching history by hand, scrolling through git log and guessing, is slow and often wrong. Git ships purpose-built tools for each kind of question.
This chapter starts with a model of what history actually is: a graph of snapshots, not a flat timeline. Then it gives each tool its own job, so you can pick the right one from the question you are asking rather than from habit. The tools are listed below, each paired with the question it answers.
| Tool | The question it answers |
|---|---|
git log searching | Which commits match a message, an author, or a piece of a diff? |
git blame | Who last touched this line, and in which commit? |
git bisect | Which commit broke the build or introduced the bug? |
Pickaxe (-S, -G) | When did this string or pattern appear, change, or vanish? |
git reflog | Where did my "lost" commit go, and how do I get it back? |
By the end you will be able to trace a regression to the exact commit, read a line's real history without being fooled by reformatting, follow a renamed function across the project, and recover work after a bad reset or a botched rebase. The final part ties these together into one full investigation.
You need Git installed and a repository with a few dozen commits to practice on; any real project you have cloned works well. You should already be comfortable with git commit, git branch, git checkout and reading git diff output. Make a throwaway clone for the recovery and bisect exercises, so that experiments with reset --hard never touch work you care about.
Part 2 · 1. The Mental Model: Reading Git History
History is a graph of snapshots
Before reaching for any search tool, it helps to know what you are searching. Git history is not a timeline of edits. It is a set of commits, and each commit is an immutable snapshot of the whole project at one moment. A commit also records one or more parent pointers, an author, a committer and a message.
Full project tree
Never edited afterwards
One for a normal commit
Two for a merge
Author wrote it
Committer applied it
Why the change exists
Each commit points back at its parent, never forward at its children. Following those pointers gives a directed acyclic graph (DAG): arrows run one way, and you can never loop back to a commit you have already left. A branch name such as main is only a label stuck on one commit, and HEAD is a label that says which commit or branch you are standing on.
- 1HEADwhere you are
- 2mainlabel on C3
- 3C3newest commit
- 4C2parent of C3
- 5C1root commit
This explains the default behaviour of git log. It starts at HEAD, prints that commit, then follows parent links and prints each commit it reaches, newest first. In other words, it shows every commit reachable from HEAD. Commits that sit only on other branches stay hidden until you ask for them.
Author, committer and SHA identity
Every commit stores two people. The author is whoever wrote the change. The committer is whoever last put the commit into the history. For a plain commit they are the same person. They differ when a commit is rewritten or moved. git cherry-pick copies one commit onto another branch. git rebase replays a series of your commits on top of a new base. A maintainer who applies your patch becomes the committer, and you stay the author.
| Situation | Author | Committer |
|---|---|---|
| Ordinary commit on your own machine | You | You |
| A maintainer applies your patch | You | The maintainer |
| A teammate cherry-picks your commit onto a release branch | You | The teammate |
| Someone rebases your branch | You | The person who ran the rebase |
You can print both people yourself. In a format string, %an is the author name, %cn is the committer name, %ad is the author date, %h is the short SHA and %s is the subject line. The command below uses the first two and passes -s so only the format line is printed, not the diff.
git show -s --format='author=%an committer=%cn' HEAD%an = author name, %cn = committer name
author=Ana Rao committer=Sam Lee
Each commit is named by its SHA-1 hash, which is computed from the commit's data: the tree, the parent SHAs, the author and committer lines with timestamps, and the message. Git keys a commit on its data, not on what you meant by it. So two commits with the same code change still get different SHAs. A cherry-pick has a different parent and a new committer timestamp, so it always gets a new hash.
After a cherry-pick or rebase, the original SHA will not appear on the new branch. Search by message, author or content (the later sections cover how), not by the old hash.
Branches, reflog and the map view
Do not confuse the history with the reflog. Branches are editable labels: a reset, rebase or force-push can move them, and what you push is the public story. The reflog is a local diary of where HEAD actually pointed, including commits, resets, rebases and checkouts. It is never pushed, so it is the place to look when the visible history no longer shows what you did.
| Branches and log | Reflog | |
|---|---|---|
| What it is | Editable labels over the commit graph | Record of every real HEAD move |
| Shared with others | Yes, when pushed | Never, it stays in your clone |
| Role | The public story | Your local trail |
For day-to-day orientation, one command draws the whole map: every branch, drawn as a graph, one line per commit. Run it first whenever you land in an unfamiliar repository.
git log --oneline --graph --all--all includes every branch, not only what HEAD reaches
Once you have the map, pick the tool by the question you are asking. Choosing wrongly is the most common way to waste an hour.
| Tool | Answers |
|---|---|
| bisect | Which commit broke it |
| blame | Where a line came from |
| -S | When content appeared or disappeared |
| reflog | Recovers anything committed locally |
History is a graph of snapshots that you read from HEAD backwards. Ask your question first, then pick the tool.
Part 3 · 2. git log Searching: Messages, Diffs, and Authors
Filtering commits by message, author and date
git log is more than a history printer. It is a query tool: every flag narrows the list of commits it walks, and the cheapest filters look only at commit metadata. The first one to learn is --grep='regex', which keeps only commits whose message matches the pattern. By default it searches the history reachable from HEAD; add --all to search every branch and tag instead.
git log --oneline --grep='timeout' git log --oneline --grep='timeout' --all
Message search is only as good as the people who wrote the messages, but it costs almost nothing, so try it first. Filters stack: each one removes more commits. You can add --author to match the author name or email, --since and --until to bound the dates, and a pathspec after a double dash to keep only commits that touched that path.
| Option | Keeps commits that... | Example |
|---|---|---|
--grep='re' | have a message matching the regex | --grep='retry' |
--author='re' | were written by a matching name or email | --author='ana' |
--since / --until | fall inside a date window | --since=2024-01 |
-- path | changed that file or directory | -- src/net.c |
--all | are searched across every branch | --grep='fix' --all |
git log --oneline --since=2024-01 -- path/file.c git log --oneline --author='ana' --since=2024-01 --until=2024-06 --grep='cache'
Two switches change how matching behaves. -i makes --grep and --author case-insensitive, so --grep='timeout' -i also finds "Timeout" and "TIMEOUT". When you give several --grep options, a commit is kept if it matches any of them, which is OR. Add --all-match and the rule flips to AND: the message must match every pattern.
| Command | Combination rule | Result |
|---|---|---|
--grep='fix' --grep='seg' | OR (default) | messages mentioning fix, or seg, or both |
--grep='fix' --grep='seg' --all-match | AND | only messages mentioning both |
--all widens the search to every branch. It has nothing to do with --all-match, which is the switch that turns several --grep options into an AND. Mixing them up gives either too many or too few commits.
Reading the results: patches, counts, stats and formats
Once a filter narrows the list, you need to look inside the commits. -p prints the full patch for every commit, --stat prints only a summary of which files changed and by how many lines, and -<n> (for example -3) stops after n commits. Combining -<n> with a filter gives you the n newest matches.
The default output is verbose, so archaeology work usually uses --format (or its alias --pretty=format:) to print one compact line per commit. The useful placeholders are %h for the short hash, %an for the author name, %ad for the author date and %s for the subject. Add --date=short to keep the date to one tidy field.
| Placeholder | Prints |
|---|---|
%h | abbreviated commit hash |
%an | author name |
%ad | author date, shaped by --date= |
%s | subject line of the message |
git log -2 --stat --format='%h %an %ad %s' --date=short --grep='timeout'
Two newest matches, with a file summary
c41d9e2 Ana Rao 2024-03-14 Raise client timeout to 30s src/net.c | 4 ++-- config/app.ini | 2 +- 2 files changed, 3 insertions(+), 3 deletions(-) 7be2a10 Li Wei 2023-11-02 Add timeout to the retry loop src/net.c | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-)
Swap --stat for -p when the summary points at something interesting and you want the actual lines. For a pure list, drop both and keep only the format string.
git log --format='%h %an %ad %s' --date=short --grep='timeout' --allSearching the diffs: -G, -S, -L and --follow
Messages describe intent; the diff records what really happened. -G'regex' shows commits whose diff adds or removes a line matching the regex. It is a regex search over the changed lines of every commit, so it finds code even when the message says only "cleanup".
git log --oneline -G'retry.*max' -- src/Its sibling -S'string' is the quick inline pickaxe: it lists commits where the number of occurrences of the string changed, so a commit that adds or removes the string shows up, while one that merely edits a line containing it does not.
git log --oneline -S'MAX_RETRIES'| -G'regex' | -S'string' | |
|---|---|---|
| Matches on | changed diff lines | change in occurrence count |
| Argument | regex | plain string |
| Fires when | a touched line matches | the count differs before and after |
| Misses | nothing it can see in the diff | edits that keep the count equal |
This is only a preview of the contrast. The Pickaxe section later in this chapter works through -S and -G in depth.
To follow one function or one range of lines through time, use -L. It prints each commit that touched that range together with the patch for just that range, so you read the evolution of a function without the noise of the rest of the file.
git log -L :parse_url:src/net.c git log -L 10,25:src/net.c
Function name form and explicit line range form
A file's history stops where it was renamed, because the log looks for the path you gave. --follow continues the search under the old name. It relies on rename detection, and it works for a single file only.
git log --follow --oneline -- src/net.c--follow tracks one file path. Pointing it at a directory, or expecting it to work together with several pathspecs, does not follow renames the way you might hope.
Keeping the search fast
Message and date filters read only commit headers, but -G, -S and -L must compute a diff for each commit they inspect. On a long history that is slow. The remedy is to shrink the candidate set first: a pathspec and --since (or a range such as v1.0..HEAD) discard commits before the expensive diff search ever runs.
- 1Range or date--since=2024-01 or v1.0..HEAD
- 2Pathspec-- src/ keeps commits touching that path
- 3Message or author--grep, --author
- 4Diff search-G, -S, -L on what is left
git log --oneline -G'retry.*max' --since=2024-01 -- src/net/Date and path cut the work before -G runs
Use this table to pick the flag from the question you are asking.
| You want to know... | Reach for | Cost |
|---|---|---|
| which commits mention a topic | --grep | cheapest |
| what a person changed in a window | --author with --since | cheap |
| which commits touched lines matching a regex | -G | slower |
| when a string appeared or disappeared | -S | slower |
| how one function evolved | -L :func:file | per-range diff |
| a file's history across renames | --follow | per file |
Filter by metadata first, search diffs second, and print with a compact --format so the results fit in your notes.
Part 4 · 3. git blame: Who Touched This Line?
Reading a blame listing
When you have found a suspicious line and want to know where it came from, git blame <file> is the first tool to reach for. It prints the whole file with one row per line, and every row is stamped with the commit SHA, the author, the date and the line number, followed by the line's content.
git blame src/retry.py
One row per line of the file as it exists in your working tree
^a1b2c3d (Ana Rao 2023-02-11 10:02:41 +0530 1) import time ^a1b2c3d (Ana Rao 2023-02-11 10:02:41 +0530 2) 3f9c2ab1 (Ben Ortiz 2024-06-03 14:20:11 +0530 3) def retry(fn, tries=3): 3f9c2ab1 (Ben Ortiz 2024-06-03 14:20:11 +0530 4) delay = 0.5 9d07e5c4 (Ana Rao 2025-01-20 09:41:57 +0530 5) for n in range(tries):
The leading ^ on a SHA marks a boundary commit, meaning blame ran out of history to inspect and stopped there. It is the oldest point blame could see, not necessarily the true first author.
Zooming in with -L
Real files run to hundreds of lines, so you almost always want to narrow the view. -L 10,25 restricts the output to lines 10 through 25. If you would rather name a function than count lines, -L :funcname blames exactly that function's body, and git works out the start and end for you.
git blame -L 10,25 src/retry.py git blame -L :retry src/retry.py
A line range, then a single function
Last touch is not the original idea
Blame answers a narrow question: which commit most recently changed each line. It does not tell you who first had the idea behind that line. If someone fixed a typo in a line written two years ago, the typo fix gets the blame and the real author disappears from the listing.
| Question | Does blame answer it? |
|---|---|
| Which commit last changed this line? | Yes, that is exactly what it prints |
| Who first wrote this logic? | Not directly; you must look further back |
| Why was the line changed? | Only through the commit message you open next |
A blame result is a starting point for reading a commit, not a verdict about a person. The sections below show how to look past commits that only moved or reformatted code.
Cutting through reformat noise
The most common way blame misleads you is a mass reformat. Someone runs a formatter over the repository, and now every touched line points at a commit titled "Run black over everything". The logic you are chasing is hidden behind it. Three flags let blame look through that kind of change.
| Flag | What it ignores or detects | Use it when |
|---|---|---|
| -w | Whitespace-only differences | Indentation or spacing was reformatted |
| -M | Lines moved within the same file | A function was reordered or relocated inside the file |
| -C | Lines moved or copied from other files | Code was split out of, or pulled in from, another file |
The flags combine freely. Repeating -C makes the search wider: one -C looks in the same commit, a second also looks at the commit that created the file, and a third searches every commit. Each extra level costs more time, so add them only when the first pass still lands on a refactor.
git blame -w -M -C -L 40,60 src/health.py
Skip whitespace churn, follow in-file and cross-file moves
Skipping known-noisy commits for good
Flags help for one lookup, but a repository that once ran a formatter wants that commit ignored every time. Put the SHAs of such mass-reformat commits in a file named .git-blame-ignore-revs, one per line, with comments allowed. GitHub reads this file automatically in its blame view, and your local git can use it too once you point a config setting at it.
# .git-blame-ignore-revs # black formatting pass c0ffee1a2b3c4d5e6f708192a3b4c5d6e7f80910 # switch to tabs 5eed00d1e2f3a4b5c6d7e8f9012345678901abcd
git config blame.ignoreRevsFile .git-blame-ignore-revs
git blame --ignore-rev c0ffee1 src/health.pySet it once per clone, or pass a single SHA ad hoc with --ignore-rev
Lines that an ignored commit touched are blamed on the commit before it, which is usually the real change you were hunting.
If blame lands on a commit called "format", "lint" or "tidy", do not conclude the author wrote the logic. Re-run with -w -M -C or add the commit to .git-blame-ignore-revs before you accuse anyone.
Cost, evolution and renames
Blame gets slower on old files
For every line, blame must walk backwards through history until it finds the commit that produced that text. The work grows roughly with lines multiplied by history, so a huge file that has lived for ten years can take noticeably long, and -C makes it slower still. If you only care about recent changes, give it a revision range so it stops early.
git blame v2.0..HEAD -- src/health.py
Only history after v2.0 is inspected; older lines show as boundary commits
When you want evolution, not origin
Blame shows a single snapshot: one commit per line. If the real question is how a piece of code changed over time, git log -p -L is clearer. It follows a line range or a function through history and prints the patch for every commit that touched it, in order.
git log -p -L :retry:src/retry.py git log -p -L 10,25:src/retry.py
Every change to one function, newest first
| git blame | git log -p -L | |
|---|---|---|
| Shows | Latest commit per line | Every commit that touched the range |
| Best for | Who and when, right now | How and why it evolved |
| Output size | One row per line | One patch per commit |
Renamed files
Blame follows a file through renames on its own while walking history, so git blame new_name.py normally credits the original author of lines that predate the rename. The catch is the path you type. It must exist at the revision you are blaming. If you blame an older revision, use the old path with it, for example git blame v1.0 -- old_name.py. The --follow flag belongs to git log, not to blame, so use it there: git log --follow -p -- new_name.py.
Output for scripts and editors
The default listing is pretty but awkward to parse. --porcelain prints a stable machine format: a header per group of lines with the SHA, original and final line numbers, then author, author-mail, author-time, summary and filename fields, and finally the line itself prefixed by a tab. Repeated commits only get their full header once. --line-porcelain repeats the full header for every line, which is wasteful but easier to read in a simple script.
git blame --line-porcelain -L 3,4 src/retry.py | grep -E '^(author|summary) 'Pull just the author and subject for each line in the range
GUI views and a worked hunt
Blame without the terminal
| Where | How to open it | What you get |
|---|---|---|
| VS Code | Command palette, then Toggle Blame Annotations | Author and age beside every line while you edit |
| GitHub | Blame button on a file page | Per-line blame with a side diff for each commit |
GitHub's view is especially handy because you can step into the commit behind a line and keep going back through its earlier versions without typing a command.
A suspect line, end to end
Suppose line 42 of src/health.py reads timeout = 5, and a health check has started failing. You want to know who set it and why. A plain git blame -L 42,42 points at a formatting commit, which tells you nothing. Adding -w -C sees past the reformat and lands on the real change.
| Command | SHA | Author | Commit subject |
|---|---|---|---|
| git blame -L 42,42 | c0ffee1a | Dev Rao | Run black over the repo |
| git blame -w -C -L 42,42 | 7e41d0c9 | Ana Rao | Lower health timeout to 5s |
- 1Spot the linetimeout = 5 looks wrong
- 2blame -w -C -Lskip whitespace and moves
- 3Copy the SHA7e41d0c9
- 4git show <sha>read the whole commit
git blame -w -C -L 42,42 src/health.py git show 7e41d0c9
Blame finds the commit; show gives its message and full diff
commit 7e41d0c9 Author: Ana Rao <ana@example.com> Date: Mon Mar 3 11:08:00 2025 +0530 Lower health timeout to 5s Probes were stalling the load balancer for 30s on cold starts. - timeout = 30 + timeout = 5
The message explains the intent, and the diff shows what else changed alongside the line. That context is what you were after from the start, and the SHA from blame is just the key to reach it.
If blame reports a path error or an unexpectedly young history, check whether the file was renamed or moved. Pass the old path together with the older revision, and use git log --follow when you want the whole story.
Part 5 · 4. git bisect: Binary Search for the Broken Commit
Why bisect finds the culprit in a handful of tests
When something worked at version 1.0 and is broken today, the culprit is one commit somewhere in between. git bisect finds it with a binary search: it checks out the commit in the middle of the range, you say whether the bug is present, and half of the remaining commits are ruled out. Repeat until one commit is left. Finding one bad commit among n takes O(log n) tests instead of up to n.
The gap between linear and logarithmic grows quickly. Even a range of ten thousand commits needs only about fourteen test runs.
| Commits in range | Bisect tests (about) | Worst case checking one by one |
|---|---|---|
| 10 | 4 | 10 |
| 100 | 7 | 100 |
| 1,000 | 10 | 1,000 |
| 10,000 | 14 | 10,000 |
To see the idea in code, the Python below models a linear history as a list. commits[hi] is known to be bad and the search repeatedly tests the middle. This is the same loop Git runs for you.
def find_first_bad(commits, is_bad): lo, hi = 0, len(commits) - 1 # commits[hi] is known bad rounds = 0 while lo < hi: mid = (lo + hi) // 2 rounds += 1 if is_bad(commits[mid]): hi = mid else: lo = mid + 1 return commits[lo], rounds commits = [f"c{i}" for i in range(1, 17)] is_bad = lambda c: int(c[1:]) >= 11 culprit, rounds = find_first_bad(commits, is_bad) print(f"culprit {culprit} after {rounds} tests")
16 commits, the bug arrives in c11
culprit c11 after 4 testsOne commit you know is good, one you know is bad, and a reliable way to tell them apart. Everything else is arithmetic.
Running a bisect: by hand and unattended
A session starts with git bisect start. You then tell Git one bad commit (usually the current one, so the commit is optional) and one known-good commit such as a release tag. As soon as it has both, Git checks out the midpoint. You build or run the program, then answer with git bisect good or git bisect bad. Git picks the next midpoint, and when only one commit is left it prints it as the first bad commit.
git bisect start git bisect bad # HEAD is broken (or: git bisect bad <commit>) git bisect good v1.0 # this tag worked # Git checks out a midpoint; you test it, then: git bisect good # or: git bisect bad # ...repeat until Git prints: <sha> is the first bad commit
Manual flow, one answer per round
Answering by hand gets tedious, so let a command answer for you. git bisect run <command> runs the command at every midpoint and reads its exit code: 0 means good, any other code from 1 to 127 means bad, and the special code 125 means skip this commit because it can't be tested. Codes above 127 abort the whole search.
| Exit code | Meaning for bisect |
|---|---|
| 0 | good, bug absent |
| 1 to 124, 126, 127 | bad, bug present |
| 125 | skip, this commit can't be tested |
| 128 and above | abort the whole bisect |
The shortcut git bisect start HEAD v1.0 gives the bad and good commits in one line. Followed by git bisect run make test, the entire search runs unattended and ends with the culprit's SHA, author and message.
git bisect start HEAD v1.0 git bisect run make test # ...Git builds and tests each midpoint on its own, then prints the culprit git bisect reset
Fully unattended search
When the project does not always build, wrap the test in a small script so a build failure turns into a skip instead of a false "bad".
#!/bin/sh make >/dev/null 2>&1 || exit 125 # can't build: skip this commit ./run-tests.sh # exit 0 = good, non-zero = bad
check.sh, used as: git bisect run ./check.sh
If the test command exits non-zero because the code did not compile, Git records that commit as bad and may blame an innocent commit. Return 125 for anything that stops you from testing.
Skip, save, look around, and clean up
Sometimes a midpoint can't be judged: it doesn't build, or it's a half-finished commit from the middle of a refactor. Run git bisect skip and Git moves to a neighbouring commit instead. Skipping has a cost. If the culprit sits among skipped commits, Git can only say that the first bad commit is one of a listed group, so the result may become approximate.
Every answer you give is recorded. git bisect log prints the session so far, and redirecting it into a file saves it. Later, or on another machine, git bisect replay <file> replays those answers. Replaying is also handy for editing the log by hand to correct one wrong answer and carry on.
Mid-search you can see what is still in play with git bisect visualize (alias git bisect view). It opens a log or graph viewer showing only the remaining candidate commits. On a machine without a graphical viewer, pass log options such as --oneline.
git bisect skip # this commit can't be tested git bisect log > bisect.log # save the session git bisect replay bisect.log # replay it later git bisect visualize --oneline # list remaining candidates
Session control
| Command | What it does |
|---|---|
| git bisect run <cmd> | automates good/bad using the exit code |
| git bisect skip | sets aside a commit you can't test |
| git bisect log | prints the answers so far; redirect it to save |
| git bisect replay <file> | re-applies a saved session |
| git bisect visualize | shows the remaining candidates (view is an alias) |
| git bisect reset | ends the session and restores your checkout |
The last step matters most. During a bisect Git has left you on a detached HEAD at whatever commit it tested last. git bisect reset returns you to the branch you started on. git bisect reset <commit> takes you to a different commit instead.
Without it you keep working on a detached HEAD at the culprit or a midpoint. New commits made there belong to no branch and are easy to lose.
Trusting the result, and bisecting more than bugs
Bisect is only as honest as the test you give it. One wrong answer sends every later round into the wrong half, and Git will still print a confident culprit. A flaky test, a dependency that changed on your machine since the old commit, or a stale build directory can all make the search lie. Make the check deterministic, rebuild from clean, and treat the answer as a lead.
If the test passes and fails at random on the same commit, bisect output is noise. Run it several times on a known commit first, then confirm the culprit with git show <sha>.
Nothing in the method says the property must be a crash. Any property that is true at old commits and false at new ones works, including performance. Let a script treat a benchmark above a threshold as bad and Git will find the commit that made things slower. The Python below reuses find_first_bad and commits from the bisect section above, with a made-up benchmark in place of a real timing.
def bench(c): return 100 if int(c[1:]) < 9 else 240 # milliseconds is_slow = lambda c: bench(c) > 150 culprit, rounds = find_first_bad(commits, is_slow) print(f"first slow commit {culprit} ({bench(culprit)} ms) after {rounds} benchmarks")
bad means slower than 150 ms
first slow commit c9 (240 ms) after 4 benchmarks
Bisect needs a good/bad answer, not a crash. Pick a measurable property, make the test trustworthy, finish with git bisect reset, and verify the culprit with git show.
Part 6 · 5. Pickaxe: -S and -G Content Archaeology
Searching by what the diff contains
Most git log filters look at metadata: who wrote a commit, when, and what the message says. The pickaxe options look inside the changes instead. They answer a different question: which commit made this piece of text show up, or disappear? You don't need to remember a message or a filename, only a fragment of the code.
git log -S'string' lists the commits where the number of occurrences of that string changed. A commit that adds the string raises the count. A commit that deletes it lowers the count. Either way the commit is listed.
git log -G'regex' lists commits whose diff has an added or removed line matching the regular expression. It doesn't compare counts. It asks only whether any changed line matches.
| -S'string' | -G'regex' | |
|---|---|---|
| Argument | Literal string by default | Always a regex |
| What it compares | Occurrence count before and after, across the whole diff | Each added or removed line, one at a time |
| Finds | The string appeared or disappeared | A line matching the pattern changed |
| Speed | Generally faster | Generally slower |
| Precision | Coarse, count-based | Exact for patterns |
There is one exception to the literal default. Adding --pickaxe-regex makes -S read its argument as a regex. It still counts occurrences, so it is not the same as -G. With -G a modified line always matches. With -S --pickaxe-regex the count of matches must change.
The difference is easiest to see on a small history. The program below imitates the two rules on three commits to one line of config code. Commit a1 adds timeout=30. Commit b2 adds a retry argument to the same line. Commit c3 changes the value to 60. This is a simulation of the matching logic, not Git itself.
import re commits = { 'a1': ([], ['fetch(url, timeout=30)']), 'b2': (['fetch(url, timeout=30)'], ['fetch(url, timeout=30, retry=2)']), 'c3': (['fetch(url, timeout=30, retry=2)'], ['fetch(url, timeout=60, retry=2)']), } for sha, (old, new) in commits.items(): before = ''.join(old).count('timeout=30') after = ''.join(new).count('timeout=30') s = 'yes' if before != after else 'no' g = 'yes' if any(re.search('timeout=30', l) for l in old + new) else 'no' print(f'{sha} -S:{s:<3} -G:{g}')
Count change (-S) versus matching changed line (-G)
a1 -S:yes -G:yes b2 -S:no -G:yes c3 -S:yes -G:yes
Commit b2 is the interesting one. It edited a line containing timeout=30, but the count stayed at one. -S skips it and -G reports it.
Putting pickaxe to work
The classic use is tracing an API rename. Suppose a helper called oldFunctionName was replaced years ago and you want to know when and by whom. Searching for the old name lists both ends of its life: the commit that introduced it and the commit that removed it. Every commit in between that used it without changing the count stays out of the list, so the output is short.
git log -S'oldFunctionName' --oneline
Introduction and removal both appear
e41c9a0 Remove oldFunctionName in favour of fetchJson
7b20d3f Add oldFunctionName helper to api clientGit lists commits newest first, so the removal comes before the introduction. You now have two SHAs bracketing the helper's lifetime.
Two options change what you get back. A pathspec after -- limits the search to certain files. That is how you find where a default value changed. By default each listed commit shows only the files that matched. Adding --pickaxe-all makes Git show the commit's whole changeset when you ask for a patch. That is useful when the string moved because of a change elsewhere, such as a rename in one file and an update in another.
git log -S'timeout=30' --oneline -- config/ git log -S'oldFunctionName' --pickaxe-all -p --stat
Narrow by path, or widen the display to every file in the commit
Pickaxe only gives you the SHA. To see what actually changed, run git show on that commit, restricted to the file you care about so unrelated edits don't bury it.
git show e41c9a0 -- src/api/client.pyRead the matching change in one file
- 1Pick a fragmenta name or value you remember
- 2git log -S'fragment' --onelineadd a pathspec if you know the area
- 3Note the SHAsone adds the string, one removes it
- 4git show <sha> -- <file>read the change in context
Cost, limits and the invisible commit
Pickaxe has to look at the diff of every commit it considers, so it is slower than a message search. It searches the entire history by default, which on a large repository can take minutes. Bound the work with --since, with a range such as v1.0..HEAD, or with a pathspec. Of the two options, -S is usually faster because it counts a plain string. -G runs a regex over every changed line, so it is slower but can match patterns.
| Way to bound the hunt | Example | Effect |
|---|---|---|
| Date | --since=2024-01-01 | Skips older commits |
| Range | v1.0..HEAD | Only commits after that tag |
| Pathspec | -- config/ | Only diffs touching those files |
| Cheaper option | -S instead of -G | Plain string count, no regex |
git log -S'timeout=30' --oneline v1.0..HEAD -- config/ git log -G'timeout\s*=\s*[0-9]+' --oneline --since=2024-01-01
A bounded -S and a bounded -G
The second command shows a case where -G is the better tool. If the value could change from 30 to 60, you want every commit that touched a timeout line, not only the commits where one particular literal came or went.
A commit that changes the lines around a string but leaves the number of occurrences identical is invisible to -S. Moving a call to a new line, or adding an argument next to timeout=30, changes nothing it can see. When a search comes back empty but you are sure the code was touched, rerun it with -G.
-S asks whether the count of a string changed. -G asks whether a changed line matches a regex. Start with -S because it is quick, switch to -G when you need a pattern or when -S stays silent, and always bound the range.
Part 7 · 6. Reflog: Recovering 'Lost' Commits
The reflog: a local diary of every move
Branches are labels that move, and the history you see with git log only shows what is reachable from where you stand now. The reflog is a different record. It is Git's private diary of where HEAD (and every branch tip) pointed over time, so it still remembers places you have since left.
Run git reflog and you see one line per move of HEAD: each commit, checkout, merge, reset and rebase step. Run git reflog <ref>, for example git reflog feature/login, to see the diary of one branch alone.
| Branch history (git log) | Reflog | |
|---|---|---|
| Records | Commits reachable from a ref | Every place a ref has pointed |
| Survives a reset --hard | No, the commits vanish from view | Yes, the old tip is still listed |
| Shared by push | Yes, it is the public story | Never, it stays in your clone |
| Lifetime | As long as a ref reaches it | About 90 days, 30 for unreachable entries |
Entries expire. By default an entry that still points at a reachable commit is kept for 90 days, and an entry pointing at a commit that is no longer reachable from any branch is kept for 30 days. When garbage collection later prunes those entries, the commits they protected can finally be deleted.
A commit you have 'removed' by reset, rebase or branch deletion stays on disk as long as some reflog entry references it. Nothing is truly lost until git gc prunes the reflog entries and then the objects.
Reading entries and rescuing a botched reset
Every reflog line has a shorthand you can feed to other commands. HEAD@{n} means where HEAD was n moves ago, HEAD@{'2 hours ago'} picks by time, and main@{n} reads that branch's own diary instead of the one for HEAD.
| Form | Meaning | Counts through |
|---|---|---|
HEAD@{3} | Where HEAD was 3 moves ago | Reflog entries, local only |
HEAD@{'2 hours ago'} | Where HEAD was at that time | Reflog timestamps |
feature@{1} | The previous tip of branch feature | That branch's own reflog |
HEAD~3 | Third ancestor of the current commit | Parent chain, shared history |
The last row is the classic trap: HEAD~3 follows parents in the commit graph, while HEAD@{3} follows your own movements. They usually name different commits.
Imagine you made three commits on feature/login and then ran git reset --hard HEAD~3 by mistake. The branch now looks as if the work never happened, but the reflog shows the whole story. The reset itself is HEAD@{0}, and the commit you were on just before it is HEAD@{1}.
2a8e5c0 HEAD@{0}: reset: moving to HEAD~3
9c7d3b8 HEAD@{1}: commit: Add password reset
b52a6e1 HEAD@{2}: commit: Add session cookie
7d0c9f4 HEAD@{3}: commit: Add login form
2a8e5c0 HEAD@{4}: checkout: moving from main to feature/loginWhat git reflog shows after the bad reset
The pre-disaster entry is the one just below the reset: 9c7d3b8, the tip you had before. You can either move the current branch back to it, or leave things alone and park the old tip on a new branch.
git reflog git reset --hard HEAD@{1} # or keep the current state and park the old tip safely: git checkout -b rescue 9c7d3b8
Two ways back to 9c7d3b8
Rebases, deleted branches and the last resort
A rebase rewrites your commits, so a rebase gone wrong feels like lost work. The reflog makes it easy to undo, because Git logs a rebase (start) entry when it begins. The entry directly older than it is your original, pre-rebase HEAD.
f31b9a7 HEAD@{0}: rebase (finish): returning to refs/heads/feature/login
e90d2c4 HEAD@{1}: rebase (pick): Add login form
c4a7710 HEAD@{2}: rebase (start): checkout main
9c7d3b8 HEAD@{3}: commit: Add password resetThe line below rebase (start) is the original tip
Here git reset --hard HEAD@{3} puts the branch back exactly as it was before the rebase touched it.
Deleted branches work the same way. A branch is only a pointer, so deleting it removes the label, not the commits. Find the last SHA the branch had in the reflog (git branch -D also prints it as 'was 9c7d3b8'), then recreate the label with git branch <name> <sha>. The branch comes back instantly.
git reflog git branch feature/login 9c7d3b8
Resurrecting a deleted branch
If the reflog entries have already expired, there is one tool left. git fsck --lost-found scans the object database itself for dangling commits, meaning commits no ref or reflog entry reaches. It works only until git gc actually deletes those objects.
Limits and prevention habits
The reflog has hard limits. It is stored in your clone and is never pushed, so it cannot rescue commits that a teammate lost after a force-push. Their reflog lives on their machine, and yours only knows what your own HEAD did.
It also records commits only. If you ran git reset --hard with uncommitted edits, those edits were never in a commit, so there is no entry to find. Your best chances are outside Git: editor swap or backup files and the local history built into many IDEs.
The cheapest protection is a habit. Before a risky rewrite, put a label on the current tip with git branch backup-before-rebase, so it stays reachable and visible. When you rebase with uncommitted changes, --autostash stashes them first and reapplies them afterwards.
git branch backup-before-rebase git rebase --autostash main git config gc.reflogExpire 180.days git config gc.reflogExpireUnreachable 60.days
Backup label, autostash, and longer retention
Retention is configurable: gc.reflogExpire sets how long reachable entries live and gc.reflogExpireUnreachable does the same for unreachable ones.
HEAD~2 walks two parents back in the commit graph. HEAD@{2} is where HEAD was two moves ago. Read the reflog listing before you pass either one to reset --hard.
Reflog entries expire and git gc removes the objects after that. Recover as soon as you notice the loss, and make the rescue branch before you do anything else.
Part 8 · 7. Comparisons, Trade-offs, and Common Mistakes
Pick the tool from the question
Every tool in this chapter answers a different question, and most wasted time comes from reaching for the familiar one. Start from what you are asking, not from the command you remember. The table below is the whole decision in one place.
| Your question | Tool | Typical command |
|---|---|---|
| Which commit broke the build? | bisect | git bisect run ./repro.sh |
| Who wrote this line? | blame | git blame -w -C -L 10,25 file |
| When did this symbol appear? | pickaxe -S | git log -S'name' --oneline |
| When did lines matching X change? | log -G | git log -G'req.*retry' -p |
| Where did my commit go? | reflog | git reflog |
When the question is fuzzy, run it through a short filter. Do you have a test that fails today and passed once? Bisect. Do you already hold a suspicious line? Blame. Are you hunting for a name or a pattern across time? Pickaxe. Did you lose something you committed yourself? Reflog.
Pick the tool from the question you are asking, not from muscle memory. The wrong tool still gives an answer, just not to your question.
Cost and precision: what each search really pays
Each tool spends a different resource. Bisect is cheap in Git work but expensive in your time, because every round runs your test. Blame, pickaxe and -G spend CPU walking history. A message search spends almost nothing, because it never opens a diff.
| Tool | What it spends | Cost shape |
|---|---|---|
| bisect | Test runs, one per round | O(log n) rounds; the test is the expense |
| blame | History walk for one file | Full walk back through that file's past |
| log -S / -G | Diff scan of each commit | Every commit in the range gets diffed |
| log --grep | Commit messages only | Metadata only, the cheapest |
The log-scale claim is easy to underestimate, so here it is in numbers. Doubling the history adds only one more test run.
import math for n in (10, 100, 1000, 100000): print(f"{n:>6} commits -> {math.ceil(math.log2(n))} test runs")
Rounds bisect needs for n commits
10 commits -> 4 test runs 100 commits -> 7 test runs 1000 commits -> 10 test runs 100000 commits -> 17 test runs
Speed trades against precision. The cheap searches trust your commit messages or a simple count; the slow ones look at actual content. Read the chain left to right as cheaper and shallower to costlier and deeper.
- 1--grepfast, only as good as the messages
- 2-Sfast, fires when the count changes
- 3-Gslower, exact regex on changed lines
- 4blame -Cslowest, deepest analysis
A good workflow starts at the cheap end. Try --grep first, then narrow with a pathspec and date before paying for -G or blame -C.
Mistakes that blame the wrong thing
The first three mistakes all end the same way: you confidently name the wrong commit. Each has a cheap guard.
A formatter sweep touches every line, so plain blame points at whoever ran it. Re-run with -w to skip whitespace and -C to follow moved code, and check blame.ignoreRevsFile before you accuse the author.
git blame -w -C -L 40,60 src/app.py git blame --ignore-revs-file .git-blame-ignore-revs src/app.py
Look past the formatting sweep
Bisect leaves you on a detached HEAD at some midpoint. If you walk away, the repo looks broken and new commits land on no branch. Always finish with git bisect reset, which returns you to the branch you started on.
git bisect start HEAD v2.1
git bisect run ./repro.sh
git bisect resetEvery session ends with reset
A test that sometimes passes on a bad commit makes bisect discard the culprit for good, and the final answer is confidently wrong. Make the check deterministic first. When bisect names a commit, reproduce the failure by hand there and at its parent, then confirm with git show.
Mistakes with scope, local state and habits
The next mistakes come from forgetting where a piece of history lives and how far a search reaches. Reflog notation and parent notation look alike but read from different records.
| HEAD@{n} | HEAD~n | |
|---|---|---|
| Reads | Reflog: where HEAD was n moves ago | Parent chain: n generations back |
| Scope | Local to your clone | Part of the commit graph, pushed with it |
| After a reset | Still lists the discarded commits | Skips them |
The model below replays four commits, a reset --hard back to c2, and a new commit c5. The two notations agree at one step and then split.
parents = {"c5": "c2", "c4": "c3", "c3": "c2", "c2": "c1", "c1": None}
reflog = ["c5", "c2", "c4", "c3", "c2", "c1"]
def tilde(commit, n):
for _ in range(n):
commit = parents[commit]
return commit
for n in (1, 2):
print(f"HEAD@{{{n}}} = {reflog[n]} HEAD~{n} = {tilde('c5', n)}")Newest reflog entry first
HEAD@{1} = c2 HEAD~1 = c2
HEAD@{2} = c4 HEAD~2 = c1After a reset the two point at different commits, as HEAD@{2} and HEAD~2 do above. Use HEAD@{n} to undo your own moves and HEAD~n to talk about the shared graph.
git log -S or -G with no pathspec or date diffs every commit in the repository and can run for minutes. Bound it with -- path, --since, or a range like v1.0..HEAD.
Reflog is per clone and never pushed. After a colleague rebases and force-pushes, your blame shows the rewritten commits, and the old ones exist only in their reflog. You cannot recover their lost commits from your machine.
Two habits prevent most of this. A graph alias gives you the whole picture in one command, and an ignore file keeps formatting sweeps from stealing blame.
git config --global alias.pl 'log --oneline --graph --all' git config blame.ignoreRevsFile .git-blame-ignore-revs echo '<sha-of-format-commit>' >> .git-blame-ignore-revs
Alias and ignore-revs setup
Run git pl before any rescue, add every formatting sweep's SHA to .git-blame-ignore-revs the day it merges, and write the pathspec first when you type a pickaxe.
Part 9 · 8. Case Study & Summary: A Full Archaeology Session
The report: a regression with no known start
Here is the situation the whole chapter has been preparing you for. A customer writes in: exports that used to include a trailing summary row now come out without it. Nobody on the team remembers touching the exporter, the release notes say nothing, and there are hundreds of commits since the last release. You have no suspect, no date and no author. This is where you apply the toolkit in order, with each tool handing its result to the next.
- 1Reproduce and boundfailing check, good tag, bad HEAD
- 2Automatebisect run pins the SHA
- 3Understandshow + blame -w -C -L
- 4Related historylog -S and log -p -L
- 5Safe fixreflog if the rebase goes wrong
Step 1: reproduce and bound the search
Bisect is only as honest as the check behind it, so first turn the customer's complaint into something a machine can answer. Write a small script that builds the export and fails if the summary row is missing. Run it by hand on your current checkout to confirm it fails, and on the last release tag to confirm it passes. Those two facts are your bounds: the tag is good, HEAD is bad.
git bisect start
git bisect bad HEAD
git bisect good v2.1Git replies with how many commits remain and about how many steps it needs
If the check is flaky or depends on your machine, bisect will confidently name the wrong commit. Run it several times on both ends first, and verify the culprit by hand at the end.
Pin the culprit, then read the story around it
Step 2: automate with bisect run
Hand-testing every midpoint is slow and error-prone. Give Git the script and let it drive: it checks out the midpoint, runs the script, and reads the exit code. 0 means good, any non-zero code from 1 to 127 except 125 means bad, and 125 means this commit cannot be tested, so skip it. With 640 commits between the tag and HEAD, the search takes about log2(640), roughly 10 test runs.
git bisect run ./repro.sh
a41c9e7d2b0f3c85e6d1f7a9b2c4d6e8f0a1b3c5 is the first bad commit commit a41c9e7d2b0f3c85e6d1f7a9b2c4d6e8f0a1b3c5 Author: Ana Ruiz <ana@example.com> Date: Tue Mar 4 16:20:11 2025 +0000 Simplify exporter row loop
That answers where: one SHA. Run git bisect reset straight away so you return to your branch instead of staying on a detached HEAD.
Step 3: understand the commit
A SHA is not an explanation. Read the whole commit first with git show, because the message and the other files in the same change often explain the intent. Then use blame on the exact lines you now care about, ignoring whitespace (-w) and detecting code copied from other files (-C), so a reformat or a move does not steal the credit.
git show a41c9e7 git blame -w -C -L 40,70 src/exporter.py
If blame points at a bulk reformat rather than a real edit, the line's true origin is further back. Check .git-blame-ignore-revs before drawing conclusions about who did what.
Related history, and rescuing a botched rebase
Step 4: find the related history
The culprit commit probably removed something the summary row depended on. Use the pickaxe to ask when that symbol appeared and when it vanished: -S lists the commits where the number of occurrences of the string changed, so both the introduction and the removal show up. Then trace the whole function with -L to watch it evolve patch by patch.
git log -S'append_summary' --oneline -- src/ git log -p -L :export_rows:src/exporter.py
If a commit changed lines near the symbol without changing its count, -S stays silent. Switch to -G'append_summary' to match any changed line against a regex.
Step 5: ship the fix, and recover if the rebase goes wrong
You write the fix and rebase your branch onto the latest main. Conflicts pile up and you resolve one of them badly. Nothing is lost: the reflog records every move of HEAD, including the pre-rebase tip. Find that entry, check it, and hard-reset to it.
git reflog
git reset --hard HEAD@{5}The number comes from the reflog listing, not from a guess
HEAD@{5} means five moves ago in your local reflog, while HEAD~5 means five parents back in the commit chain. Read the listing and pick the entry by its message before resetting. Also remember the reflog only holds commits, never uncommitted edits.
The summary rule and the six one-liners
After one full session you can see that each tool answers a different question. Choose the tool by the question, not by habit.
| Question | Answered by | Tool |
|---|---|---|
| WHERE did it break? | Binary search over commits | bisect |
| WHEN and WHO? | Last touch of a line, first appearance of content | blame, log -S |
| HOW did it evolve? | Patch-by-patch changes to text or a function | log -G, log -L |
| WHAT happened locally? | Every move of HEAD in this clone | reflog |
The cheatsheet worth memorizing
git bisect run ./repro.sh # pin the first bad commit git blame -w -C -L 40,70 file # who, ignoring whitespace and moves git log -S'symbol' --oneline # when the count of a string changed git log -G'regex' --oneline # when a changed line matched git log -p -L :func:file # how one function evolved git reset --hard HEAD@{n} # undo, after reading git reflog
Six commands, one per job
| One-liner | Use it when |
|---|---|
| bisect run | you know a good and a bad point but not the culprit |
| blame -w -C | you need context for specific lines |
| log -S | a symbol or string appeared or disappeared |
| log -G | a changed line matched a pattern but the count stayed the same |
| log -L | you want a function's or range's history |
| reflog + reset --hard | a rebase, reset or deletion went wrong locally |
Prevention and the final reminder
Make the next hunt cheaper
Every tool in this chapter works better on clean history. Small commits shrink the diff bisect lands on, so the culprit is readable at a glance. Meaningful messages make --grep and git show useful. Tags on releases hand you a reliable good boundary for step one, the one thing you otherwise have to guess.
- Commit one logical change at a time, with its test.
- Write a subject that says what changed, and a body that says why.
- Tag every release so
goodis never a guess. - Keep
.git-blame-ignore-revsfor bulk reformats.
Git history is append-mostly and reflogged. With bisect, blame, the pickaxe, log -L and the reflog, you can find where something broke, who and when, how it changed, and how to get back. Only uncommitted work and commits beyond the reflog's retention are truly out of reach.
Part 10 · Check yourself
Quiz
Each question gives you a situation and asks you to predict what Git will do or to spot what went wrong. Try to answer in your head before reading the answer.
A bisect session runs git bisect start HEAD v1.0 followed by git bisect run ./check.sh. The script crashes with exit code 1 on every commit, including ones that predate the bug because a build tool is missing on your machine. What does bisect report, and what should you have done?
- It reports the oldest commit after v1.0 (or a nonsense commit) as the culprit, because every exit code that is non-zero (other than 125) means bad.
- The test measured your environment, not the bug, so the result is a lie.
- Make the script exit
125when the build tool is missing so those commits are skipped, and reproduce manually at the reported commit withgit showbefore trusting it. - Run
git bisect resetafterwards to get back to your branch.
#!/bin/sh command -v make >/dev/null || exit 125 make test
A line in app.py shows git blame pointing at a commit titled "Run formatter over repo" by Dana. Dana says they never wrote that logic. What is the next command, and why is accusing Dana premature?
- Blame reports the last commit that touched each line, and a reformat touches every line it re-indents.
- Rerun with
git blame -w -C -M -L 40,55 app.pyso whitespace-only edits and moved or copied lines are looked through. - Better still, add the formatting commit's SHA to
.git-blame-ignore-revsso it stops stealing blame. - Then use
git show <sha>on the real commit to read it in context.
git blame -w -C -M -L 40,55 app.py git config blame.ignoreRevsFile .git-blame-ignore-revs
A file starts with timeout = 30. Commit A changes the line to timeout = 30 # seconds. Commit B changes it to timeout = 60. Which commits does git log -S'timeout = 30' --oneline list, and which does git log -G'timeout' --oneline list?
-S'timeout = 30'lists only commit B, plus the commit that first added the line. In commit A the line is rewritten but the stringtimeout = 30still appears once on each side, so the count is unchanged.- Commit B removes the string, so the count drops from one to zero and
-Sfires. -G'timeout'lists A and B (and the introducing commit), because both diffs add or remove a changed line matching the regex.- Lesson:
-Ssees counts change,-Gsees matching lines change.
You run git reset --hard HEAD~3 and then realize you needed those commits. A teammate suggests git reset --hard HEAD~3 in reverse, i.e. HEAD~-3. Why does that not work, and what does?
HEAD~nwalks the parent chain from where HEAD is now, and after the reset those three commits are no longer ancestors of HEAD.HEAD@{n}walks the reflog, the local record of where HEAD actually was.- Run
git reflog, find the entry from before the reset, and rungit reset --hard HEAD@{1}(or the exact SHA). - This only recovers committed work; uncommitted edits discarded by
--hardwere never in a commit and the reflog cannot return them.
git reflog
git reset --hard HEAD@{1}A teammate force-pushes a rebased branch and their old commits vanish from the remote. You ask them to "just check your reflog", but they cloned fresh yesterday. Can you recover the lost commits from your own clone?
- Only if your clone had fetched those commits before the force-push: the reflog is per-clone and is never pushed.
- Check
git reflog origin/featureorgit reflog show origin/featurein your clone to find the old tip SHA, thengit branch rescue <sha>. - A teammate's fresh clone has no history of those commits at all, so their reflog cannot help.
- If nobody still has the objects,
git fsck --lost-foundon a clone that once saw them is the last resort.
Summary
- Pick the tool from the question:
bisectfinds WHERE it broke,blameand-Sfind WHO and WHEN,-Gand-Lshow HOW it evolved,reflogshows what happened locally. git bisect runneeds a trustworthy script: exit 0 is good, non-zero is bad, 125 skips, and always finish withgit bisect reset.git blameshows the last toucher, so use-w -C -Mand a.git-blame-ignore-revsfile before blaming anyone.-Sfires when a string's occurrence count changes,-Gfires when a changed line matches a regex; bound both with a pathspec or range.git log -L :func:filetraces a function's whole evolution, which is often clearer than blame.- The reflog is local and records where HEAD really moved, so
HEAD@{n}rescues botched resets and rebases but never uncommitted work. - Clean history makes every tool sharper: small commits, clear messages, and tags on releases.