Measuring where a transcription factor binds DNA, and what loss-of-function looks like at fragment-scale resolution.
The earlier views show that TP53 is frequently mutated, that CDKN1A can also be repressed through promoter methylation, and that the downstream pathway can fail through more than one route. ChIP-seq (chromatin immunoprecipitation followed by sequencing) provides the direct physical evidence: where a specific protein is bound on the genome at the moment the experiment was fixed.
The assay cross-links proteins to DNA, shears the DNA, immunoprecipitates fragments bound by the target protein, and sequences the fragment ends. Mapping reads back to the reference produces a pileup: a histogram of inferred fragment coverage. The summit can be narrow, but the assay still sees fragments, not a protein footprint base by base.
ChIP-seq is never read in isolation. Any antibody pull-down picks up background: accessible chromatin, copy-number changes, repetitive regions, sticky sequences, and library amplification. The input control (sonicated DNA from the same cells, no immunoprecipitation) estimates that local background. A good transcription-factor peak rises above input, appears in biological replicates, and usually has the expected motif near the summit.
RRRCWWGYYY RRRCWWGYYY, the canonical two-half-site p53 response-element pattern. The letters are degenerate, so a motif match supports a peak but does not prove binding by itself.Figure 1. A 10-kb window around the CDKN1A transcription start site. Three tracks stacked on the same coordinate axis: wildtype TP53 ChIP, the R175H mutant ChIP, and the input control. Hover anywhere on the track panel for a vertical crosshair and per-track signal readout. The gene body at the bottom shows exon structure; the tall peak sits about 2 kb upstream of the TSS, at the canonical TP53-responsive element.
The wildtype track shows a sharp pileup centered on the p53 response element, roughly 200 bp wide and rising well above local background. Input is flat. The mutant track is also near background: no convincing peak at the CDKN1A promoter in this synthetic example. R175H is a canonical structural hotspot. It sits near the zinc-stabilized DNA-binding surface, so the mutation can distort the fold needed for sequence-specific binding. Contact hotspots such as R248 and R273 are different: they more directly disrupt residues that touch DNA.
A pipeline cannot eyeball peaks. Peak calling turns a continuous pileup into a discrete list of bound regions consumed by downstream analyses: motif enrichment, target annotation, differential binding. Tools such as MACS compare ChIP signal to local background, estimate enrichment, and report peak coordinates, summits, p-values, fold enrichment, and false-discovery measures. In practice, the threshold should be applied to adjusted q-values and reproducibility, not to raw p-values alone.
Figure 2. The same wildtype signal from Figure 1, re-expressed as Benjamini-Hochberg adjusted local enrichment scores. Bins above the line are called as belonging to a peak; contiguous runs of called bins are merged into discrete peak regions, outlined in green. Drag the threshold down and noise fragments appear. Drag it up and the peak narrows to the summit. Synthetic data, MACS-style calling.
Every ChIP-seq paper depends on thresholds, but there is no universal raw \(-\log_{10}(p)\) cutoff. Too permissive publishes noise; too stringent misses biology. A defensible workflow controls false discovery, removes known problematic regions, checks library complexity and fraction of reads in peaks, and asks whether independently prepared replicates rank the same peaks highly. ENCODE-style transcription-factor pipelines use IDR for that last step: reproducible peaks matter more than impressive single-sample scores.
Genome-wide comparison shows whether R175H is a broken key or a differently shaped one. The figure compares WT and mutant ChIP signal at ten canonical TP53 targets.
Figure 3. Differential binding summary for ten validated TP53 target genes. The log₂ fold-change column uses a diverging scale centered on zero: values left of center (orange→red) mean the mutant binds less. Hover any row for the numeric values. Sort by fold change to see that the mutant binds essentially none of its canonical targets.
The mutant is not simply a weaker version of wildtype in this example; it is near background at most canonical targets. A few of the strongest wildtype sites retain marginal signal. With the DNA-binding surface destabilized, R175H often fails to hold a stable, sequence-specific contact with p53 response elements, and the dependent transcriptional program weakens.
ChIP-seq answers where a protein was bound at fixation. It does not say whether binding caused transcriptional change. Binding can be functional, decoy, or inconsequential. ChIP-seq is therefore paired with RNA-seq from the same cells so bound sites can be matched against expression changes.