Understand the relationship
The reasoning behind the result
Window choice changes the observations
overlap starts 1,2,3,…; non-overlap starts 1,3,5,…
Overlapping counting measures every adjacent pair. Non-overlapping counting partitions one declared frame of the record into pairs and can leave one trailing symbol.
A circular overlapping record adds the boundary pair from the last symbol back to the first.
Ambiguity is not fractional evidence
A window containing an IUPAC ambiguity symbol is counted as unresolved. It is not distributed over possible canonical pairs using an unstated probability model.
Only retained canonical windows contribute to either position marginal; unresolved windows and trailing unpaired symbols remain disclosed outside that denominator.
Expected frequency is a declared independence comparison
E(XY) = pfirst(X)psecond(Y)
The expectation multiplies the first-position and second-position symbol shares of the retained canonical pair windows. It is a descriptive independence comparison, not an organism, genome or strand-symmetry reference.
Observed/expected is undefined when either position marginal is zero.
Composition does not establish mechanism
Pair enrichment can depend on region selection, strand convention, sequence length, coding frame and many biological processes.
This workbench does not infer methylation, mutation pressure, codon bias, motifs, statistical significance or causation.
Follow the numbers
Count overlapping pairs
- For AACGCGTTAA, advance a two-symbol window one base at a time.
- Retain AA, AC, CG, GC, CG, GT, TT, TA and AA: nine possible canonical windows.
- Count each of the 16 pair types, including zero rows.
- Divide each count by nine for observed frequency.
- Multiply each symbol's first-position share by the second symbol's second-position share for the displayed independence expectation.
The ratio describes these retained windows under one explicit marginal-product comparison; it is not a population enrichment test.
Quick guide
How to use this calculator
- Declare the polymer, record topology and whether adjacent windows overlap.
- Enter one sequence and keep ambiguity symbols when they are part of the record.
- Interpret observed/expected only against the retained-window marginal-product comparison, not as a universal biological baseline.
Calculation method
Calculation and interpretation
Keep literal pair counts, window choice, unresolved ambiguity and the chosen independence comparison visible instead of returning one unexplained dinucleotide percentage.
Observed pair frequency = canonical pair count ÷ canonical pair windows. Independence expectation = p(first-position symbol) × p(second-position symbol), using the retained canonical pair windows.
Worked example
Count overlapping pairs
The ratio describes these retained windows under one explicit marginal-product comparison; it is not a population enrichment test.
Observed pair frequency = canonical pair count ÷ canonical pair windows. Independence expectation = p(first-position symbol) × p(second-position symbol), using the retained canonical pair windows.
Supported inputs
Precision and limits
Literal word composition
The workflow counts entered symbols; it does not search motifs, align records or infer a reference genome.
One circular convention
Circular mode is restricted to overlapping adjacent pairs so the last-to-first boundary is unambiguous.
Ambiguity remains unresolved
Windows containing IUPAC ambiguity are disclosed and excluded from canonical pair frequencies.
No significance inference
Observed/expected ratios have no confidence interval, p-value or multiple-testing interpretation here.
Numerical support
One sequence of 2–200,000 nucleotide symbols is supported; all 16 canonical pair rows are retained.
Continue calculating
Related calculators