Bioinformatics & Sequence Analysis

Dinucleotide Composition and Expected-Frequency Workbench

Count all 16 canonical dinucleotides in an overlapping, non-overlapping or circular record and compare observed frequencies with the retained windows' first- and second-position marginal expectations.

Biology · experimental measurements

Keep literal pair counts, window choice, unresolved ambiguity and the chosen independence comparison visible instead of returning one unexplained dinucleotide percentage.

Private calculations in your browser · explicit inputs and model boundaries
Example preview · Overlapping DNA recordLiteral canonical pair counts
AA2 windows
AC1 windows
AG0 windows
AT0 windows
CA0 windows
CC0 windows
CG2 windows
CT0 windows
GA0 windows
GC1 windows
GG0 windows
GT1 windows
TA1 windows
TC0 windows
TG0 windows
TT1 windows

Bars retain all 16 canonical dinucleotides in alphabetic order. Ambiguous windows remain in the summary and are not allocated among bars.

  1. 1EnterProvide the known values
  2. 2CalculateResults update automatically
  3. 3VerifyReview the details and units
Try an example

One raw or single-record FASTA sequence. IUPAC ambiguity is retained as unresolved pair windows rather than divided among bases.

Calculation result

Enter valid values to see the result.

Your entries are calculated in this browser and are not submitted to 365CALCS.COM.

Feedback

Understand the relationship

The reasoning behind the result

Window choice changes the observations

overlap starts 1,2,3,…; non-overlap starts 1,3,5,…

Overlapping counting measures every adjacent pair. Non-overlapping counting partitions one declared frame of the record into pairs and can leave one trailing symbol.

A circular overlapping record adds the boundary pair from the last symbol back to the first.

Ambiguity is not fractional evidence

A window containing an IUPAC ambiguity symbol is counted as unresolved. It is not distributed over possible canonical pairs using an unstated probability model.

Only retained canonical windows contribute to either position marginal; unresolved windows and trailing unpaired symbols remain disclosed outside that denominator.

Expected frequency is a declared independence comparison

E(XY) = pfirst(X)psecond(Y)

The expectation multiplies the first-position and second-position symbol shares of the retained canonical pair windows. It is a descriptive independence comparison, not an organism, genome or strand-symmetry reference.

Observed/expected is undefined when either position marginal is zero.

Composition does not establish mechanism

Pair enrichment can depend on region selection, strand convention, sequence length, coding frame and many biological processes.

This workbench does not infer methylation, mutation pressure, codon bias, motifs, statistical significance or causation.

Follow the numbers

Count overlapping pairs

  1. For AACGCGTTAA, advance a two-symbol window one base at a time.
  2. Retain AA, AC, CG, GC, CG, GT, TT, TA and AA: nine possible canonical windows.
  3. Count each of the 16 pair types, including zero rows.
  4. Divide each count by nine for observed frequency.
  5. Multiply each symbol's first-position share by the second symbol's second-position share for the displayed independence expectation.

The ratio describes these retained windows under one explicit marginal-product comparison; it is not a population enrichment test.

Quick guide

How to use this calculator

  1. Declare the polymer, record topology and whether adjacent windows overlap.
  2. Enter one sequence and keep ambiguity symbols when they are part of the record.
  3. Interpret observed/expected only against the retained-window marginal-product comparison, not as a universal biological baseline.

Calculation method

Calculation and interpretation

Keep literal pair counts, window choice, unresolved ambiguity and the chosen independence comparison visible instead of returning one unexplained dinucleotide percentage.

Observed pair frequency = canonical pair count ÷ canonical pair windows. Independence expectation = p(first-position symbol) × p(second-position symbol), using the retained canonical pair windows.

Worked example

Count overlapping pairs

The ratio describes these retained windows under one explicit marginal-product comparison; it is not a population enrichment test.

Observed pair frequency = canonical pair count ÷ canonical pair windows. Independence expectation = p(first-position symbol) × p(second-position symbol), using the retained canonical pair windows.

Supported inputs

Precision and limits

Literal word composition

The workflow counts entered symbols; it does not search motifs, align records or infer a reference genome.

One circular convention

Circular mode is restricted to overlapping adjacent pairs so the last-to-first boundary is unambiguous.

Ambiguity remains unresolved

Windows containing IUPAC ambiguity are disclosed and excluded from canonical pair frequencies.

No significance inference

Observed/expected ratios have no confidence interval, p-value or multiple-testing interpretation here.

Numerical support

One sequence of 2–200,000 nucleotide symbols is supported; all 16 canonical pair rows are retained.

Continue calculating

Related calculators