Skip to content

Information Theory and the Shannon–Hartley Law

Information theory quantifies the uncertainty removed when a message is received. If an event has probability PP, its self-information is

Thus a less probable event conveys more information. Independent events add information because −log⁡2(P1P2)=−log⁡2P1−log⁡2P2-\log_2(P_1P_2)=-\log_2P_1-\log_2P_2.

For a discrete source with symbols xix_i and probabilities pip_i, the average information per source symbol is its entropy:

Entropy is maximized when all source symbols are equally likely. It gives the minimum achievable average number of bits per symbol for lossless source coding, approached by sufficiently long codes.

Channel capacity is the supremum of information rates that can be transmitted reliably through a channel. Two limits must be distinguished.

A channel band-limited to BB Hz can convey at most 2B2B independent pulses per second without intersymbol interference. This is the Nyquist signalling rate,

Rs=2Bsymbols/s.R_s=2B\quad\text{symbols/s}.

If each pulse can assume one of MM equally likely distinguishable levels, one symbol carries log⁡2M\log_2M bits. Therefore

C=Rslog⁡2M=2Blog⁡2M.C=R_s\log_2M=\boxed{2B\log_2M}.

In a mathematically noiseless channel there is no bound on MM, so levels could be packed arbitrarily closely. Noise removes that freedom and leads to the Shannon–Hartley limit.

Model one independent channel sample as

Y=X+N,Y=X+N,

where XX is the transmitted sample and NN is independent additive white Gaussian noise. Mutual information per sample is

I(X;Y)=h(Y)−h(Y∣X)=h(Y)−h(N),I(X;Y)=h(Y)-h(Y\mid X)=h(Y)-h(N),

because knowing XX leaves only the uncertainty due to NN. Among all random variables of a fixed variance, the Gaussian has maximum differential entropy. For signal power SS and noise power NN,

h(N)=12log⁡2(2πeN),h(Y)≤12log⁡2 ⁣(2πe(S+N)),h(N)=\frac12\log_2(2\pi eN),\qquad h(Y)\leq\frac12\log_2\!\bigl(2\pi e(S+N)\bigr),

with equality for a Gaussian input. Hence the greatest information per independent real sample is

Csample=12log⁡2 ⁣(S+NN)=12log⁡2 ⁣(1+SN)bits/sample.\begin{aligned} C_{\mathrm{sample}} &=\frac12\log_2\!\left(\frac{S+N}{N}\right)\\ &=\frac12\log_2\!\left(1+\frac{S}{N}\right) \quad\text{bits/sample}. \end{aligned}

A channel of bandwidth BB supplies 2B2B independent real samples per second, so

C=2B Csample=Blog⁡2 ⁣(1+SN).C=2B\,C_{\mathrm{sample}} =\boxed{B\log_2\!\left(1+\frac{S}{N}\right)}.

This is a proof sketch: a rigorous argument decomposes the band-limited Gaussian channel into independent orthogonal modes and invokes the channel coding theorem. For every rate R<CR<C, sufficiently long suitable codes can make error probability arbitrarily small; finite codes are not promised zero error. Rates above CC cannot be made reliable.

Shannon–Hartley capacity rises logarithmically with the linear signal-to-noise ratio; every additional 3 dB does not double capacity.

Shannon–Hartley capacity rises logarithmically with the linear signal-to-noise ratio; every additional 3 dB does not double capacity.

Interpretation and Bandwidth–SNR Trade-off

Section titled “Interpretation and Bandwidth–SNR Trade-off”
  • Fixed numerical S/NS/N: Capacity increases linearly with bandwidth BB.

  • Fixed bandwidth: Capacity increases only logarithmically with S/NS/N.

  • Coding significance: Capacity is a theoretical reliability limit; practical modulation and coding schemes seek to approach it.

When signal power SS and the two-sided white-noise description are represented by an equivalent in-band noise density N0N_0 such that N=N0BN=N_0B, increasing bandwidth also admits more noise. Then

C=Blog⁡2 ⁣(1+SN0B),C=B\log_2\!\left(1+\frac{S}{N_0B}\right),

which still increases with BB, but with diminishing returns rather than linearly.