
Chapter 6, Deep Feedforward Networks, separates what a network can represent from how efficiently it can represent it and whether optimization can find the parameters.
Existence Does Not Imply Learnability
The common shorthand for the Universal Approximation Theorem is that neural networks can approximate any function. Chapter 6 makes the important constraint explicit: this is an existence result about representational capacity, not a statement about efficiency or trainability.
A single hidden layer network can approximate complex structured functions, but the width required may scale exponentially for certain compositional forms. Depth can reduce parameter count by reusing intermediate computations. In that sense, depth changes scaling behavior, not just capacity.
A linear bottleneck imposes a rank constraint
When we stack linear layers without nonlinearities between them, two consecutive linear transformations collapse into a single linear transformation. Functionally, nothing changes, but the parameterization does.
If a weight matrix W∈Rm×n is factored as W=AB with A∈Rm×r and B∈Rr×n, we have expressed the same linear map with a rank constraint and potentially far fewer parameters when r≪min(m,n). In other words, a linear path with a narrow intermediate layer induces a low-rank factorization.
The same structure shows up in LoRA-style adaptation for large language models. Inserting a bottleneck linear path imposes a low-rank constraint on the effective weight update. The connection is basic linear algebra: the architecture changes the parameterization of the update.
Softplus vs ReLU
The comparison between softplus and ReLU corrected an intuition I had inherited from smooth optimization: smoother functions should be easier to train. Softplus is differentiable everywhere with nonzero gradient, while ReLU is nondifferentiable at zero and flat for negative inputs. By a classical smoothness criterion, softplus seems preferable.
Chapter 6 cites experiments in which ReLU outperformed softplus. It also explains why ReLU's linear behavior on positive inputs can help optimization: its derivative is constant there, so the activation does not shrink the gradient as its input grows. On negative inputs, it outputs zero, effectively selecting a subnetwork for each input. These properties matter alongside smoothness; being differentiable everywhere does not by itself make an activation easier to train.
Get new posts by email
Occasional updates when I publish something new.
Follow.it sends the emails and includes an unsubscribe link.
Related posts
Dropout as Shared-Parameter Bagging
March 7, 2026
A concrete way to understand dropout in deep learning: sampled subnetworks with shared weights, standing in for a much more expensive ensemble.
Math Foundations for Deep Learning
February 7, 2026
Part I of Deep Learning reviews the linear algebra, probability, and numerical computation that later chapters build on.
Camping Indoors in San Francisco
April 27, 2026
My first week in San Francisco: paperwork, an unfurnished apartment, two cats, Bay Trail runs, and waiting for furniture.