Formalisms:
Given data X={x(1),…,x(n)} corresponding to labels y={y(1),…,y(n)}, find h:Rd→R, where h∈H (i.e. the set of all hypothesis) performs best according with the evaluation criteria.
h(x)=j=1∑dθjxj+θ0x0=θ⊤x, where x0 = 1
And then we can define the Cost Function as:
J(θ)=2n1i=1∑n(hθ(x(i))−y(i))2
where we minimize J(θ) to find the model parameters θ.
Gradient Descent
- initialize θ
- Repeat until convergence
θj←θj−α∂θj∂J(θ)
For linear regression:
∂θj∂J(θ)=∂θj∂2n1i=1∑n(hθ(x(i))−y(i))2=∂θj∂2n1i=1∑n(hθ(x(i))−y(i))2→consider it as (u2)′=2×u×u′=n1i=1∑n(hθ(x(i))−y(i))×∂θj∂(hθ(x(i))−y(i))=n1i=1∑n(k=0∑dhθ(x(i))−y(i))×∂θj∂(k=0∑dθkxk(i)−y(i))=n1i=1∑n(k=0∑dhθ(x(i))−y(i))(xj(i))
- Assume convergence when ∣∣θnew−θold∣∣2<ϵ
page 53.