Good old ReLU. I’ve heard that CNN’s perform better when using ReLU as the activation func, and that’s probably why; ReLU acts as a filter on the image’s features.
Nah. Without a nonlinearity you just get a linear combination of inputs instead of the output of a deep neural network. ReLU or Ramp is just the simplest possible non linearity. Using a simple function can enable using deeper networks yielding even better performance.
It’s actually somewhat of a headache, numerically. Works well enough tho.
Max(input*weights, 0) is an if in a sense, I guess.
Good old ReLU. I’ve heard that CNN’s perform better when using ReLU as the activation func, and that’s probably why; ReLU acts as a filter on the image’s features.
Nah. Without a nonlinearity you just get a linear combination of inputs instead of the output of a deep neural network. ReLU or Ramp is just the simplest possible non linearity. Using a simple function can enable using deeper networks yielding even better performance.
It’s actually somewhat of a headache, numerically. Works well enough tho.