-

- -

Showing posts with label deep-learning. Show all posts
Showing posts with label deep-learning. Show all posts

Thursday, January 30, 2020

[ebook] Deep Learning with JavaScript Neural networks in TensorFlow.js



.

Deep learning has transformed the fields of computer vision, image processing, and natural language applications. Thanks to TensorFlow.js, now JavaScript developers can build deep learning apps without relying on Python or R. Deep Learning with JavaScript shows developers how they can bring DL technology to the web. Written by the main authors of the TensorFlow library, this new book provides fascinating use cases and in-depth instruction for deep learning apps in JavaScript in your browser or on Node.

about the technology

Running deep learning applications in the browser or on Node-based backends opens up exciting possibilities for smart web applications. With the TensorFlow.js library, you build and train deep learning models with JavaScript. Offering uncompromising production-quality scalability, modularity, and responsiveness, TensorFlow.js really shines for its portability. Its models run anywhere JavaScript runs, pushing ML farther up the application stack.

about the book

In Deep Learning with JavaScript, you’ll learn to use TensorFlow.js to build deep learning models that run directly in the browser. This fast-paced book, written by Google engineers, is practical, engaging, and easy to follow. Through diverse examples featuring text analysis, speech processing, image recognition, and self-learning game AI, you’ll master all the basics of deep learning and explore advanced concepts, like retraining existing models for transfer learning and image generation.

what's inside

  • Image and language processing in the browser
  • Tuning ML models with client-side data
  • Text and image creation with generative deep learning
  • Source code samples to test and modify

about the reader

For JavaScript programmers interested in deep learning.

about the author

Shanging CaiStanley Bileschi and Eric D. Nielsen are software engineers with experience on the Google Brain team, and were crucial to the development of the high-level API of TensorFlow.js. This book is based in part on the classic, Deep Learning with Python by François Chollet.

.

https://www.manning.com/books/deep-learning-with-javascript 

Saturday, January 4, 2020

DeepLearning With TensorFlowJS 4 - The intuitions behind Gradient-Descent Optimization

One-layer model is fitting a linear function f(input), defined as output = kernel * input + bias

The kernel and bias are tunable parameters (the weights) of the dense layer.

These weights contain the information learned by the network from exposure to the training data.

Initially, these weights are filled with small random values (a step called random initialization).

To find a good setting for the kernel and bias (collectively, the weights) we need two things:

  • A measure that tells us how well we are doing at a given setting of the weights. This is represented by a loss function measurement. 
  • A method to update the weights’ values so that next time we will do better than we currently are doing, according to the measure previously mentioned. This is accomplished by an optimizer method i.e. the algorithm by which the network will update its weights (kernel and bias, in this case) based on the data and the loss function.
  • The compile() method specifies 'sgd' as the optimizer and 'meanAbsoluteError' as the loss.

    'meanAbsoluteError' means that the loss function will calculate how far the predictions are from the targets, take their absolute values (making them all positive), and then return the average of those values:

    meanAbsoluteError = average( absolute(modelOutput - targets))

    'sgd' stands for stochastic gradient descent, a calculus formula to determine what adjustments should be made to the weights in order to reduce the loss.

    The fit() method is the training process of a model in TensorFlow.js. It can often be long-running, lasting for seconds or minutes. Therefore, the async/await feature is used.

    The evaluate() method calculates the loss function as applied to the provided example features and targets. It is similar to the fit() method in that it calculates the same loss, but evaluate() does not update the model’s weights.

    The training loop iterates through the following steps:

    1. Draw a batch of training samples x and corresponding targets y_true. A batch is simply a number of input examples put together as a tensor. The number of examples in a batch is called the batch size. In practical deep learning, it is often set to be a power of 2, such as 128 or 256. Examples are batched together to take advantage of the GPU’s parallel processing power and to make the calculated values of the gradients more stable.

    2. Run the network on x (a step called the forward pass) to obtain predictions y_pred.

    3. Compute the loss of the network on the batch, a measure of the mismatch between y_true and y_pred. Recall that the loss function is specified when model.compile() is called.

    4. Update all the weights (parameters) in the network in a way that slightly reduces the loss on this batch. The detailed updates to the individual weights are managed by the optimizer, which was specified during the model.compile() call.

    The loss as a function of all tunable parameters is known as the loss surface concept.

    The loss surface for this example has a bowl shape, with a global minimum at the bottom of the bowl representing the best parameter settings. 

    In general, however, the loss surface of a deep-learning model is much more complex. It will have many more than two dimensions and could have many local minima i.e. points that are lower than anything nearby but not the lowest overall.



    For larger problems i.e. when optimizing millions of weights, the likelihood of randomly selecting a good direction becomes vanishingly small. 

    A much better approach is to take advantage of the fact that all operations used in the network are differentiable and hence, to compute the gradient of the loss with regard to the network’s parameters. 

    The mathematical definition of a gradient specifies a direction along which the loss function increases. When training neural networks, the loss should gradually decrease. Therefore the weights should be moved in the direction opposite the gradient. This training process is aptly named gradient descent.

    One of the most desirable properties of deep neural networks are that they are universal approximators. Which means they should be able to cover non-convex functions as well. The problem with non-convex functions is that your initial guess might not be near the global minima and gradient descent might converge to a local minima. A solution to this problem is the stochastic gradient descent  approach.

    The term “stochastic” means drawing random samples from the training data during each gradient-descent step for efficiency, as opposed to using every training data sample at every step. In short, stochastic gradient descent is simply a modification of gradient descent for computational efficiency.

    Stochastic means nondeterministic or unpredictable. Random generally means unrecognizable, not adhering to a pattern. A random variable is also called a stochastic variable. (https://math.stackexchange.com/questions/114373/whats-the-difference-between-stochastic-and-random)


    .

    Friday, January 3, 2020

    DeepLearning With TensorFlowJS 3 - Fitting The Model


     

    This tutorial is based on the book Deep Learning With JavaScript (TensorFlowJS).



    https://codepen.io/tfjs-book/pen/VEVMMd

    Thursday, January 2, 2020

    DeepLearning With TensorFlowJS 2 - Plotting Tensor Data

    This tutorial is based on the book Deep Learning With JavaScript (TensorFlowJS).


    Tensors

    Tensors are the core data structure of TensorFlow.js 

    Tensors can also be thought of as containers for numbers.

    They are a generalization of vectors and matrices to potentially higher dimensions. 

    The number of dimensions and size of each dimension is called the tensor’s shape.
     
    Declaring a tensor

    // Pass an array of values to create a vector.
    tf.tensor([1, 2, 3, 4]).print();

    // Pass a nested array of values to make a matrix or a higher
    // dimensional tensor.
    tf.tensor([[1, 2], [3, 4]]).print();

    //Creates rank-1 tf.Tensor with the provided values, shape and dtype.
    tf.tensor1d([1, 2, 3]).print();

    //Creates rank-2 tf.Tensor with the provided values, shape and dtype.
    // Pass a nested array.
    tf.tensor2d([[1, 2], [3, 4]]).print();




    Plotly.js is a charting library that comes with over 40 chart types, 3D charts, statistical graphs, and SVG maps.





    https://codepen.io/tfjs-book/pen/dgQVze

    Tuesday, December 31, 2019

    DeepLearning With TensorFlowJS 1 - Train Data and Test Data

    This tutorial is based on the book Deep Learning With JavaScript (TensorFlowJS).

    The first script loads the TensorFlow package and defines the symbol tf, which provides a way to refer to names in TensorFlow.

    The second script creates two constants, trainData and testData, each representing 20 samples of how long it took to download a file (timeSec) and the size of that file (sizeMB). The elements in sizeMB and those in timeSec have one-to-one correspondence. For example, the first element of sizeMB in trainData is 0.080 MB, and downloading that file took 0.135 seconds—that is, the first element of timeSec—and so forth.

    The goal in this example will be to estimate timeSec, given just sizeMB.

    https://codepen.io/tfjs-book/pen/VEVMbx

    Wednesday, December 25, 2019

    [ebook] Neural Networks and Deep Learning free online book.

    .

    CHAPTER 1: Using neural nets to recognize handwritten digits

    In this chapter we'll write a computer program implementing a neural network that learns to recognize handwritten digits. 

    .

    CHAPTER 2: How the backpropagation algorithm works

    In this chapter I'll explain a fast algorithm for computing such gradients, an algorithm known as backpropagation.

    .

    CHAPTER 3:Improving the way neural networks learn

    In this chapter I explain a suite of techniques which can be used to improve on our vanilla implementation of backpropagation, and so improve the way our networks learn.

    .

    CHAPTER 4:A visual proof that neural nets can compute any function

    In this chapter I give a simple and mostly visual explanation of the universality theorem. 

    .

    CHAPTER 5:Why are deep neural networks hard to train?

    In this chapter, we'll try training deep networks using our workhorse learning algorithm - stochastic gradient descent by backpropagation.

    .

    CHAPTER 6:Deep learning

    In this chapter, we'll develop techniques which can be used to train deep networks, and apply them in practice.

    .

    Wednesday, August 8, 2018

    ML5JS vs KERAS



     .

    .

    .

    https://towardsdatascience.com/introduction-to-ml5-js-3fe51d6a4661

    .

    Thursday, September 18, 2014

    Neural Network Optimization Algorithm Visualisation



    .
    Optimization is a mathematical discipline that determines the “best” solution in a quantitatively well-defined sense. Mathematical optimization of the processes governed by partial differential equations has seen considerable progress in the past decade, and since then it has been applied to a wide variety of disciplines e.g., science, engineering, mathematics, economics, and even commerce. Optimization theory provides algorithms to solve well-structured optimization problems along with the analysis of those algorithms. A typical optimization problem includes an objective function that is to be minimized or maximized with the given constraints. Optimization theory provides algorithms to solve well-structured optimization problems along with the analysis of those algorithms. Optimization algorithms in machine learning (especially in neural networks) aim at minimizing an objective function (generally called loss or cost function), which is intuitively the difference between the predicted data and the expected values

    Stochastic gradient-based optimization is of core practical importance in many fields of science and engineering. Many problems in these fields can be cast as the optimization of some scalar parameterized objective function requiring maximization or minimization with respect to its parameters. Gradient descent is an optimization algorithm that uses the gradient of the objective function to navigate the search space. Several optimization algorithms based on gradient descent exist in the literature, but just to name a few the classification of Gradient descent optimization algorithms goes as follows ...

    (futher reading: https://medium.com/analytics-vidhya/a-complete-guide-to-adam-and-rmsprop-optimizer-75f4502d83be Feb 2021)



    .
    Visualizing Optimization Algorithms (algos)


    Algos without scaling based on gradient information really struggle to break symmetry here - SGD gets no where and Nesterov Accelerated Gradient (NAG) / Momentum exhibits oscillations until they build up velocity in the optimization direction.

    Algos that scale step size based on the gradient quickly break symmetry and begin descent.




    Due to the large initial gradient, velocity based techniques shoot off and bounce around - adagrad almost goes unstable for the same reason.

    Algos that scale gradients/step sizes like adadelta and RMSProp proceed more like accelerated SGD and handle large gradients with more stability.


    Behavior around a saddle point.

    NAG/Momentum again like to explore around, almost taking a different path. 

    Adadelta/Adagrad/RMSProp proceed like accelerated SGD.
    ..
    Reference: