Researcher profile

Vincent Vanhoucke

· Google (United States)

0Publications
0KT Citations
0KT h-index
0KT i10-index

KT metrics are calculated only from papers uploaded or published on KnowledgeTrend and citations matched between those KnowledgeTrend papers. Imported metadata and external citation counts are excluded.

Research interests

Research interests have not yet been added.

Academic profiles & contact

Publications

3 research records shown

Rethinking the Inception Architecture for Computer Vision
2016 · DOI 10.1109/cvpr.2016.308

Convolutional networks are at the core of most state of-the-art computer vision solutions for a wide variety of tasks. Since 2014 very deep convolutional networks started to become mainstream, yielding substantial gains in various benchmarks. Although increased model size and computational cost tend to translate to immediate quality gains for most tasks (as long as enough labeled data is provided for training), computational efficiency and low parameter count are still enabling factors for various use cases such as mobile vision and big-data scenarios. Here we are exploring ways to scale up networks in ways that aim at utilizing the added computation as efficiently as possible by suitably factorized convolutions and aggressive regularization. We benchmark our methods on the ILSVRC 2012 classification challenge validation set demonstrate substantial gains over the state of the art: 21:2% top-1 and 5:6% top-5 error for single frame evaluation using a network with a computational cost of 5 billion multiply-adds per inference and with using less than 25 million parameters. With an ensemble of 4 models and multi-crop evaluation, we report 3:5% top-5 error and 17:3% top-1 error on the validation set and 3:6% top-5 error on the official test set.

Read paper
Going deeper with convolutions
2015 · DOI 10.1109/cvpr.2015.7298594

We propose a deep convolutional neural network architecture codenamed Inception that achieves the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC14). The main hallmark of this architecture is the improved utilization of the computing resources inside the network. By a carefully crafted design, we increased the depth and width of the network while keeping the computational budget constant. To optimize quality, the architectural decisions were based on the Hebbian principle and the intuition of multi-scale processing. One particular incarnation used in our submission for ILSVRC14 is called GoogLeNet, a 22 layers deep network, the quality of which is assessed in the context of classification and detection.

Read paper
Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
2012 · IEEE Signal Processing Magazine · DOI 10.1109/msp.2012.2205597

Most current speech recognition systems use hidden Markov models (HMMs) to deal with the temporal variability of speech and Gaussian mixture models (GMMs) to determine how well each state of each HMM fits a frame or a short window of frames of coefficients that represents the acoustic input. An alternative way to evaluate the fit is to use a feed-forward neural network that takes several frames of coefficients as input and produces posterior probabilities over HMM states as output. Deep neural networks (DNNs) that have many hidden layers and are trained using new methods have been shown to outperform GMMs on a variety of speech recognition benchmarks, sometimes by a large margin. This article provides an overview of this progress and represents the shared views of four research groups that have had recent successes in using DNNs for acoustic modeling in speech recognition.

Read paper

Co-authors

Christian Szegedy

Google (United States)

2 shared publications
Geoffrey E. Hinton

University of Toronto

1 shared publication
Li Deng

University of Waterloo

1 shared publication
Dong Yu

Microsoft (United States)

1 shared publication
George E. Dahl

University of Toronto

1 shared publication
Abdelrahman Mohamed

University of Toronto

1 shared publication
Navdeep Jaitly

University of Toronto

1 shared publication
Andrew Senior

Google (United States)

1 shared publication
Patrick Nguyen

Google (United States)

1 shared publication
Tara N. Sainath

IBM Research - Thomas J. Watson Research Center

1 shared publication
Brian Kingsbury

Michigan State University

1 shared publication
Wei Liu

University of North Carolina at Chapel Hill

1 shared publication