Princeton University
US
Researchers
Public research profiles associated with Princeton University.
Li Fei-Fei
Kai Li
Li-Jia Li
Richard Socher
Wei Dong
Jia Deng
B Santra
H-Y Ko
R Car
Eric F. Wood
Miroslav Dudı́k
David M. Blei
Abhijit Banerjee
Robert E. Schapire
Evan M. Cofer
Robert M. May
Stanislas Leibler
J. J. Hopfield
John Wheeler
Charles W. Misner
Clifford P. Brangwynne
S. B. Treiman
P. M. Chaikin
Research from this institution
Publications linked through researcher authorship records.
ImageNet: A large-scale hierarchical image database
The explosion of image data on the Internet has the potential to foster more sophisticated and robust models and algorithms to index, retrieve, organize and interact with images and multimedia data. But exactly how such data can be harnessed and organized remains a critical problem. We introduce here a new database called “ImageNet”, a large-scale ontology of images built upon the backbone of the WordNet structure. ImageNet aims to populate the majority of the 80,000 synsets of WordNet with an average of 500–1000 clean and full resolution images. This will result in tens of millions of annotated images organized by the semantic hierarchy of WordNet. This paper offers a detailed analysis of ImageNet in its current state: 12 subtrees with 5247 synsets and 3.2 million images in total. We show that ImageNet is much larger in scale and diversity and much more accurate than the current image datasets. Constructing such a large-scale database is a challenging task. We describe the data collection scheme with Amazon Mechanical Turk. Lastly, we illustrate the usefulness of ImageNet through three simple applications in object recognition, image classification and automatic object clustering. We hope that the scale, accuracy, diversity and hierarchical structure of ImageNet can offer unparalleled opportunities to researchers in the computer vision community and beyond.
ImageNet Large Scale Visual Recognition Challenge
Glove: Global Vectors for Word Representation
Recent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using vector arith-metic, but the origin of these regularities has remained opaque. We analyze and make explicit the model properties needed for such regularities to emerge in word vectors. The result is a new global log-bilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods. Our model efficiently leverages statistical information by training only on the nonzero elements in a word-word co-occurrence matrix, rather than on the en-tire sparse matrix or on individual context windows in a large corpus. The model pro-duces a vector space with meaningful sub-structure, as evidenced by its performance of 75 % on a recent word analogy task. It also outperforms related models on simi-larity tasks and named entity recognition. 1
QUANTUM ESPRESSO: a modular and open-source software project for quantum simulations of materials
QUANTUM ESPRESSO is an integrated suite of computer codes for electronic-structure calculations and materials modeling, based on density-functional theory, plane waves, and pseudopotentials (norm-conserving, ultrasoft, and projector-augmented wave). The acronym ESPRESSO stands for opEn Source Package for Research in Electronic Structure, Simulation, and Optimization. It is freely available to researchers around the world under the terms of the GNU General Public License. QUANTUM ESPRESSO builds upon newly-restructured electronic-structure codes that have been developed and tested by some of the original authors of novel electronic-structure algorithms and applied in the last twenty years by some of the leading materials modeling groups worldwide. Innovation and efficiency are still its main focus, with special attention paid to massively parallel architectures, and a great effort being devoted to user friendliness. QUANTUM ESPRESSO is evolving towards a distribution of independent and interoperable codes in the spirit of an open-source project, where researchers active in the field of electronic-structure calculations are encouraged to participate in the project by contributing their own codes or by implementing their own ideas into existing codes.
Latent dirichlet allocation
We describe latent Dirichlet allocation (LDA), a generative probabilistic model for collections of discrete data such as text corpora. LDA is a three-level hierarchical Bayesian model, in which each item of a collection is modeled as a finite mixture over an underlying set of topics. Each topic is, in turn, modeled as an infinite mixture over an underlying set of topic probabilities. In the context of text modeling, the topic probabilities provide an explicit representation of a document. We present efficient approximate inference techniques based on variational methods and an EM algorithm for empirical Bayes parameter estimation. We report results in document modeling, text classification, and collaborative filtering, comparing to a mixture of unigrams model and the probabilistic LSI model.
Maximum entropy modeling of species geographic distributions
Advanced capabilities for materials modelling with Q <scp>uantum</scp> ESPRESSO
Quantum EXPRESSO is an integrated suite of open-source computer codes for quantum simulations of materials using state-of-the-art electronic-structure techniques, based on density-functional theory, density-functional perturbation theory, and many-body perturbation theory, within the plane-wave pseudopotential and projector-augmented-wave approaches. Quantum EXPRESSO owes its popularity to the wide variety of properties and processes it allows to simulate, to its performance on an increasingly broad array of hardware architectures, and to a community of researchers that rely on its capabilities as a core open-source development platform to implement their ideas. In this paper we describe recent extensions and improvements, covering new methodologies and property calculators, improved parallelization, code modularization, and extended interoperability both within the distribution and with external software.
Simple mathematical models with very complicated dynamics
Modeling of species distributions with Maxent: new extensions and a comprehensive evaluation
Accurate modeling of geographic distributions of species is crucial to various applications in ecology and conservation. The best performing techniques often require some parameter tuning, which may be prohibitively time‐consuming to do separately for each species, or unreliable for small or biased datasets. Additionally, even with the abundance of good quality data, users interested in the application of species models need not have the statistical knowledge required for detailed tuning. In such cases, it is desirable to use “default settings”, tuned and validated on diverse datasets. Maxent is a recently introduced modeling technique, achieving high predictive accuracy and enjoying several additional attractive properties. The performance of Maxent is influenced by a moderate number of parameters. The first contribution of this paper is the empirical tuning of these parameters. Since many datasets lack information about species absence, we present a tuning method that uses presence‐only data. We evaluate our method on independently collected high‐quality presence‐absence data. In addition to tuning, we introduce several concepts that improve the predictive accuracy and running time of Maxent. We introduce “hinge features” that model more complex relationships in the training data; we describe a new logistic output format that gives an estimate of probability of presence; finally we explore “background sampling” strategies that cope with sample selection bias and decrease model‐building time. Our evaluation, based on a diverse dataset of 226 species from 6 regions, shows: 1) default settings tuned on presence‐only data achieve performance which is almost as good as if they had been tuned on the evaluation data itself; 2) hinge features substantially improve model performance; 3) logistic output improves model calibration, so that large differences in output values correspond better to large differences in suitability; 4) “target‐group” background sampling can give much better predictive performance than random background sampling; 5) random background sampling results in a dramatic decrease in running time, with no decrease in model performance.
A Simple Model of Herd Behavior
We analyze a sequential decision model in which each decision maker looks at the decisions made by previous decision makers in taking her own decision. This is rational for her because these other decision makers may have some information that is important for her. We then show that the decision rules that are chosen by optimizing individuals will be characterized by herd behavior; i.e., people will be doing what others are doing rather than using their information. We then show that the resulting equilibrium is inefficient.
Probabilistic topic models
Surveying a suite of algorithms that offer a solution to managing large document archives.
A simple hydrologically based model of land surface water and energy fluxes for general circulation models
A generalization of the single soil layer variable infiltration capacity (VIC) land surface hydrological model previously implemented in the Geophysical Fluid Dynamics Laboratory general circulation model (GCM) is described. The new model is comprised of a two‐layer characterization of the soil column, and uses an aerodynamic representation of the latent and sensible heat fluxes at the land surface. The infiltration algorithm for the upper layer is essentially the same as for the single layer VIC model, while the lower layer drainage formulation is of the form previously implemented in the Max‐Planck‐Institut GCM. The model partitions the area of interest (e.g., grid cell) into multiple land surface cover types; for each land cover type the fraction of roots in the upper and lower zone is specified. Evapotranspiration consists of three components: canopy evaporation, evaporation from bare soils, and transpiration, which is represented using a canopy and architectural resistance formulation. Once the latent heat flux has been computed, the surface energy balance is iterated to solve for the land surface temperature at each time step. The model was tested using long‐term hydrologic and climatological data for Kings Creek, Kansas to estimate and validate the hydrological parameters, and surface flux data from three First International Satellite Land Surface Climatology Project Field Experiment intensive field campaigns in the summer‐fall of 1987 to validate the surface energy fluxes.
From molecular to modular cell biology
Principles of Condensed Matter Physics
Now in paperback, this book provides an overview of the physics of condensed matter systems. Assuming a familiarity with the basics of quantum mechanics and statistical mechanics, the book establishes a general framework for describing condensed phases of matter, based on symmetries and conservation laws. It explores the role of spatial dimensionality and microscopic interactions in determining the nature of phase transitions, as well as discussing the structure and properties of materials with different symmetries. Particular attention is given to critical phenomena and renormalization group methods. The properties of liquids, liquid crystals, quasicrystals, crystalline solids, magnetically ordered systems and amorphous solids are investigated in terms of their symmetry, generalised rigidity, hydrodynamics and topological defect structure. In addition to serving as a course text, this book is an essential reference for students and researchers in physics, applied physics, chemistry, materials science and engineering, who are interested in modern condensed matter physics.
Population biology of infectious diseases: Part I
<i>The Feynman Lectures on Physics</i>
Share Icon Share Twitter Facebook Reddit LinkedIn Reprints and Permissions Cite Icon Cite Search Site Citation Richard P. Feynman, Robert B. Leighton, Matthew Sands, S. B. Treiman; The Feynman Lectures on Physics. Physics Today 1 August 1964; 17 (8): 45–46. https://doi.org/10.1063/1.3051743 Download citation file: Ris (Zotero) Reference Manager EasyBib Bookends Mendeley Papers EndNote RefWorks BibTex toolbar search Search Dropdown Menu toolbar search search input Search input auto suggest filter your search All ContentPhysics Today Search Advanced Search
<i>Principles of Condensed Matter Physics</i>
Share Icon Share Twitter Facebook Reddit LinkedIn Reprints and Permissions Cite Icon Cite Search Site Citation Paul M. Chaikin, Thomas C. Lubensky, Thomas A. Witten; Principles of Condensed Matter Physics. Physics Today 1 November 1995; 48 (11): 82. https://doi.org/10.1063/1.2808258 Download citation file: Ris (Zotero) Reference Manager EasyBib Bookends Mendeley Papers EndNote RefWorks BibTex toolbar search Search Dropdown Menu toolbar search search input Search input auto suggest filter your search All ContentPhysics Today Search Advanced Search
Opportunities and obstacles for deep learning in biology and medicine
Deep learning describes a class of machine learning algorithms that are capable of combining raw inputs into layers of intermediate features. These algorithms have recently shown impressive results across a variety of domains. Biology and medicine are data-rich disciplines, but the data are complex and often ill-understood. Hence, deep learning techniques may be particularly well suited to solve problems of these fields. We examine applications of deep learning to a variety of biomedical problems-patient classification, fundamental biological processes and treatment of patients-and discuss whether deep learning will be able to transform these tasks or if the biomedical sphere poses unique challenges. Following from an extensive literature review, we find that deep learning has yet to revolutionize biomedicine or definitively resolve any of the most pressing challenges in the field, but promising advances have been made on the prior state of the art. Even though improvements over previous baselines have been modest in general, the recent progress indicates that deep learning methods will provide valuable means for speeding up or aiding human investigation. Though progress has been made linking a specific neural network's prediction to input features, understanding how users should interpret these models to make testable hypotheses about the system under study remains an open challenge. Furthermore, the limited amount of labelled data for training presents problems in some domains, as do legal and privacy constraints on work with sensitive health records. Nonetheless, we foresee deep learning enabling changes at both bench and bedside with the potential to transform several areas of biology and medicine.