University of California, Los Angeles
US
Researchers
Public research profiles associated with University of California, Los Angeles.
Peter Langfelder
Lieven Vandenberghe
Wayne W. Grody
Ashraf S. Ibrahim
Maria Punchak
Volker Hartenstein
Simon Potter
James C. McWilliams
Alexander F. Shchepetkin
Bengt Muthén
Karen Nylund‐Gibson
Chih-Ping Chou
Alan L. Carsrud
Kenneth Lange
John Novembre
David H. Alexander
Peter M. Bentler
Marc A. Suchard
Nelson B. Freimer
Minsoo Kim
Lori L. Bonnycastle
Harry V. Vinters
Michael V. Sofroniew
Francis F. Chen
Andrew J. Larkoski
Thomas B. Smith
Nelson B. Freimer
Minsoo Kim
Michael J. Gandal
Research from this institution
Publications linked through researcher authorship records.
Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives
This article examines the adequacy of the “rules of thumb” conventional cutoff criteria and several new alternatives for various fit indexes used to evaluate model fit in practice. Using a 2‐index presentation strategy, which includes using the maximum likelihood (ML)‐based standardized root mean squared residual (SRMR) and supplementing it with either Tucker‐Lewis Index (TLI), Bollen's (1989) Fit Index (BL89), Relative Noncentrality Index (RNI), Comparative Fit Index (CFI), Gamma Hat, McDonald's Centrality Index (Mc), or root mean squared error of approximation (RMSEA), various combinations of cutoff values from selected ranges of cutoff criteria for the ML‐based SRMR and a given supplemental fit index were used to calculate rejection rates for various types of true‐population and misspecified models; that is, models with misspecified factor covariance(s) and models with misspecified factor loading(s). The results suggest that, for the ML method, a cutoff value close to .95 for TLI, BL89, CFI, RNI, and Gamma Hat; a cutoff value close to .90 for Mc; a cutoff value close to .08 for SRMR; and a cutoff value close to .06 for RMSEA are needed before we can conclude that there is a relatively good fit between the hypothesized model and the observed data. Furthermore, the 2‐index presentation strategy is required to reject reasonable proportions of various types of true‐population and misspecified models. Finally, using the proposed cutoff criteria, the ML‐based TLI, Mc, and RMSEA tend to overreject true‐population models at small sample size and thus are less preferable when sample size is small.
Fiji: an open-source platform for biological-image analysis
Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology
Convex Optimization
Convex optimization problems arise frequently in many different fields. This book provides a comprehensive introduction to the subject, and shows in detail how such problems can be solved numerically with great efficiency. The book begins with the basic elements of convex sets and functions, and then describes various classes of convex optimization problems. Duality and approximation techniques are then covered, as are statistical estimation techniques. Various geometrical problems are then presented, and there is detailed discussion of unconstrained and constrained minimization problems, and interior-point methods. The focus of the book is on recognizing convex optimization problems and then finding the most appropriate technique for solving them. It contains many worked examples and homework exercises and will appeal to students, researchers and practitioners in fields such as engineering, computer science, mathematics, statistics, finance and economics.
WGCNA: an R package for weighted correlation network analysis
BACKGROUND: Correlation networks are increasingly being used in bioinformatics applications. For example, weighted gene co-expression network analysis is a systems biology method for describing the correlation patterns among genes across microarray samples. Weighted correlation network analysis (WGCNA) can be used for finding clusters (modules) of highly correlated genes, for summarizing such clusters using the module eigengene or an intramodular hub gene, for relating modules to one another and to external sample traits (using eigengene network methodology), and for calculating module membership measures. Correlation networks facilitate network based gene screening methods that can be used to identify candidate biomarkers or therapeutic targets. These methods have been successfully applied in various biological contexts, e.g. cancer, mouse genetics, yeast genetics, and analysis of brain imaging data. While parts of the correlation network methodology have been described in separate publications, there is a need to provide a user-friendly, comprehensive, and consistent software implementation and an accompanying tutorial. RESULTS: The WGCNA R software package is a comprehensive collection of R functions for performing various aspects of weighted correlation network analysis. The package includes functions for network construction, module detection, gene selection, calculations of topological properties, data simulation, visualization, and interfacing with external software. Along with the R package we also present R software tutorials. While the methods development was motivated by gene expression data, the underlying data mining approach can be applied to a variety of different settings. CONCLUSION: The WGCNA package provides R functions for weighted correlation network analysis, e.g. co-expression network analysis of gene expression data. The R package along with its source code and additional material are freely available at http://www.genetics.ucla.edu/labs/horvath/CoexpressionNetwork/Rpackages/WGCNA.
MrBayes 3.2: Efficient Bayesian Phylogenetic Inference and Model Choice Across a Large Model Space
Since its introduction in 2001, MrBayes has grown in popularity as a software package for Bayesian phylogenetic inference using Markov chain Monte Carlo (MCMC) methods. With this note, we announce the release of version 3.2, a major upgrade to the latest official release presented in 2003. The new version provides convergence diagnostics and allows multiple analyses to be run in parallel with convergence progress monitored on the fly. The introduction of new proposals and automatic optimization of tuning parameters has improved convergence for many problems. The new version also sports significantly faster likelihood calculations through streaming single-instruction-multiple-data extensions (SSE) and support of the BEAGLE library, allowing likelihood calculations to be delegated to graphics processing units (GPUs) on compatible hardware. Speedup factors range from around 2 with SSE code to more than 50 with BEAGLE for codon problems. Checkpointing across all models allows long runs to be completed even when an analysis is prematurely terminated. New models include relaxed clocks, dating, model averaging across time-reversible substitution models, and support for hard, negative, and partial (backbone) tree constraints. Inference of species trees from gene trees is supported by full incorporation of the Bayesian estimation of species trees (BEST) algorithms. Marginal model likelihoods for Bayes factor tests can be estimated accurately across the entire model space using the stepping stone method. The new version provides more output options than previously, including samples of ancestral states, site rates, site d(N)/d(S) rations, branch rates, and node dates. A wide range of statistics on tree parameters can also be output for visualization in FigTree and compatible software.
Comparative fit indexes in structural models.
Normed and nonnormed fit indexes are frequently used as adjuncts to chi-square statistics for evaluating the fit of a structural model. A drawback of existing indexes is that they estimate no known population parameters. A new coefficient is proposed to summarize the relative reduction in the noncentrality parameters of two nested models. Two estimators of the coefficient yield new normed (CFI) and nonnormed (FI) fit indexes. CFI avoids the underestimation of fit often noted in small samples for Bentler and Bonett's (1980) normed fit index (NFI). FI is a linear function of Bentler and Bonett's non-normed fit index (NNFI) that avoids the extreme underestimation and overestimation often found in NNFI. Asymptotically, CFI, FI, NFI, and a new index developed by Bollen are equivalent measures of comparative fit, whereas NNFI measures relative fit by comparing noncentrality per degree of freedom. All of the indexes are generalized to permit use of Wald and Lagrange multiplier statistics. An example illustrates the behavior of these indexes under conditions of correct specification and misspecification. The new fit indexes perform very well at all sample sizes.
Fit indices in covariance structure modeling: Sensitivity to underparameterized model misspecification.
This study evaluated the sensitivity of maximum likelihood (ML)-, generalized least squares (GLS)-, and asymptotic distribution-free (ADF)-based fit indices to model misspecification, under conditions that varied sample size and distribution. The effect of violating assumptions of asymptotic robustness theory also was ex-amined. Standardized root-mean-square residual (SRMR) was the most sensitive index to models with misspecified factor covariance(s), and Tucker-Lewis Index (1973; TLI), Bollen's fit index (1989; BL89), relative noncentrality index (RNI), comparative fit index (CFI), and the ML- and GLS-based gamma hat, McDonald's centrality index (1989; Me), and root-mean-square error of approximation (RMSEA) were the most sensitive indices to models with misspecified factor loadings. With ML and GLS methods, we recommend the use of SRMR, supple-mented by TLI, BL89, RNI, CFI, gamma hat, Me, or RMSEA (TLI, Me, and RMSEA are less preferable at small sample sizes). With the ADF method, we recommend the use of SRMR, supplemented by TLI, BL89, RNI, or CFI. Finally, most of the ML-based fit indices outperformed those obtained from GLS and ADF
Deciding on the Number of Classes in Latent Class Analysis and Growth Mixture Modeling: A Monte Carlo Simulation Study
Mixture modeling is a widely applied data analysis technique used to identify unobserved heterogeneity in a population. Despite mixture models' usefulness in practice, one unresolved issue in the application of mixture models is that there is not one commonly accepted statistical indicator for deciding on the number of classes in a study population. This article presents the results of a simulation study that examines the performance of likelihood-based tests and the traditionally used Information Criterion (ICs) used for determining the number of classes in mixture modeling. We look at the performance of these tests and indexes for 3 types of mixture models: latent class analysis (LCA), a factor mixture model (FMA), and a growth mixture models (GMM). We evaluate the ability of the tests and indexes to correctly identify the number of classes at three different sample sizes (n = 200, 500, 1,000). Whereas the Bayesian Information Criterion performed the best of the ICs, the bootstrap likelihood ratio test proved to be a very consistent indicator of classes across all of the models considered.
Fast model-based estimation of ancestry in unrelated individuals
Population stratification has long been recognized as a confounding factor in genetic association studies. Estimated ancestries, derived from multi-locus genotype data, can be used to perform a statistical correction for population stratification. One popular technique for estimation of ancestry is the model-based approach embodied by the widely applied program structure. Another approach, implemented in the program EIGENSTRAT, relies on Principal Component Analysis rather than model-based estimation and does not directly deliver admixture fractions. EIGENSTRAT has gained in popularity in part owing to its remarkable speed in comparison to structure. We present a new algorithm and a program, ADMIXTURE, for model-based estimation of ancestry in unrelated individuals. ADMIXTURE adopts the likelihood model embedded in structure. However, ADMIXTURE runs considerably faster, solving problems in minutes that take structure hours. In many of our experiments, we have found that ADMIXTURE is almost as fast as EIGENSTRAT. The runtime improvements of ADMIXTURE rely on a fast block relaxation scheme using sequential quadratic programming for block updates, coupled with a novel quasi-Newton acceleration of convergence. Our algorithm also runs faster and with greater accuracy than the implementation of an Expectation-Maximization (EM) algorithm incorporated in the program FRAPPE. Our simulations show that ADMIXTURE's maximum likelihood estimates of the underlying admixture coefficients and ancestral allele frequencies are as accurate as structure's Bayesian estimates. On real-world data sets, ADMIXTURE's estimates are directly comparable to those from structure and EIGENSTRAT. Taken together, our results show that ADMIXTURE's computational speed opens up the possibility of using a much larger set of markers in model-based ancestry estimation and that its estimates are suitable for use in correcting for population stratification in association studies.
Competing models of entrepreneurial intentions
Practical Issues in Structural Modeling
Practical problems that are frequently encountered in applications of covariance structure analysis are discussed and solutions are suggested. Conceptual, statistical, and practical requirements for structural modeling are reviewed to indicate how basic assumptions might be violated. Problems associated with estimation, results, and model fit are also mentioned. Various issues in each area are raised, and possible solutions are provided to encourage more appropriate and successful applications of structural modeling.
The regional oceanic modeling system (ROMS): a split-explicit, free-surface, topography-following-coordinate oceanic model
Astrocytes: biology and pathology
Astrocytes are specialized glial cells that outnumber neurons by over fivefold. They contiguously tile the entire central nervous system (CNS) and exert many essential complex functions in the healthy CNS. Astrocytes respond to all forms of CNS insults through a process referred to as reactive astrogliosis, which has become a pathological hallmark of CNS structural lesions. Substantial progress has been made recently in determining functions and mechanisms of reactive astrogliosis and in identifying roles of astrocytes in CNS disorders and pathologies. A vast molecular arsenal at the disposal of reactive astrocytes is being defined. Transgenic mouse models are dissecting specific aspects of reactive astrocytosis and glial scar formation in vivo. Astrocyte involvement in specific clinicopathological entities is being defined. It is now clear that reactive astrogliosis is not a simple all-or-none phenomenon but is a finely gradated continuum of changes that occur in context-dependent manners regulated by specific signaling events. These changes range from reversible alterations in gene expression and cell hypertrophy with preservation of cellular domains and tissue structure, to long-lasting scar formation with rearrangement of tissue structure. Increasing evidence points towards the potential of reactive astrogliosis to play either primary or contributing roles in CNS disorders via loss of normal astrocyte functions or gain of abnormal effects. This article reviews (1) astrocyte functions in healthy CNS, (2) mechanisms and functions of reactive astrogliosis and glial scar formation, and (3) ways in which reactive astrocytes may cause or contribute to specific CNS disorders and lesions.
Impulse response analysis in nonlinear multivariate models
Genetic studies of body mass index yield new insights for obesity biology
Estimating the global incidence of traumatic brain injury
OBJECTIVE: Traumatic brain injury (TBI)-the "silent epidemic"-contributes to worldwide death and disability more than any other traumatic insult. Yet, TBI incidence and distribution across regions and socioeconomic divides remain unknown. In an effort to promote advocacy, understanding, and targeted intervention, the authors sought to quantify the case burden of TBI across World Health Organization (WHO) regions and World Bank (WB) income groups. METHODS: Open-source epidemiological data on road traffic injuries (RTIs) were used to model the incidence of TBI using literature-derived ratios. First, a systematic review on the proportion of RTIs resulting in TBI was conducted, and a meta-analysis of study-derived proportions was performed. Next, a separate systematic review identified primary source studies describing mechanisms of injury contributing to TBI, and an additional meta-analysis yielded a proportion of TBI that is secondary to the mechanism of RTI. Then, the incidence of RTI as published by the Global Burden of Disease Study 2015 was applied to these two ratios to generate the incidence and estimated case volume of TBI for each WHO region and WB income group. RESULTS: Relevant articles and registries were identified via systematic review; study quality was higher in the high-income countries (HICs) than in the low- and middle-income countries (LMICs). Sixty-nine million (95% CI 64-74 million) individuals worldwide are estimated to sustain a TBI each year. The proportion of TBIs resulting from road traffic collisions was greatest in Africa and Southeast Asia (both 56%) and lowest in North America (25%). The incidence of RTI was similar in Southeast Asia (1.5% of the population per year) and Europe (1.2%). The overall incidence of TBI per 100,000 people was greatest in North America (1299 cases, 95% CI 650-1947) and Europe (1012 cases, 95% CI 911-1113) and least in Africa (801 cases, 95% CI 732-871) and the Eastern Mediterranean (897 cases, 95% CI 771-1023). The LMICs experience nearly 3 times more cases of TBI proportionally than HICs. CONCLUSIONS: Sixty-nine million (95% CI 64-74 million) individuals are estimated to suffer TBI from all causes each year, with the Southeast Asian and Western Pacific regions experiencing the greatest overall burden of disease. Head injury following road traffic collision is more common in LMICs, and the proportion of TBIs secondary to road traffic collision is likewise greatest in these countries. Meanwhile, the estimated incidence of TBI is highest in regions with higher-quality data, specifically in North America and Europe.
Introduction to Plasma Physics and Controlled Fusion
Mapping genomic loci implicates genes and synaptic biology in schizophrenia
Schizophrenia has a heritability of 60–80%1, much of which is attributable to common risk alleles. Here, in a two-stage genome-wide association study of up to 76,755 individuals with schizophrenia and 243,649 control individuals, we report common variant associations at 287 distinct genomic loci. Associations were concentrated in genes that are expressed in excitatory and inhibitory neurons of the central nervous system, but not in other tissues or cell types. Using fine-mapping and functional genomic data, we identify 120 genes (106 protein-coding) that are likely to underpin associations at some of these loci, including 16 genes with credible causal non-synonymous or untranslated region variation. We also implicate fundamental processes related to neuronal function, including synaptic organization, differentiation and transmission. Fine-mapped candidates were enriched for genes associated with rare disruptive coding variants in people with schizophrenia, including the glutamate receptor subunit GRIN2A and transcription factor SP4, and were also enriched for genes implicated by such variants in neurodevelopmental disorders. We identify biological processes relevant to schizophrenia pathophysiology; show convergence of common and rare variant associations in schizophrenia and neurodevelopmental disorders; and provide a resource of prioritized genes and variants to advance mechanistic studies. A genome-wide association study including over 76,000 individuals with schizophrenia and over 243,000 control individuals identifies common variant associations at 287 genomic loci, and further fine-mapping analyses highlight the importance of genes involved in synaptic processes.
Global guideline for the diagnosis and management of mucormycosis: an initiative of the European Confederation of Medical Mycology in cooperation with the Mycoses Study Group Education and Research Consortium
Genome-wide association study of more than 40,000 bipolar disorder cases provides new insights into the underlying biology
Bipolar disorder is a heritable mental illness with complex etiology. We performed a genome-wide association study of 41,917 bipolar disorder cases and 371,549 controls of European ancestry, which identified 64 associated genomic loci. Bipolar disorder risk alleles were enriched in genes in synaptic signaling pathways and brain-expressed genes, particularly those with high specificity of expression in neurons of the prefrontal cortex and hippocampus. Significant signal enrichment was found in genes encoding targets of antipsychotics, calcium channel blockers, antiepileptics and anesthetics. Integrating expression quantitative trait locus data implicated 15 genes robustly linked to bipolar disorder via gene expression, encoding druggable targets such as HTR6, MCHR1, DCLK3 and FURIN. Analyses of bipolar disorder subtypes indicated high but imperfect genetic correlation between bipolar disorder type I and II and identified additional associated loci. Together, these results advance our understanding of the biological etiology of bipolar disorder, identify novel therapeutic leads and prioritize genes for functional follow-up studies. Genome-wide association analyses of 41,917 bipolar disorder cases and 371,549 controls of European ancestry provide new insights into the etiology of this disorder and identify novel therapeutic leads and potential opportunities for drug repurposing.
Applying evolutionary biology to address global challenges
BACKGROUND Differences among species in their ability to adapt to environmental change threaten biodiversity, human health, food security, and natural resource availability. Pathogens, pests, and cancers often quickly evolve resistance to control measures, whereas crops, livestock, wild species, and human beings often do not adapt fast enough to cope with climate change, habitat loss, toxicants, and lifestyle change. To address these challenges, practices based on evolutionary biology can promote sustainable outcomes via strategic manipulation of genetic, developmental, and environmental factors. Successful strategies effectively slow unwanted evolution and reduce fitness in costly species or improve performance of valued organisms by reducing phenotype-environment mismatch or increasing group productivity. Tactics of applied evolutionary biology range broadly, from common policies that promote public health or preserve habitat for threatened species—but are easily overlooked as having an evolutionary rationale, to the engineering of new genomes. ADVANCES The scope and development of current tactics vary widely. In particular, genetic engineering attracts much attention (and controversy) but now is used mainly for traits under simple genetic control. Human gene therapy, which mainly involves more complex controls, has yet to be applied successfully at large scales. In contrast, other methods to alter complex traits are improving. These include artificial selection for drought- and flood-tolerant crops through bioinformatics and application of “life course” approaches in medicine to reduce human metabolic disorders. Successful control of unwanted evolution depends on governance initiatives that address challenges arising from both natural and social factors. Principal among these challenges are (i) global transfer of genes and selection agents; (ii) interlinked evolution across traditional sectors of society (environment, food, and health); and (iii) conflicts between individual and group incentives that threaten regulation of antibiotic use and crop refuges. Evolutionarily informed practices are a newer prospect in some fields and require more systematic research, as well as ethical consideration—for example, in attempts to protect wild species through assisted migration, in the choice of source populations for restoration, or in genetic engineering. OUTLOOK A more unified platform will better convey the value of evolutionary methods to the public, scientists, and decision-makers. For researchers and practitioners, applications may be expanded to other disciplines, such as in the transfer of refuge strategies that slow resistance evolution in agriculture to slow unwanted evolution elsewhere (for example, cancer resistance or harvest-induced evolution). For policy-makers, adoption of practices that minimize unwanted evolution and reduce phenotype-environment mismatch in valued species is likely essential to achieve the forthcoming Sustainable Development Goals and the 2020 Aichi Biodiversity Targets.
Is infrared-collinear safe information all you need for jet classification?
A bstract Machine learning-based jet classifiers are able to achieve impressive tagging performance in a variety of applications in high-energy and nuclear physics. However, it remains unclear in many cases which aspects of jets give rise to this discriminating power, and whether jet observables that are tractable in perturbative QCD such as those obeying infrared-collinear (IRC) safety serve as sufficient inputs. In this article, we introduce a new classifier, Jet Flow Networks (JFNs), in an effort to address the question of whether IRC unsafe information provides additional discriminating power in jet classification. JFNs are permutation-invariant neural networks (deep sets) that take as input the kinematic information of reconstructed subjets. The subjet radius and a cut on the subjet’s transverse momenta serve as tunable hyperparameters enabling a controllable sensitivity to soft emissions and nonperturbative effects. We demonstrate the performance of JFNs for quark vs. gluon and Z vs. QCD jet tagging. For small subjet radii and transverse momentum cuts, the performance of JFNs is equivalent to the IRC-unsafe Particle Flow Networks (PFNs), demonstrating that infrared-collinear unsafe information is not necessary to achieve strong discrimination for both cases. As the subjet radius is increased, the performance of the JFNs remains essentially unchanged until physical thresholds that we identify are crossed. For relatively large subjet radii, we show that the JFNs may offer an increased model independence with a modest tradeoff in performance compared to classifiers that use the full particle information of the jet. These results shed new light on how machines learn patterns in high-energy physics data.