skip to main content
Resource type Show Results with: Show Results with: Index

Large-Scale Bayesian Logistic Regression for Text Categorization

Genkin, Alexander ; Lewis, David D ; Madigan, David

Technometrics, 01 August 2007, Vol.49(3), pp.291-304 [Peer Reviewed Journal]

Full text available online

Citations Cited by
  • Title:
    Large-Scale Bayesian Logistic Regression for Text Categorization
  • Author/Creator: Genkin, Alexander ; Lewis, David D ; Madigan, David
  • Language: English
  • Subjects: Information Retrieval ; Lasso ; Penalization ; Ridge Regression ; Support Vector Classifier ; Variable Selection ; Engineering ; Statistics ; Mathematics
  • Is Part Of: Technometrics, 01 August 2007, Vol.49(3), pp.291-304
  • Description: Logistic regression analysis of high-dimensional data, such as natural language text, poses computational and statistical challenges. Maximum likelihood estimation often fails in these applications. We present a simple Bayesian logistic regression approach that uses a Laplace prior to avoid overfitting and produces sparse predictive models for text data. We apply this approach to a range of document classification problems and show that it produces compact predictive models at least as effective as those produced by support vector machine classifiers or ridge logistic regression combined with feature selection. We describe our model fitting algorithm, our open source implementations (BBR and BMR), and experimental results.
  • Identifier: ISSN: 0040-1706 ; E-ISSN: 1537-2723 ; DOI: 10.1198/004017007000000245