University of Florida :: Department of Computer and Information Science and Engineering (CISE)

Departmental Report : REP-2012-548

Search Departmental Reports
By Author: 
By Time: 
Keyword search:
Report ID:REP-2012-548
Title:A Sampling Algebra for Aggregate Estimation
Authors:Supriya Nirkhiwale, PhD candidate

University of Florida
Dept of Computer and Information Science,
University of Florida
650 283 9690

Alin Dobra, Associate Professor

University of Florida
Dept of Computer and Information Science

Chris Jermaine, Associate Professor

Rice University
3028 Duncan Hall
6100 Main St, Houston, TX, 77005
713 348 5690
Abstract:

As of 2005, sampling has been incorporated in a ll major databases. While efficient sampling techniques are easily realizable, determining the accuracy of an estimate obtained from the sample is still an unresolved problem. In this paper, we present a theoretical framework that allows an elegant treatment of the problem. We base our work on generalized uniform sampling (GUS), a class of sampling methods that subsumes a wide variety of sampling techniques. We introduce a key notion of equivalence that allows GUS sampling operators to commute with selection and join, and derivation of confidence intervals. We illustrate the theory through extensive examples and give indications on how to use it to provide meaningful estimations in database systems.

Posted:May 4, 2012