Hive - A Warehousing Solution Over a Map-Reduce Framework

Keyword search

Guided search

Click a term to initiate a search.

Hive - A Warehousing Solution Over a Map-Reduce Framework

Mon, 10/19/2009 - 10:15 — admin

Authors:

Thusoo, Ashish; Sarma, Joydeep Sen; Jain, Namit; Shao, Zheng; Chakka, Prasad; Anthony, Suresh; Liu, Hao; Wyckoff, Pete; Murthy, Raghotham

Author:

Liu, H

Wyckoff, P

Murthy, R

Anthony, S

Chakka, P

Jain, N

Shao, Z

Thusoo, A

Sarma, J

The size of data sets being collected and analyzed in the
industry for business intelligence is growing rapidly, mak-
ing traditional warehousing solutions prohibitively expen-
sive. Hadoop [3] is a popular open-source map-reduce im-
plementation which is being used as an alternative to store
and process extremely large data sets on commodity hard-
ware. However, the map-reduce programming model is very
low level and requires developers to write custom programs
which are hard to maintain and reuse.
In this paper, we present Hive, an open-source data ware-
housing solution built on top of Hadoop. Hive supports
queries expressed in a SQL-like declarative language - HiveQL,
which are compiled into map-reduce jobs executed on Hadoop.
In addition, HiveQL supports custom map-reduce scripts to
be plugged into queries. The language includes a type sys-
tem with support for tables containing primitive types, col-
lections like arrays and maps, and nested compositions of
the same. The underlying IO libraries can be extended to
query data in custom formats. Hive also includes a system
catalog, Hive-Metastore, containing schemas and statistics,
which is useful in data exploration and query optimization.
In Facebook, the Hive warehouse contains several thousand
tables with over 700 terabytes of data and is being used ex-
tensively for both reporting and ad-hoc analyses by more
than 100 users.

Year:

2009

Venue:

VLDB 2009

URL:

http://www.vldb.org/pvldb/2/vldb09-938.pdf

Citations:

Citations range:

n/a

Attachment	Size
Liu2009HiveAWarehousingSolutionOveraMapReduceFramework.pdf	250.44 KB

websearch

Cloud Computing publication categorizer

Keyword search

Guided search

Author

Year

Topic

Tags

mailpart

Citations range

Hive - A Warehousing Solution Over a Map-Reduce Framework

Navigation

Related categories

User login