Hive 8、Hive2 beeline 和 Hive jdbc,Hive的UDF、UDAF、UDTF

来源：互联网发布：java log日志编辑：程序博客网时间：2024/06/05 10:18

1、Hive2 beeline

Beeline 要与HiveServer2配合使用，支持嵌入模式和远程模式

启动beeline

打开两个Shell窗口，一个启动Hive2 一个beeline连接hive2

#启动HiverServer2 , ./bin/hiveserver2 [root@node5 ~]# hiveserver216/02/23 22:55:25 WARN conf.HiveConf: HiveConf of name hive.metastore.local does not exist

#启动Beeline # $ ./bin/beeline 下面的方式需要配置Hive的环境变量[root@node5 ~]# beelineBeeline version 1.2.1 by Apache Hivebeeline>

启动beeline之后可以尝试连接hiveserver2

beeline> !connect jdbc:hive2://node5:10000Connecting to jdbc:hive2://node5:10000Enter username for jdbc:hive2://node5:10000: #默认 用户名就是登录账号 密码为空

2、Hive jdbc

打开Eclipse 新建一个Java 项目：

public class Demo {    public static void main(String[] args) {        try {            Class.forName("org.apache.hive.jdbc.HiveDriver");            Connection conn = DriverManager.getConnection("jdbc:hive2://node5:10000/hive","root","123456");            String sql = "select * from news";            Statement sment = conn.createStatement();            ResultSet rs = sment.executeQuery(sql);            while(rs.next()){                HiveQueryResultSet hqrs = (HiveQueryResultSet)rs;                System.out.println(hqrs.getString(1)+"\t"+hqrs.getString(2));            }            conn.close();        } catch (Exception e) {            e.printStackTrace();        }    }}

在Hive数据库中有这样一个表 news

hive> use hive    > ;OKTime taken: 4.244 secondshive> select * from news;OK1    I'm tom2    what are you doing3    i'm okTime taken: 1.247 seconds, Fetched: 3 row(s)

执行完Java代码以后，可以看到数据正常查询出来了：

Hive的UDF、UDAF、UDTF

Hive自定义函数包括三种UDF、UDAF、UDTF

　　UDF(User-Defined-Function) 一进一出

　　UDAF(User- Defined Aggregation Funcation) 聚集函数，多进一出。Count/max/min

　　UDTF(User-Defined Table-Generating Functions) 一进多出，如lateral view explore()

　　使用方式：在HIVE会话中add 自定义函数的jar文件，然后创建function继而使用函数

UDF

1、UDF函数可以直接应用于select语句，对查询结构做格式化处理后，再输出内容。

2、编写UDF函数的时候需要注意一下几点：

　　a）自定义UDF需要继承org.apache.hadoop.hive.ql.UDF。

　　b）需要实现evaluate函数，evaluate函数支持重载。

　　例：写一个返回字符串长度的Demo:

import org.apache.hadoop.hive.ql.exec.UDF;public class GetLength extends UDF{    public int evaluate(String str) {        try{            return str.length();        }catch(Exception e){            return -1;        }    }}

3、步骤

　　a）把程序打包放到目标机器上去；

　　b）进入hive客户端，添加jar包：

hive> add jar /root/hive_udf.jar

　　c）创建临时函数：

hive> create temporary function getLen as 'com.raphael.len.GetLength';

　　d）查询HQL语句：

hive> select getLen(info) from apachelog;OK6029871026960677966Time taken: 0.072 seconds, Fetched: 9 row(s)

　　e）销毁临时函数：

hive> DROP TEMPORARY FUNCTION getLen;

UDAF

多行进一行出，如sum()、min()，用在group by时

1.必须继承

　　org.apache.hadoop.hive.ql.exec.UDAF(函数类继承)

　　org.apache.hadoop.hive.ql.exec.UDAFEvaluator(内部类Evaluator实现UDAFEvaluator接口)

2.Evaluator需要实现 init、iterate、terminatePartial、merge、terminate这几个函数

　　init():类似于构造函数，用于UDAF的初始化

　　iterate():接收传入的参数，并进行内部的轮转，返回boolean

　　terminatePartial():无参数，其为iterate函数轮转结束后，返回轮转数据，类似于hadoop的Combiner

　　merge():接收terminatePartial的返回结果，进行数据merge操作，其返回类型为boolean

　　terminate():返回最终的聚集函数结果

　　#开发一个功能同：

　　#Oracle的wm_concat()函数

　　#Mysql的group_concat()

　　UDAF 详细文档：http://www.cnblogs.com/ggjucheng/archive/2013/02/01/2888051.html

UDTF

　　UDTF 详细文档: http://www.cnblogs.com/ggjucheng/archive/2013/02/01/2888819.html

阅读全文

0 0