Amazon Athena JDBC Driver Wrapper Supporting the 'metis' Package
Vous ne pouvez pas sélectionner plus de 25 sujets Les noms de sujets doivent commencer par une lettre ou un nombre, peuvent contenir des tirets ('-') et peuvent comporter jusqu'à 35 caractères.
boB Rudis a01ab3351f
initial commit
il y a 7 ans
R initial commit il y a 7 ans
inst initial commit il y a 7 ans
man initial commit il y a 7 ans
tests initial commit il y a 7 ans
.Rbuildignore initial commit il y a 7 ans
.codecov.yml initial commit il y a 7 ans
.gitignore initial commit il y a 7 ans
.travis.yml initial commit il y a 7 ans
DESCRIPTION initial commit il y a 7 ans
NAMESPACE initial commit il y a 7 ans
NEWS.md initial commit il y a 7 ans
README.Rmd initial commit il y a 7 ans
README.md initial commit il y a 7 ans
metis.Rproj initial commit il y a 7 ans

README.md

metis : Helpers for Accessing and Querying Amazon Athena

Including a lightweight RJDBC shim.

THIS IS SUPER ALPHA QUALITY. NOTHING TO SEE HERE. MOVE ALONG.

The goal will be to get around enough of the "gotchas" that are preventing raw RJDBC Athena connecitons from "just working" with dplyr v0.6.0+ and also get around the fetchSize problem without having to not use dbGetQuery().

It will also support more than the vanilla id/secret auth mechism (it currently support the default basic auth and temp token auth, the latter via environment variables).

See the Usage section for an example.

The following functions are implemented:

  • athena_connect: Make a JDBC connection to Athena (this returns an AthenaConnection object which is a super-class of it's RJDBC vanilla counterpart)
  • Athena: AthenaJDBC`
  • AthenaConnection-class: AthenaJDBC
  • AthenaDriver-class: AthenaJDBC
  • AthenaResult-class: AthenaJDBC
  • dbConnect-method: AthenaJDBC
  • dbGetQuery-method: AthenaJDBC
  • dbSendQuery-method: AthenaJDBC

Installation

devtools::install_github("hrbrmstr/metis")

Usage

library(metis)
library(dplyr)

# current verison
packageVersion("metis")
## [1] '0.1.0'
ath <- athena_connect("your_schema_name")

res <- dbGetQuery(ath, "
SELECT format_datetime(timestamp, 'yyyy-MM-dd HH:00:00') timestamp,
        port as field, count(port) cnt_field FROM your_schema_name.your_table_name
        WHERE CONTAINS(ARRAY['201705'], date)
        AND port IN (445, 139, 3389)
        AND timestamp > date '2017-05-01'
        AND timestamp <= date '2017-05-22'
GROUP BY format_datetime(timestamp, 'yyyy-MM-dd HH:00:00'), port LIMIT 1000000
")