Using Spark to Load Oracle Data into Cassandra (Jim Hatcher, IHS Markit) | C* Summit 2016

Jim Hatcher

Using Spark to Load Oracle Data into Cassandra

1 Introduction

2 Problem Description

3 Methods of loading external data into Cassandra

4 What is Spark?

5 Lessons Learned

6 Resources

Introduction

At IHS Markit, we take raw data and turn it into information and insights for our customers.

Automotive Systems (CarFax)Defense Systems (Jane’s)Oil & Gas Systems (Petra, Kingdom)Maritime SystemsTechnology Systems (Electronic Parts Database, Root Metrics)ChemicalsFinancial Systems (Wall Street on Demand)Lots of others

Problem Description

Cluster

Factory

Oracle

Back-end Applications Customer-facing Systems

Load Files

Customer-facing

Applications

Oracle

Cassandra+

SolrFactory

Applications

Data Updates

Cassandra+

Methods of loading external data into Cassandra

Methods of Loading External Data into C*

1. CQL Copy command2. Sqoop3. Write a custom program that uses the CQL driver4. Write a Spark program

What is Spark?

What is Spark?Spark is a processing framework designed to work with distributed data.

“up to 100X faster than MapReduce” according to spark.apache.org

Used in any ecosystem where you want to work with distributed data (Hadoop, Cassandra, etc.)

Includes other specialized libraries:• SparkSQL• Spark Streaming• MLLib• GraphX

Spark Facts

Conceptually Similar To MapReduce

Written In Scala

Supported By DataBricks

Supported Languages Scala, Java, Python, R

Spark Client

DriverSpark

Context

Spark Master

Spark Worker

Executor

1. Request Resources2. Allocate Resources

Credit: https://academy.datastax.com/courses/ds320-analytics-apache-spark/introduction-spark-architecture

Spark Architecture

Spark with CassandraCredit: https://academy.datastax.com/courses/ds320-analytics-apache-spark/introduction-spark-architecture

Cassandra Cluster

Spark Worker

Spark WorkerSpark Worker

Spark Master

Spark Client

Spark Cassandra Connector – open source, supported by DataStaxhttps://github.com/datastax/spark-cassandra-connector

ETL (Extract, Transform, Load)

Text File

JDBC Data Source

Cassandra

Hadoop

Extract Data

Spark: Create RDD or Data

Data Source(s)Spark Code

Transform Data

Spark: Map function

Spark Code

Cassandra

Data Source(s)

Load Data

Spark: Save

Spark Code

Typical Code - Example// Extractval extracted = sqlContext .read .format("jdbc") .options( Map[String, String]( "url" -> "jdbc:oracle:thin:username/password@//hostname:port/oracle_svc", "dbtable" -> "table_name" ) ) .load()

// Transformval transformed = extracted.map { dbRow => (dbRow.getAs[String](“field_one"), dbRow.getAs[Integer](“field_two"))}

// Loadtransformed.saveToCassandra(“keyspace_name", “table_name", SomeColumns(“field_one“, “field_two"))

Lessons Learned

Lesson #1 - Spark SQL handles Oracle NUMBER fields with no precision incorrectlyhttps://issues.apache.org/jira/browse/SPARK-10909

All of our Oracle tables have ID fields defined as NUMBER(15,0).

When you use Spark SQL to access an Oracle table, there is a piece of code in the JDBC driver that reads the metadata and creates a dataframe with the proper schema. If your schema has a NUMBER(*, 0) field defined in it, you get a “Overflowed precision” error.

This is fixed in Spark 1.5, but we don’t have the option of adopting a new version of Spark since we’re using Spark bundled with DSE 4.8.6 (which uses spark 1.4.2). We were able to fix this by stealing the fix from the Spark 1.5 code and applying it to our code (yay, open source!).

At some point, we’ll update to DSE 5.* which uses Spark 1.6, and we can remove this code.

import java.sql.Typesimport org.apache.spark.sql.jdbc.{JdbcDialect, JdbcType}import org.apache.spark.sql.types._

private case object OracleDialect extends JdbcDialect {

override def canHandle(url: String): Boolean = url.startsWith("jdbc:oracle")

override def getCatalystType(sqlType: Int, typeName: String, size: Int, md: MetadataBuilder): Option[DataType] = { // Handle NUMBER fields that have no precision/scale in special way // because JDBC ResultSetMetaData converts this to 0 precision and -127 scale // For more details, please see // https://github.com/apache/spark/pull/8780#issuecomment-145598968 // and // https://github.com/apache/spark/pull/8780#issuecomment-144541760 if (sqlType == Types.NUMERIC && size == 0) { // This is sub-optimal as we have to pick a precision/scale in advance whereas the data // in Oracle is allowed to have different precision/scale for each value. Option(DecimalType(38, 10)) } else { None } }

override def getJDBCType(dt: DataType): Option[JdbcType] = dt match { case StringType => Some(JdbcType("VARCHAR2(255)", java.sql.Types.VARCHAR)) case _ => None }

}org.apache.spark.sql.jdbc.JdbcDialects.registerDialect(OracleDialect)

Lesson #1 - Spark SQL handles Oracle NUMBER fields with no precision incorrectly

Lesson #2 - Spark SQL doesn’t handle timeuuid fields correctly https://issues.apache.org/jira/browse/SPARK-10501

Spark SQL doesn’t know what to do with a timeuuid field when reading a table from Cassandra. This is an issue since we commonly use timeuuid columns in our Cassandra key structures.

We got this error: scala.MatchError: UUIDType (of class org.apache.spark.sql.cassandra.types.UUIDType$)

We are able to work around this issue by casting the timeuuid values to strings, like this:

val dataFrameRaw = sqlContext .read .format("org.apache.spark.sql.cassandra") .options(Map("table" -> "table_name", "keyspace" -> "keyspace_name")) .load()

val dataFrameFixed = dataFrameRaw .withColumn(“timeuuid_column", dataFrameRaw("timeuuid_column").cast(StringType))

Lesson #3 – Careful when generating ID fields

We created an RDD:

val baseRdd = rddInsertsAndUpdates.map { dbRow =>

val keyColumn = { if (!dbRow.isNullAt(dbRow.fieldIndex(“timeuuid_key_column"))) { dbRow.getAs[String]("timeuuid_key_column") } else { UUIDs.timeBased().toString } }

//do some further processing

(keyColumn, …other values)}

Then, we took that RDD and transformed it into another RDD:

val invertedIndexTable = baseRdd.map { entry => (entry.getString(“timeuuid_key_column"), entry.getString(“fld_1"))}

Then we wrote them both to C*, like this:

baseRdd.saveToCassandra(“keyspace_name", “table_name", SomeColumns(“key_column“, “fld_1“, “fld_2"))

invertedIndexTable.saveToCassandra(“keyspace_name", “inverted_index_table_name" SomeColumns(“key_column“, “fld_1“)

Lesson #3 – Careful when generating ID fields

We kept finding that the ID values in the inverted index table had slightly different ID values than the values in the base table.

We fixed this by adding a cache() to our first RDD.

val baseRdd = rddInsertsAndUpdates.map { dbRow =>

val keyColumn = { if (!dbRow.isNullAt(dbRow.fieldIndex(“timeuuid_key_column"))) { dbRow.getAs[String]("timeuuid_key_column") } else { UUIDs.timeBased().toString } }

//do some further processing

(keyColumn, …other values)}.cache()

Lesson #4 – You can only return an RDD of a tuple if you have 22 items or less.

It’s pretty common in Spark to return an RDD of tuplesval myNewRdd = myOldRdd.map { dbRow =>

val firstName = dbRow.getAs[String](“FirstName") val lastName = dbRow.getAs[String](“LastName") val calcField1 = dbRow.getAs[Intger](“SomeColumn") * 3.14

(firstName, lastName, calcField1)}

This works great until you get to 22 fields in your tuple, and then Scala throws an error. (Later versions of Scala lift this restriction, but it’s a problem for our version of Scala.)

Lesson #4 – You can only return an RDD of a tuple if you have 22 items or less.

You can fix this by returning an RDD of CassandraRows instead. (especially if your goal is to save them to C*)val myNewRdd = myOldRdd.map { dbRow =>

val firstName = dbRow.getAs[String](“FirstName") val lastName = dbRow.getAs[String](“LastName") val calcField1 = dbRow.getAs[Integer](“SomeColumn") * 3.14 val allValues = IndexedSeq[AnyRef](firstName, lastName, calcField1) val allColumnNames = Array[String]( “first_name", “last_name", “calc_field_1“) new CassandraRow(allColumnNames, allValues)}

Lesson #5 – Getting a JDBC dataframe based on a SQL statement is not very intuitive.To get a dataframe from a JDBC source, you do this:

val exampleDataFrame = sqlContext .read .format("jdbc") .options( Map[String, String]( "url" -> "jdbc:oracle:thin:username/password@//hostname:port/oracle_svc", "dbtable" -> "table_name" ) ) .load()

You would think there would be a version of this call that lets you pass in a SQL statement but there is not.

However, when JDBC creates your query from the above syntax, all it does is prepend your dbtable value with “SELECT * FROM”.

Lesson #5 – Getting a JDBC dataframe based on a SQL statement is not very intuitive.So, the workaround is to do this:

val sql = "( " + " SELECT S.* " + " FROM Sample S " + " WHERE ID = 11111 " + " ORDER BY S.SomeField " + ")" val exampleDataFrame = sqlContext .read .format("jdbc") .options( Map[String, String]( "url" -> "jdbc:oracle:thin:username/password@//hostname:port/oracle_svc", "dbtable" -> sql ) ) .load()

Lesson #6 – Creating a partitioned JDBC dataframe is not very intuitive.The code to get a JDBC dataframe looks like this:val basePartitionedOracleData = sqlContext .read .format("jdbc") .options( Map[String, String]( "url" -> "jdbc:oracle:thin:username/password@//hostname:port/oracle_svc", "dbtable" -> "ExampleTable", "lowerBound" -> "1", "upperBound" -> "10000", "numPartitions" -> "10", "partitionColumn" -> “KeyColumn" ) ) .load()

The last four arguments in that map are there for the purpose of getting a partitioned dataset. If you pass any of them, you have to pass all of them.

Lesson #6 – Creating a partitioned JDBC dataframe is not very intuitive.When you pass these additional arguments in, here’s what it does:

It builds a SQL statement template in the format “SELECT * FROM {tableName} WHERE {partitionColumn} >= ? AND {partitionColumn} < ?”

It sends {numPartitions} statements to the DB engine. If you suppled these values: {dbTable=ExampleTable, lowerBound=1, upperBound=10,000, numPartitions=10, partitionColumn=KeyColumn}, it would create these ten statements:

SELECT * FROM ExampleTable WHERE KeyColumn >= 1 AND KeyColumn < 1001

Lesson #7 – JDBC *really* wants you to get your partitioned dataframe using a sequential ID column.In our Oracle database, we don’t have sequential integer ID columns.

We tried to get around that by doing a query like this and passing “ROW_NUMBER” as the partitioning column:SELECT ST.*, ROW_NUMBER() OVER (ORDER BY ID_FIELD ASC) AS ROW_NUMBER FROM SourceTable STWHERE …my criteriaORDER BY ID_FIELDBut, this didn’t perform well.

We ended up creating a processing table:

CREATE TABLE SPARK_ETL_BATCH_SEQUENCE (SEQ_ID NUMBER(15,0) NOT NULL, //this has a sequence that gets auto-incrementedBATCH_ID NUMBER(15,0) NOT NULL,ID_FIELD NUMBER(15,0) NOT NULL

Lesson #7 – JDBC *really* wants you to get your partitioned dataframe using a sequential ID column.We insert into this table first:INSERT INTO SPARK_ETL_BATCH_SEQUENCE ( BATCH_ID, ID_FIELD ) //SEQ_ID gets auto-populatedSELECT {NextBatchID}, ID_FIELDFROM SourceTable STWHERE …my criteriaORDER BY ID_FIELD

Then, we join to it in the query where we get our data which provides us with a sequential ID:SELECT ST.*, SEQ.SEQ_IDFROM SourceTable STINNER JOIN SPARK_ETL_BATCH_SEQUENCE SEQ ON ST.ID_FIELD = SEQ.ID_FIELDWHERE …my criteriaORDER BY ID_FIELDAnd, we use SEQ_ID as our Partitioning Column.Despite its need to talk to Oracle twice, this approach has proven to perform much faster than having uneven partitions.

Resources

Spark• Books

• Learning Sparkhttp://shop.oreilly.com/product/0636920028512.do

Scala (Knowing Scala with really help you progress in Spark)• Functional Programming Principles in Scala (videos)

https://www.youtube.com/user/afigfigueira/playlists?shelf_id=9&view=50&sort=dd• Books

http://www.scala-lang.org/documentation/books.html

Spark and Cassandra• DataStax Academy

http://academy.datastax.com/• Self-paced course: DS320: DataStax Enterprise Analytics with Apache Spark – Really Good!• Tutorials

• Spark Cassandra Connector website – lots of good exampleshttps://github.com/datastax/spark-cassandra-connector

Using Spark to Load Oracle Data into Cassandra (Jim Hatcher, IHS Markit) | C* Summit 2016

Software

Transcript of Using Spark to Load Oracle Data into Cassandra (Jim Hatcher, IHS Markit) | C* Summit 2016

Technical Guide | November 2018 · The Navigator is a GUI for the IHS Markit Securities Finance Toolkit. A IHS Markit icon or menu item will be added to the Excel menu when the Toolkit

IHS Markit Software Installation Guide (PDF)onlinehelp.ihs.com/Energy/License_Manager/1.0.0/IHS... · 2018. 4. 3. · IHS Markit Software Installation Guide (PDF) Created Date: 4/3/2018

Menlo Park, California - IHS Markit

IHS MARKIT / BME GERMANY MANUFACTURING PMI®...IHS Markit / BME Germany Manufacturing PMI® 2018 IHS Markit OUTPUT INDEX 20 30 40 50 60 70 '96 '98 '00 '02 '04 '06 '08 '10 '12 '14 '16

REVENUE 2016 - iab Interactive Advertising Bureau Austria€¦ · GLOBAL MOBILE REVENUE 2016 IAB EUROPE & IAB USA & IHS Markit Die Zahlen wurden von IAB Europe, IAB USA und IHS Markit

Week Ahead Economic Overview - IHS Markit · By Joe Hayes Economist, IHS Markit, London Email: joseph.hayes@ihsmarkit.com Key data for Europe in the week ahead includes flash PMI

汽车显示面板市场竞争格局变化 - Markit来源：IHS Markit Automotive Display Market Tracker IHS Markit 汽车显示器市场跟踪 机密。© 2018 IHS MarkitTM。版权所有。市场重新洗牌

7 in 2017 - IHS Markit · 7 in 2017: Top global technology trends in 2017 T he 7 big themes identified in this paper represent what we at IHS Markit Technology believe will be the

The Big Prize: How Collaboration Transforms Workflow - IHS Markit

IHS Markit Ltd. - Climate Change 2020€¦ · IHS Markit Ltd. - Climate Change 2020 C0. Introduction C0.1 (C0.1) Give a general description and introduction to your organization.

Pennsylvania’s Opportunities in Petrochemical ... · PDF fileMarch 2017 8 © 2017 IHS Markit IHS Markit | Prospects to Enhance Pennsylvania’s Opportunities in Petrochemical Manufacturing

Country Intelligence Monitor - IHS Markit › www › pdf › 0819 › Connect...Confidential. © 2020 IHS Markit®.All rights reserved. Connect Login Instructions 3

Presentation to MBS 2017 Forecast · PDF file© 2017 IHS Markit. All Rights Reserved.© 2017 IHS Markit. All Rights Reserved. Presentation to MBS 2017 – Forecast Panel Michael Robinet

Monthly Model Performance Report - IHS Markit · 2018. 12. 5. · Monthly Performance Recap. IHS Markit Model Matrix. Deep Value. Earnings Momentum. Price Momentum. Relative Value.

IHS ENERGY US Land Data - IHS Markit · 2015-03-19 · IHS ENERGY US Land Data Identify opportunities and monitor competitor activity through comprehensive land and lease data from

Chartis risk tech 100 vendor highlights IHS Markit

Cargo Risks Outlook 2020 - IHS Markit · 2/12/2020 · IHS Markit owns all IHS Markit logos and trade names contained in this presentation that are subject to license. Opinions,

November 2017 Author: Daniel Knapp, IHS Markit - IAB Slovakia · IAB Europe and IHS Markit. All data & analysis must be quoted as “IAB Europe ... This is an interim update of the

OCTOBER 2016 OLEFINS AND METHANOL - IHS Markit · © 2016 IHS. ALL RIGHTS RESERVED. Ethylene The Market 34th Annual World Methanol Conference/ Oct 2016

The Economic Contribution of Independent …...Bob Fryklund Vice President, IHS Markit Upstream Energy Curtis Smith Director, IHS Markit Upstream Consulting Bob Flanagan Director,

汽车显示面板市场竞争格局变化 - Markit来源：IHS Markit Automotive Display Market Tracker IHS Markit 汽车显示器市场跟踪机密。© 2018 IHS MarkitTM。版权所有。市场重新洗牌