Sunday, June 18, 2017

Pinebook!

Not Hadoop-related, but awesome all the same. A few months ago I stumbled on PINE64's website and saw the pinebook, a linux arm64 laptop. That and a PocketCHIP made a great late-birthday, early-father's-day set of presents.

Build and shipping takes a couple months, shipping was almost 1/3 of the laptop cost, and performance and keyboard quality is exactly what you would expect :) But it is still a fun bit of hardware.

If you decide to get one, make sure to add on a USB-to-H-barrel power cord (or make your own). The pinebook does come with a power supply, but no point in carting around yet another wall-wart when the pinebook happily charges off a phone charger.

Mine powered right up into Xenial. I'm normally RH-based since everywhere I've been employed in the last couple decades has been, so it's nice to jump back into Debian-based.

aarch64 wasn't in mainline rust, but was in nursery, so
curl -sSf https://raw.githubusercontent.com/rust-lang-nursery/rustup.rs/master/rustup-init.sh | bash
worked just fine and got me up and going with rust.




Update: HackADay has a great write-up. I didn't experience any of the screen issues they had since I have the 14", but the page has a great tear-down and overview of performance (which is not much :) )

Thursday, February 23, 2017

Writing ORC files is easier than a few years ago

Several years ago I was asked to compare writing Parquet and ORCFile formats from standalone java (without using the Hadoop libraries). At the time ORC was not separated from Hive and it was much more involved than writing Parquet from java. It looks like that changed in 2015 but I only revisited the issue within the past few months.

To build ORC:
Download the current release (currently 1.3.2)
tar xzvf orc-1.3.2.tar.gz && cd ./orc-1.3.2/
cd ./java
mvn package

ls -la ./tools/target/orc-tools-1.3.2-uber.jar

A simple example of writing is:


And a simple example of reading is:


Tuesday, January 6, 2015

Github PrintGZHeader

I receive some data that is a gzip file made of concatenated smaller gzip files. The larger file is a valid gzip according to the rfc. Everything works fine unless I need to look at the original filename or mtime of the gz streams. So... I created a program to print out the original filename and mtime from the header information in all the gzip streams in a gz file.

https://github.com/awcoleman/PrintGZHeader

Compile: gcc -o printGZHeader printGZHeader.c -lz
And run: ./printGZHeader myGZfile.gz

My C is very rusty, I will happily accept any patches to clean up bad practices. My test files do not have header comments or extra fields, please send patches if you find that the code does not work appropriately on them (or send me a test file and I will try).

Hopefully this saves someone else a little bit of time.

Wednesday, December 31, 2014

Nostalgia in Clustering

The close of 2014 made me remember an old clustering project I did around 2004 on a shoestring budget. The project was correlating customers into families, with a sub-task of deduplicating customer records (from typos and other issues). The entire project team was… me.

I gathered up a server with a couple of old hard drives as a mySQL server and PXE boot server, and four other computers PXE-booting into linux with openMOSIX for clustering. I didn’t have budget for cases for the four slaves, so used old cookie sheets to mount them. I used wooden dowels to fix two cookie sheet nodes together so they could sit vertically.

My processing was done in perl. Once OpenMOSIX reported a slave was free, a perl process would spawn and grab a workload from mySQL. OpenMOSIX would migrate the process to the open slave.

Fortunately I was able to complete the project with only four slaves. I had figured out my power supplies could power two slaves. I was working on converting a couple of ATX power supply extension cables into a Y-splitter and only using one power supply per "cookie".

I found some pictures of the nodes from an old presentation:


Sunday, July 6, 2014

Custom Writable

I have never tackled a custom Writable before. I am a huge Avro (http://avro.apache.org/) fan, so I usually try to get my data converted to avro early. A discussion got me interested in tackling it and I had my Bouncy Castle ASN.1 Hadoop example open, so I extended that to a basic custom Writable example.

This thread by Oded Rosen was invaluable:
http://mail-archives.apache.org/mod_mbox/hadoop-general/201005.mbox/%3CAANLkTinzP8-nnGg8Q5aaJ8gXCCg6Som7e8Xarc_2PGDD@mail.gmail.com%3E
(also at http://osdir.com/ml/general-hadoop-apache/2010-05/msg00073.html if above is down)

I put the code in package com.awcoleman.BouncyCastleGenericCDRHadoopWithWritable in github.

The basics from the thread above and a bit of other reading are:
If your class will only be used as a value and not a key, implement the Writable interface.
If your class will be used as a key (and possibly a value), implement the WritableComparable interface (which extends Writable).

A Writable must have 3 things:
An empty contructor. There can be other contructors with arguments, but there must be a no argument one as well.
An overridden write method to write variables out.
An overridden readFields method to populate an object from a previous write method output.

Hadoop reuses Writable objects, so cleaning all variables before populating them in readFields will stop surprises.

WritableComparable adds to Writable:
An overridden hashcode method to partition keys.
An overridden compareTo method.

The advice given in the 'How to write a complex Writable' thread adds:
Override the equals method
Implement RawComparator for your type. This post (http://vangjee.wordpress.com/2012/03/30/implementing-rawcomparator-will-speed-up-your-hadoop-mapreduce-mr-jobs-2/) has an example that extends WritableComparator, which implements RawComparator.

In my example in github, I only tested Writable since I pull individual fields and wrap them as Text or LongWritable for the keys.


Wednesday, July 2, 2014

Processing ASN.1 Call Detail Records with Hadoop (using Bouncy Castle) Part 3

Finally we get to the Hadoop Map/Reduce job...

We created the data and created a simple decoder to test, so now we can take the decoding logic and put it in a RecordReader.

The InputFormat we create is very simple - set isSplitable false and use our RecordReader named RawFileRecordReader.


The RecordReader does the bulk of the work.


RawFileRecordReader simply returns the filename and the count of the ASN.1 records in the file. We can change that to something more useful in a later post.

The Driver is also simple.


Get the full code on github. The code here is the simplest way to handle binary data files. There are lots of things to add for better performance. If the data files are large enough, adding in splitting logic may be worthwhile. If the data files are small, it may be worth using a map job to group them into sequence files, or convert them into avro files.

Update: Links to Part 1, Part 2, Part 3.

Monday, June 16, 2014

Processing ASN.1 Call Detail Records with Hadoop (using Bouncy Castle) Part 2

The Stand-alone Decoder

Now that we have created sample data, we can create a simple decoder with the Bouncy Castle library.


The decompressStream method is a little overkill, but will let the sample data be compressed and handle it fine. This causes a dependency on commons-compress but can also be removed easily (just change to return input).

To iterate through the ASN.1 file, we keep grabbing objects from ASN1InputStream with readObject. Once we have an object, we use it to create a CallDetailRecord instance.


Using Bouncy Castle requires some digging into the data format to get the expected set of classes. Now that the decoder is complete, we can move on to the Map/Reduce job. We didn't have to create a decoder and could have jumped straight into the Map/Reduce job, but creating a simple decoder for the first time I tackle a binary format has always saved me time.

Update: Links to Part 1, Part 2, Part 3.