I needed to test certain scenarios for a client against a Microsoft Active Directory Domain Controller and Intermediate Certificate Authority. The easiest way was to use Vagrant with the mwrock/Windows2012R2 box.
I wasn't able to automate the complete install, but did get it to a set of cut-and-paste lines.
Code is at [ https://github.com/awcoleman/vagrant_win_ad_dc_ca_test ]
Copy Vagrantfile into new directory
Follow directions in README.txt
The next iteration will probably use Ansible support for Windows (unfortunately there is no CA module)
Saturday, February 24, 2018
Monday, July 3, 2017
Giraph Error: Could not find or load main class org.apache.giraph.yarn.GiraphApplicationMaster
Very old post from 2014 that got lost in my drafts. Posting so hopefully this helps out someone.
Often Google acts like magic for me: type in my error, and out pops the solution. Not so for a Giraph error I recently hit. Hopefully this post lets Google work like magic for someone else :)
After installing Giraph on a BigTop 0.7 VM, I was able to run the benchmark that takes no input or output but nothing more complicated.
This works:
hadoop jar /usr/share/doc/giraph-1.0.0.5/giraph-examples-1.0.0-for-hadoop-2.0.6-alpha-jar-with-dependencies.jar org.apache.giraph.benchmark.PageRankBenchmark -Dgiraph.zkList=127.0.0.1:2181 -libjars /usr/lib/giraph/giraph-1.0.0-for-hadoop-2.0.6-alpha-jar-with-dependencies.jar -e 1 -s 3 -v -V 50 -w 1
But this:
hadoop jar /usr/share/doc/giraph-1.0.0.5/giraph-examples-1.0.0-for-hadoop-2.0.6-alpha-jar-with-dependencies.jar org.apache.giraph.GiraphRunner -Dgiraph.zkList=127.0.0.1:2181 -libjars /usr/lib/giraph/giraph.jar org.apache.giraph.examples.SimpleShortestPathsVertex -vif org.apache.giraph.io.formats.JsonLongDoubleFloatDoubleVertexInputFormat -vip /user/acoleman/giraphtest/tiny_graph.txt -of org.apache.giraph.io.formats.IdWithValueTextOutputFormat -op /user/acoleman/giraphtest/shortestpathsC2 -ca SimpleShortestPathsVertex.source=2 -w 1
Often Google acts like magic for me: type in my error, and out pops the solution. Not so for a Giraph error I recently hit. Hopefully this post lets Google work like magic for someone else :)
After installing Giraph on a BigTop 0.7 VM, I was able to run the benchmark that takes no input or output but nothing more complicated.
This works:
hadoop jar /usr/share/doc/giraph-1.0.0.5/giraph-examples-1.0.0-for-hadoop-2.0.6-alpha-jar-with-dependencies.jar org.apache.giraph.benchmark.PageRankBenchmark -Dgiraph.zkList=127.0.0.1:2181 -libjars /usr/lib/giraph/giraph-1.0.0-for-hadoop-2.0.6-alpha-jar-with-dependencies.jar -e 1 -s 3 -v -V 50 -w 1
But this:
hadoop jar /usr/share/doc/giraph-1.0.0.5/giraph-examples-1.0.0-for-hadoop-2.0.6-alpha-jar-with-dependencies.jar org.apache.giraph.GiraphRunner -Dgiraph.zkList=127.0.0.1:2181 -libjars /usr/lib/giraph/giraph.jar org.apache.giraph.examples.SimpleShortestPathsVertex -vif org.apache.giraph.io.formats.JsonLongDoubleFloatDoubleVertexInputFormat -vip /user/acoleman/giraphtest/tiny_graph.txt -of org.apache.giraph.io.formats.IdWithValueTextOutputFormat -op /user/acoleman/giraphtest/shortestpathsC2 -ca SimpleShortestPathsVertex.source=2 -w 1
does not.
Looking at the latest container logs with:
cat $(ls -1rtd $(ls -1rtd /var/log/hadoop-yarn/containers/application_* | tail -1)/container_* | tail -1)/*
I find:
Error: Could not find or load main class org.apache.giraph.yarn.GiraphApplicationMaster
I beat my head against the wall trying to add to libjars, to -yj, copying jars into every directory i could find.
I stumbled across
http://mail-archives.apache.org/mod_mbox/giraph-user/201312.mbox/%3C198091226.KO6f1kuK42@chronos7%3E
which gives the answer. If https://issues.apache.org/jira/browse/GIRAPH-814 hasn't been applied, then mapreduce.application.classpath has to be hard set or Giraph simply won't work.
vi /etc/hadoop/conf.pseudo/mapred-site.xml
<property>
<name>mapreduce.application.classpath</name>
<value>/usr/lib/hadoop-mapreduce/*,/usr/lib/hadoop-mapreduce/lib/*,/usr/lib/giraph/giraph-1.0.0-for-hadoop-2.0.6-alpha-jar-with-dependencies.jar
</value>
</property>
I did not need to restart yarn-resourcemanager or yarn-nodemanager for this to get picked up.
DropWizard and Hive (and/or Impala)
I have a small DropWizard/D3.js/jqGrid application to visualize the results of some analysis. I had been taking the results of the analysis from hdfs and shoveling it into mySQL (with sqoop) to examine samples. This is working well enough that I wanted to go straight to the source. With DropWizard this should be easy enough to wrap my data in a Hive external table and use the Hive JDBC driver instead of mySQL.
To pull in Hive JDBC and its dependencies, add to pom.xml:
If you are already familiar with DropWizard and just need an example, examine the pom.xml and config-hive.yaml files in my example application on GitHub.
To pull in Hive JDBC and its dependencies, add to pom.xml:
<dependency>
<groupid>org.apache.hive</groupid>
<artifactid>hive-jdbc</artifactid>
<version>1.1.0</version>
<exclusions>
<exclusion>
<groupid>org.slf4j</groupid>
<artifactid>slf4j-log4j12</artifactid>
</exclusion>
<exclusion>
<groupid>com.sun.jersey</groupid>
<artifactid>*</artifactid>
</exclusion>
</exclusions>
</dependency>
<dependency>
<groupid>org.apache.hadoop</groupid>
<artifactid>hadoop-common</artifactid>
<version>2.6.0</version>
<exclusions>
<exclusion>
<groupid>org.slf4j</groupid>
<artifactid>slf4j-log4j12</artifactid>
</exclusion>
<exclusion>
<groupid>com.sun.jersey</groupid>
<artifactid>*</artifactid>
</exclusion>
</exclusions>
</dependency>
Sunday, June 18, 2017
Pinebook!
Not Hadoop-related, but awesome all the same. A few months ago I stumbled on PINE64's website and saw the pinebook, a linux arm64 laptop. That and a PocketCHIP made a great late-birthday, early-father's-day set of presents.
Build and shipping takes a couple months, shipping was almost 1/3 of the laptop cost, and performance and keyboard quality is exactly what you would expect :) But it is still a fun bit of hardware.
If you decide to get one, make sure to add on a USB-to-H-barrel power cord (or make your own). The pinebook does come with a power supply, but no point in carting around yet another wall-wart when the pinebook happily charges off a phone charger.
Mine powered right up into Xenial. I'm normally RH-based since everywhere I've been employed in the last couple decades has been, so it's nice to jump back into Debian-based.
aarch64 wasn't in mainline rust, but was in nursery, so
curl -sSf https://raw.githubusercontent.com/rust-lang-nursery/rustup.rs/master/rustup-init.sh | bash
Build and shipping takes a couple months, shipping was almost 1/3 of the laptop cost, and performance and keyboard quality is exactly what you would expect :) But it is still a fun bit of hardware.
If you decide to get one, make sure to add on a USB-to-H-barrel power cord (or make your own). The pinebook does come with a power supply, but no point in carting around yet another wall-wart when the pinebook happily charges off a phone charger.
Mine powered right up into Xenial. I'm normally RH-based since everywhere I've been employed in the last couple decades has been, so it's nice to jump back into Debian-based.
aarch64 wasn't in mainline rust, but was in nursery, so
curl -sSf https://raw.githubusercontent.com/rust-lang-nursery/rustup.rs/master/rustup-init.sh | bash
worked just fine and got me up and going with rust.
Update: HackADay has a great write-up. I didn't experience any of the screen issues they had since I have the 14", but the page has a great tear-down and overview of performance (which is not much :) )
Thursday, February 23, 2017
Writing ORC files is easier than a few years ago
Several years ago I was asked to compare writing Parquet and ORCFile formats from standalone java (without using the Hadoop libraries). At the time ORC was not separated from Hive and it was much more involved than writing Parquet from java. It looks like that changed in 2015 but I only revisited the issue within the past few months.
To build ORC:
Download the current release (currently 1.3.2)
tar xzvf orc-1.3.2.tar.gz && cd ./orc-1.3.2/
cd ./java
And a simple example of reading is:
To build ORC:
Download the current release (currently 1.3.2)
tar xzvf orc-1.3.2.tar.gz && cd ./orc-1.3.2/
cd ./java
mvn package
ls -la ./tools/target/orc-tools-1.3.2-uber.jar
A simple example of writing is:
A simple example of writing is:
And a simple example of reading is:
Tuesday, January 6, 2015
Github PrintGZHeader
I receive some data that is a gzip file made of concatenated smaller gzip files. The larger file is a valid gzip according to the rfc. Everything works fine unless I need to look at the original filename or mtime of the gz streams. So... I created a program to print out the original filename and mtime from the header information in all the gzip streams in a gz file.
https://github.com/awcoleman/PrintGZHeader
Compile: gcc -o printGZHeader printGZHeader.c -lz
And run: ./printGZHeader myGZfile.gz
My C is very rusty, I will happily accept any patches to clean up bad practices. My test files do not have header comments or extra fields, please send patches if you find that the code does not work appropriately on them (or send me a test file and I will try).
Hopefully this saves someone else a little bit of time.
https://github.com/awcoleman/PrintGZHeader
Compile: gcc -o printGZHeader printGZHeader.c -lz
And run: ./printGZHeader myGZfile.gz
My C is very rusty, I will happily accept any patches to clean up bad practices. My test files do not have header comments or extra fields, please send patches if you find that the code does not work appropriately on them (or send me a test file and I will try).
Hopefully this saves someone else a little bit of time.
Wednesday, December 31, 2014
Nostalgia in Clustering
The close of 2014 made me remember an old clustering project I did around 2004 on a shoestring budget. The project was correlating customers into families, with a sub-task of deduplicating customer records (from typos and other issues). The entire project team was… me.
I gathered up a server with a couple of old hard drives as a mySQL server and PXE boot server, and four other computers PXE-booting into linux with openMOSIX for clustering. I didn’t have budget for cases for the four slaves, so used old cookie sheets to mount them. I used wooden dowels to fix two cookie sheet nodes together so they could sit vertically.
My processing was done in perl. Once OpenMOSIX reported a slave was free, a perl process would spawn and grab a workload from mySQL. OpenMOSIX would migrate the process to the open slave.
Fortunately I was able to complete the project with only four slaves. I had figured out my power supplies could power two slaves. I was working on converting a couple of ATX power supply extension cables into a Y-splitter and only using one power supply per "cookie".
I found some pictures of the nodes from an old presentation:
Subscribe to:
Posts (Atom)


