‏הצגת רשומות עם תוויות HDFS. הצג את כל הרשומות
‏הצגת רשומות עם תוויות HDFS. הצג את כל הרשומות

יום שלישי, 11 ביוני 2013

Executing the hadoop samples

I keep on follow the Single node setup document.
I call to copy from the local system to the HDFS th config folder
bin/hadoop fs -put conf input
The result can be found using the
NameNode page - http://localhost:50070/ 
The files are list in: /user/zvika/input directory
Note that each file Block Size is 64MB the user is zvika and the Group is supergroup

I execute the samples using :
 zvika@ubuntu:~/myStaff/Hadoop/hadoop-1.1.2$ bin/hadoop jar hadoop-examples-*.jar grep input output 'dfs[a-z.]+'

The result is a very long list :

13/06/11 21:51:08 INFO util.NativeCodeLoader: Loaded the native-hadoop library
13/06/11 21:51:08 WARN snappy.LoadSnappy: Snappy native library not loaded
13/06/11 21:51:08 INFO mapred.FileInputFormat: Total input paths to process : 20
13/06/11 21:51:09 INFO mapred.JobClient: Running job: job_201306112139_0001
13/06/11 21:51:10 INFO mapred.JobClient:  map 0% reduce 0%
13/06/11 21:51:14 INFO mapred.JobClient:  map 10% reduce 0%
13/06/11 21:51:16 INFO mapred.JobClient:  map 15% reduce 0%
13/06/11 21:51:17 INFO mapred.JobClient:  map 20% reduce 0%
13/06/11 21:51:18 INFO mapred.JobClient:  map 25% reduce 0%
13/06/11 21:51:19 INFO mapred.JobClient:  map 30% reduce 0%
13/06/11 21:51:20 INFO mapred.JobClient:  map 40% reduce 0%
13/06/11 21:51:21 INFO mapred.JobClient:  map 45% reduce 0%
13/06/11 21:51:22 INFO mapred.JobClient:  map 50% reduce 0%
13/06/11 21:51:23 INFO mapred.JobClient:  map 55% reduce 13%
13/06/11 21:51:24 INFO mapred.JobClient:  map 60% reduce 13%
13/06/11 21:51:25 INFO mapred.JobClient:  map 65% reduce 13%
13/06/11 21:51:26 INFO mapred.JobClient:  map 70% reduce 13%
13/06/11 21:51:27 INFO mapred.JobClient:  map 80% reduce 13%
13/06/11 21:51:28 INFO mapred.JobClient:  map 85% reduce 13%
13/06/11 21:51:29 INFO mapred.JobClient:  map 90% reduce 13%
13/06/11 21:51:30 INFO mapred.JobClient:  map 100% reduce 13%
13/06/11 21:51:32 INFO mapred.JobClient:  map 100% reduce 23%
13/06/11 21:51:34 INFO mapred.JobClient:  map 100% reduce 100%
13/06/11 21:51:34 INFO mapred.JobClient: Job complete: job_201306112139_0001
13/06/11 21:51:34 INFO mapred.JobClient: Counters: 30
13/06/11 21:51:34 INFO mapred.JobClient:   Job Counters
13/06/11 21:51:34 INFO mapred.JobClient:     Launched reduce tasks=1
13/06/11 21:51:34 INFO mapred.JobClient:     SLOTS_MILLIS_MAPS=33979
13/06/11 21:51:34 INFO mapred.JobClient:     Total time spent by all reduces waiting after reserving slots (ms)=0
13/06/11 21:51:34 INFO mapred.JobClient:     Total time spent by all maps waiting after reserving slots (ms)=0
13/06/11 21:51:34 INFO mapred.JobClient:     Launched map tasks=20
13/06/11 21:51:34 INFO mapred.JobClient:     Data-local map tasks=20
13/06/11 21:51:34 INFO mapred.JobClient:     SLOTS_MILLIS_REDUCES=19571
13/06/11 21:51:34 INFO mapred.JobClient:   File Input Format Counters
13/06/11 21:51:34 INFO mapred.JobClient:     Bytes Read=29676
13/06/11 21:51:34 INFO mapred.JobClient:   File Output Format Counters
13/06/11 21:51:34 INFO mapred.JobClient:     Bytes Written=180
13/06/11 21:51:34 INFO mapred.JobClient:   FileSystemCounters
13/06/11 21:51:34 INFO mapred.JobClient:     FILE_BYTES_READ=82
13/06/11 21:51:34 INFO mapred.JobClient:     HDFS_BYTES_READ=31840
13/06/11 21:51:34 INFO mapred.JobClient:     FILE_BYTES_WRITTEN=1081971
13/06/11 21:51:34 INFO mapred.JobClient:     HDFS_BYTES_WRITTEN=180
13/06/11 21:51:34 INFO mapred.JobClient:   Map-Reduce Framework
13/06/11 21:51:34 INFO mapred.JobClient:     Map output materialized bytes=196
13/06/11 21:51:34 INFO mapred.JobClient:     Map input records=843
13/06/11 21:51:34 INFO mapred.JobClient:     Reduce shuffle bytes=196
13/06/11 21:51:34 INFO mapred.JobClient:     Spilled Records=6
13/06/11 21:51:34 INFO mapred.JobClient:     Map output bytes=70
13/06/11 21:51:34 INFO mapred.JobClient:     Total committed heap usage (bytes)=3346661376
13/06/11 21:51:34 INFO mapred.JobClient:     CPU time spent (ms)=5060
13/06/11 21:51:34 INFO mapred.JobClient:     Map input bytes=29676
13/06/11 21:51:34 INFO mapred.JobClient:     SPLIT_RAW_BYTES=2164
13/06/11 21:51:34 INFO mapred.JobClient:     Combine input records=3
13/06/11 21:51:34 INFO mapred.JobClient:     Reduce input records=3
13/06/11 21:51:34 INFO mapred.JobClient:     Reduce input groups=3
13/06/11 21:51:34 INFO mapred.JobClient:     Combine output records=3
13/06/11 21:51:34 INFO mapred.JobClient:     Physical memory (bytes) snapshot=4006256640
13/06/11 21:51:34 INFO mapred.JobClient:     Reduce output records=3
13/06/11 21:51:34 INFO mapred.JobClient:     Virtual memory (bytes) snapshot=22421676032
13/06/11 21:51:34 INFO mapred.JobClient:     Map output records=3
13/06/11 21:51:34 INFO mapred.FileInputFormat: Total input paths to process : 1
13/06/11 21:51:34 INFO mapred.JobClient: Running job: job_201306112139_0002
13/06/11 21:51:35 INFO mapred.JobClient:  map 0% reduce 0%
13/06/11 21:51:38 INFO mapred.JobClient:  map 100% reduce 0%
13/06/11 21:51:45 INFO mapred.JobClient:  map 100% reduce 33%
13/06/11 21:51:47 INFO mapred.JobClient:  map 100% reduce 100%
13/06/11 21:51:47 INFO mapred.JobClient: Job complete: job_201306112139_0002
13/06/11 21:51:47 INFO mapred.JobClient: Counters: 30
13/06/11 21:51:47 INFO mapred.JobClient:   Job Counters
13/06/11 21:51:47 INFO mapred.JobClient:     Launched reduce tasks=1
13/06/11 21:51:47 INFO mapred.JobClient:     SLOTS_MILLIS_MAPS=3134
13/06/11 21:51:47 INFO mapred.JobClient:     Total time spent by all reduces waiting after reserving slots (ms)=0
13/06/11 21:51:47 INFO mapred.JobClient:     Total time spent by all maps waiting after reserving slots (ms)=0
13/06/11 21:51:47 INFO mapred.JobClient:     Launched map tasks=1
13/06/11 21:51:47 INFO mapred.JobClient:     Data-local map tasks=1
13/06/11 21:51:47 INFO mapred.JobClient:     SLOTS_MILLIS_REDUCES=8441
13/06/11 21:51:47 INFO mapred.JobClient:   File Input Format Counters
13/06/11 21:51:47 INFO mapred.JobClient:     Bytes Read=180
13/06/11 21:51:47 INFO mapred.JobClient:   File Output Format Counters
13/06/11 21:51:47 INFO mapred.JobClient:     Bytes Written=52
13/06/11 21:51:47 INFO mapred.JobClient:   FileSystemCounters
13/06/11 21:51:47 INFO mapred.JobClient:     FILE_BYTES_READ=82
13/06/11 21:51:47 INFO mapred.JobClient:     HDFS_BYTES_READ=297
13/06/11 21:51:47 INFO mapred.JobClient:     FILE_BYTES_WRITTEN=101471
13/06/11 21:51:47 INFO mapred.JobClient:     HDFS_BYTES_WRITTEN=52
13/06/11 21:51:47 INFO mapred.JobClient:   Map-Reduce Framework
13/06/11 21:51:47 INFO mapred.JobClient:     Map output materialized bytes=82
13/06/11 21:51:47 INFO mapred.JobClient:     Map input records=3
13/06/11 21:51:47 INFO mapred.JobClient:     Reduce shuffle bytes=82
13/06/11 21:51:47 INFO mapred.JobClient:     Spilled Records=6
13/06/11 21:51:47 INFO mapred.JobClient:     Map output bytes=70
13/06/11 21:51:47 INFO mapred.JobClient:     Total committed heap usage (bytes)=220528640
13/06/11 21:51:47 INFO mapred.JobClient:     CPU time spent (ms)=790
13/06/11 21:51:47 INFO mapred.JobClient:     Map input bytes=94
13/06/11 21:51:47 INFO mapred.JobClient:     SPLIT_RAW_BYTES=117
13/06/11 21:51:47 INFO mapred.JobClient:     Combine input records=0
13/06/11 21:51:47 INFO mapred.JobClient:     Reduce input records=3
13/06/11 21:51:47 INFO mapred.JobClient:     Reduce input groups=1
13/06/11 21:51:47 INFO mapred.JobClient:     Combine output records=0
13/06/11 21:51:47 INFO mapred.JobClient:     Physical memory (bytes) snapshot=296771584
13/06/11 21:51:47 INFO mapred.JobClient:     Reduce output records=3
13/06/11 21:51:47 INFO mapred.JobClient:     Virtual memory (bytes) snapshot=2142064640
13/06/11 21:51:47 INFO mapred.JobClient:     Map output records=3


To examine the hadoop job processing go to http://ubuntu:50060/tasktracker.jsp and refresh while executing  the JOB

יום ראשון, 9 ביוני 2013

temp HDFS nodename

In the single node setup a temporary HDFS is created and located in:
 tmp/hadoop-zvika
The following command
bin/hadoop namenode –format
Should be execute on every machine restart in order to create the tmp HDFS  
Only after executing the command  there is an access to nodename: http://localhost:50070/

HDFS
commands :
~/myStaff/Hadoop/hadoop-1.1.2/bin$ ./hadoop fs –help
Create directory
./hadoop fs -mkdir /user/hadoop/dir1
List the created directory
zvika@ubuntu:~/myStaff/Hadoop/hadoop-1.1.2/bin$ ./hadoop fs -ls /user/hadoop
Found 1 items
drwxr-xr-x   - zvika supergroup          0 2013-06-09 22:26 /user/hadoop/dir1

יום רביעי, 29 במאי 2013

Hbase

A great hello world tutorial explaining about how to start with Hbase and Hadoop can be found Here.
This is my summery and notes about the post :
Installing the SSH server:
sudo apt-get install openssh-server

Create the Hadoop user:
sudo addgroup hadoop
sudo adduser --ingroup hadoop huser

Generate the user public keys:
#login as hadoop user
sudo -i -u huser 
#Create the hadoop user public key
ssh-keygen -t dsa -P '' -f ~/.ssh/id_dsa
#Copy the generated public key onto the ssh/authorized_keys
cat ~/.ssh/id_dsa.pub >> ~/.ssh/authorized_keys 

Setting Up HDFS
#Create a directory used to contain the HDFS file
mkdir /home/huser/my_hdfs_folder

Update the hadoop config with the hdfs directrory
Note:hadoop.tmp.dir is used as the base for temporary directories locally, and also in HDFS.
The following configuration set the created directory as the HDFS directory.

  1: <?xml version=”1.0”?>
  2:  <?xml-stylesheet type=”text/xsl” href=”configuration.xsl”?>
  3:  <configuration>
  4:  <property>
  5:  <name>hadoop.tmp.dir</name>
  6:  <value>/home/huser/my_hdfs_folder</value>
  7:  </property>
  8:  <property>
  9:  <name>fs.default.name</name>
 10:  <value>hdfs://ubuntu:8020</value>
 11:  </property>
 12:  </configuration>

#Format the HDFS
/usr/local/hadoop/bin/hadoop namenode –format
#start the hadoop single instance
/usr/local/hadoop/bin/start-all.sh

View the lifeness of the hdoop instance in the following url:
http://ubuntu:50070/dfshealth.jsp


Setting up HBase
HBase need a directory inside of the HDFS
We create it using the HDFS fs –mkdir command for example
/usr/local/hadoop/bin/hadoop fs -mkdir myHbase
The new hdfs directoy should be point out in the Hbase site configuration file :
hbase-site.xml.

  1: configuration> 
  2:  <property> 
  3:  <name>hbase.rootdir</name> 
  4:  <value>hdfs://ubuntu:8020/user/huser/myHbase</value> 
  5:  <description> 
  6:  </description> 
  7:  </property> 
  8:  <property> 
  9:  <name>hbase.master</name> 
 10:  <value>ubuntu:60000</value> 
 11:  <description> 
 12:  </description> 
 13:  </property> 
 14:  </configuration>

Start the HBase DB
/usr/local/hbase/bin/start-hbase.sh
Monitor its lifeness
http://ubuntu:60010/master-status
Starting the Shell
/usr/local/hbase/bin/hbase shell

Create and Update a simple DB
#Create a new table named myBlogs along with a column family BlogText 
create ‘myBlogs','BlogText'
#insert some data

  1: put ‘myBlogs','Ruby','BlogText:1','About ruby bla bla.'
  2: 
  3: put ‘myBlogs','Ruby','BlogText:2','about {|X| bla bal.'
  4: 
  5: put ‘myBlogs','Ruby','BlogText:3','for loops.'
  6: 
  7: put ‘myBlogs','Python','BlogText:1','iter tools .'

The following code is used to query the created  hbase DB

  1: package my.learn.hbase;
  2: 
  3: import java.util.NavigableMap;
  4: import java.util.NavigableSet;
  5: 
  6: import org.apache.hadoop.conf.Configuration;
  7: import org.apache.hadoop.hbase.HBaseConfiguration;
  8: import org.apache.hadoop.hbase.client.HBaseAdmin;
  9: import org.apache.hadoop.hbase.client.HTableFactory;
 10: import org.apache.hadoop.hbase.client.HTableInterface;
 11: import org.apache.hadoop.hbase.client.Result;
 12: import org.apache.hadoop.hbase.client.ResultScanner;
 13: import org.apache.hadoop.hbase.client.Scan;
 14: import org.apache.hadoop.hbase.util.Bytes;
 15: 
 16: public class HBaseReadMyBlogsData  {
 17: 
 18:   public static final byte[] TablemyBlogs = Bytes.toBytes("myBlogs");
 19:   // The column family
 20:   public static final byte[] BlogText_FAMILY = Bytes.toBytes("BlogText");
 21: 
 22:   
 23:   private void ShowTheBlogsText() throws Exception {
 24: 
 25:     // Load's the hbase-site.xml config
 26:     Configuration config = HBaseConfiguration.create();
 27:     //Factory for creating HTable instances.
 28:     HTableFactory factory = new HTableFactory();
 29:     
 30:     HBaseAdmin.checkHBaseAvailable(config);
 31: 
 32:     // Link to table
 33:     HTableInterface table = factory.createHTableInterface(config,
 34:         TablemyBlogs);
 35: 
 36:     // Used to retrieve rows from the table
 37:     Scan scan = new Scan();
 38: 
 39:     // Scan through each row in the table
 40:     ResultScanner rs = table.getScanner(scan);
 41:     try {
 42:       // Loop through each retrieved row
 43:       for (Result r = rs.next(); r != null; r = rs.next()) {
 44:         //print out the row key
 45:         System.out.println("Key: " + new String(r.getRow()));
 46:       
 47:         //For each key loop over its qualifier for "ruby" key we will have 1 , 2 , 3 
 48:         
 49:         NavigableMap familyMap = r
 50:             .getFamilyMap(BlogText_FAMILY);
 51:         // This is a list of the qualifier keys
 52:         NavigableSet keySet = familyMap.navigableKeySet();
 53: 
 54:         // Print out each value within each qualifier
 55:         for (byte[] key : keySet) {
 56:           System.out.println("\t Definition: " + (new String(key))
 57:               + ", Value:"
 58:               + new String(r.getValue(BlogText_FAMILY, key)));
 59:         }
 60:       }
 61:     } catch (Exception e) {
 62:       throw e;
 63:     } finally {
 64:       rs.close();
 65:     }
 66: 
 67:   }
 68: }

Notes:
The HBaseAdmin provides an interface to manage HBase database table metadata + general administrative functions. like create, drop, list, enable and disable tables.
The HBaseAdmin can be used to add and drop table column families.

יום שבת, 25 במאי 2013

Accessing HDFS

Copying file from and to HDFS

In order to copy a file from the local file system to HDFS we can use the FS  command: copyFromLocal

For an example:

  1: % hadoop fs -copyFromLocal LabsVideos/Videos/VideoSample164.mpeg hdfs://VideosDataNode1/user/Lab1/VideoSample164.mpeg

In order  to copy a file from the HDFS to local directory we can use the FS command :copyToLocal

For an example :


  1: % hadoop fs -copyToLocal LabsVideos/BenchMarksResutls/Results080513.xls hdsf://DataResultsNode/user/Lab1/Results080513.xls

Accessing HDFS  by code

The HDFS can be accessed an manipulate by several way :
1.Java interface 
The new api  :abstract file system and  File context  allow easy interface to access files  in all notes in the cluster .
The main class  in the new api is  the  context   class 
An example  of the new api usage :

  1: FileContext myFContext = FileContext.getFileContext(); // uses the default config the default FS
  2: //Set the working directory to other data node 
  3: myFContext.setWorkingDir("hdsf://NotLocalDataResultsNode/user/Lab1/VideosList");
  4: //Opens an FSDataInputStream at the indicated Path.
  5:  FSDataInputStream theFSDataInputStream = myFContext.open ("VideoSample2");
  6:  //Read from the FSDataInputStream
  7:  byte[] Databuffer = new byte[1000];
  8:  theFSDataInputStream.read (0 ,Databuffer, 0 , 1000);
  9:  
 10:  //Use the theFSDataInputStream
 11:  
 12:  theFSDataInputStream.close(); 
 13:  
 14: //Create new directory 
 15: myFContext.create("NewVideoSamplesDir");


The namespace for the  FileContext  interface  is:  org.apache.hadoop.fs  .Documentation  about the FileContext interface can be found here.



2.Accessing HDFS using C  LibHDFS library . Documentation can be found  here


3.Accessing HDFS  using http calls .