Showing posts with label oozie. Show all posts
Showing posts with label oozie. Show all posts
Monday, May 8, 2017
Sunday, May 7, 2017
Monday, February 15, 2016
Oozie : Running a jobflow using java api
This is sample program to run the map-reduce job flow from oozie java client Api.copy the example directory shipped with the oozie installation on the dfs and create a sample program as follows:
import java.util.Properties;
import org.apache.oozie.client.OozieClient;
import org.apache.oozie.client.WorkflowJob;
public class OozieWFJavaApi{
public static void main(String args[]){
OozieClient wc = new OozieClient("http://ip-10-0-0-233:11000/oozie");
Properties conf = wc.createConfiguration();
conf.setProperty(OozieClient.APP_PATH, "maprfs:///user/mapr/examples/apps/map-reduce/workflow.xml");
conf.setProperty("jobTracker", "maprfs:///");
conf.setProperty("nameNode", "maprfs:///");
conf.setProperty("queueName", "default");
conf.setProperty("oozie.use.system.libpath", "true");
conf.setProperty("oozie.wf.rerun.failnodes", "true");
conf.setProperty("outputDir","map-reduce");
try {
String jobId = wc.run(conf);
System.out.println("job, " + jobId + " submitted");
while (wc.getJobInfo(jobId).getStatus() == WorkflowJob.Status.RUNNING) {
System.out.println("Workflow job in Running State");
Thread.sleep(1000);
}
System.out.println("WF job Completed");
System.out.println(wc.getJobInfo(jobId));
} catch (Exception r) {
System.out.println("Job submission failed with exception " + r.getLocalizedMessage());
}
}
}
Compile and run using oozie client api[mapr@ip-10-0-0-233 tmp]$ javac -cp .:/opt/mapr/oozie/oozie-4.2.0/lib/oozie-client-4.2.0-mapr-1510.jar OozieWFJavaApi.java [mapr@ip-10-0-0-233 tmp]$ java -cp .:/opt/mapr/oozie/oozie-4.2.0/lib/* OozieWFJavaApi job, 0000007-160215031423760-oozie-mapr-W submitted Workflow job in Running State Workflow job in Running State Workflow job in Running State Workflow job in Running State WF job Completed Workflow id[0000007-160215031423760-oozie-mapr-W] status[SUCCEEDED]
Tuesday, November 19, 2013
Apache Oozie Workflow : Configure and Running a MapReduce job
In this post I will demonstrate you how to configure the Oozie workflow. let's develop a simple MapReduce program using java, if you find any difficulties in doing it then download the code from my git location.Download
Please follow my earlier post to install and run oozie server, create a job directory say SimpleOozieMR as per following directory structure
---SimpleOozieMR
----workflow
-----lib
------workflow.xml
in the lib folder copy the you hadoop job jar and related jars.
let's configure our workflow.xml and keep it into the workflow directory as shown.
Now configure your properties file PatentCitation.properties as follows
lets create a shell script which will run your first oozie job:
Please follow my earlier post to install and run oozie server, create a job directory say SimpleOozieMR as per following directory structure
---SimpleOozieMR
----workflow
-----lib
------workflow.xml
in the lib folder copy the you hadoop job jar and related jars.
let's configure our workflow.xml and keep it into the workflow directory as shown.
<workflow-app name="WorkFlowPatentCitation" xmlns="uri:oozie:workflow:0.1">
<start to="JavaMR-Job"/>
<action name="JavaMR-Job">
<java>
<job-tracker>${jobTracker}</job-tracker>
<name-node>${nameNode}</name-node>
<prepare>
<delete path="${outputDir}"/>
</prepare>
<configuration>
<name>mapred.queue.name</name>
<value>default</value>
</configuration>
<main-class>com.rjkrsinghhadoop.App</main-class>
<arg>${citationIn}</arg>
<arg>${citationOut}</arg>
</java>
<ok to="end"/>
<error to="fail"/>
</action>
<kill name="fail">
<message>"Killed job due to error: ${wf:errorMessage(wf:lastErrorNode())}"</message>
</kill>
<end name="end" />
</workflow-app>
Now configure your properties file PatentCitation.properties as follows
nameNode=hdfs://master:8020 jobTracker=master:8021 queueName=default citationIn=citationIn-hdfs citationOut=citationOut-hdfs oozie.wf.application.path=$(namenode)/user/rks/oozieworkdir/SimpleOozieMR/workflow
lets create a shell script which will run your first oozie job:
#!/bin/sh # export OOZIE_URL="http://localhost:11000/oozie" #copy your input data to the hdfs hadoop fs -copyFromLocal /home/rks/CitationInput.txt citationIn-hdfs #copy SimpleOozieMR to hdfs hadoop fs -put /home/rks/SimpleOozieMR SimpleOozieMR #running the oozie job cd /usr/lib/oozie/bin/ oozie job -config /home/rks/SimpleOozieMR/PatentCitation.properties -run
Apache oozie : Getting Started
Apache oozie Introduction:
--- Started by Yahoo, currenly managed by Apache open source project.
--- Oozie is a workflow scheduler system to manage Apache Hadoop jobs.
-- MapReduce
-- Pig,Hive
-- Streaming
-- Standard Applications
--- Oozie is a scalable, reliable and extensible system.
--- User specifies action flow as Directed Acyclic Graph (DAG)
--- DAG: is a collection of vertices and directed edge configured so that one may not traverse the same vertex twice
--- Each node signifies eighter a Job or Script,Execution and branching can be parameterized by time, decision, data availability,
file size etc.
--- Client specifies process flow in webflow XML
--- Oozie is an extra level of abstraction between user and Hadoop
--- Oozie has its own server application which talks to it's own database(Apache Derby(default),MySql,Oracle etc.
--- User must load required component into the HDFS prior to the execution like input data, flow XML,JARs, resource files.

Interaction with Oozie through command line
Web Interface

Installation
-Download Oozie from the Apache oozie official site
-Download ExtJS
-Configure core-site.xml
-restart namenode
-Copy Hadoop jars into a directory
-Extract ExtJS into Oozie's webapp
-Run oozie-setup.sh
-Relocalt newly generated war file
-Configure oozie-site.xml
-Initialize the databse
-Start oozie server
it's done, in the next course of action we will run MapReduce job configured using. stay tuned
--- Started by Yahoo, currenly managed by Apache open source project.
--- Oozie is a workflow scheduler system to manage Apache Hadoop jobs.
-- MapReduce
-- Pig,Hive
-- Streaming
-- Standard Applications
--- Oozie is a scalable, reliable and extensible system.
--- User specifies action flow as Directed Acyclic Graph (DAG)
--- DAG: is a collection of vertices and directed edge configured so that one may not traverse the same vertex twice
--- Each node signifies eighter a Job or Script,Execution and branching can be parameterized by time, decision, data availability,
file size etc.
--- Client specifies process flow in webflow XML
--- Oozie is an extra level of abstraction between user and Hadoop
--- Oozie has its own server application which talks to it's own database(Apache Derby(default),MySql,Oracle etc.
--- User must load required component into the HDFS prior to the execution like input data, flow XML,JARs, resource files.

Interaction with Oozie through command line
$oozie job --oozie http://localhost:11000/oozie -config /user/rks/spjob/job.properties -run
Web Interface

Installation
-Download Oozie from the Apache oozie official site
-Download ExtJS
-Configure core-site.xml
-restart namenode
-Copy Hadoop jars into a directory
-Extract ExtJS into Oozie's webapp
-Run oozie-setup.sh
-Relocalt newly generated war file
-Configure oozie-site.xml
-Initialize the databse
-Start oozie server
it's done, in the next course of action we will run MapReduce job configured using. stay tuned
Subscribe to:
Posts (Atom)