Logo Questions Linux Laravel Mysql Ubuntu Git Menu
 

Implementing an Algorithm in Hadoop and java

Hi I am trying to implement a newly built bioinformatics algorithm in Hadoop and Java (I am not sure if it could be done). I have searched a lot over internet for implementing the algorithm on Hadoop. However all I find is "Identify the parallel tasks and execute them over hadoop". I would really appreciate if you guys can guide me to the right resources of Hadoop with Java over internet where I could find some solid example other than word count. I know Java well, but hadoop is my first time. Any help would be appreciated.

This is what I want to do

I have a very large text file (approx 100 MB) which have lines of characters (A,G,T,C) randomized.Long sequences of randomized A,G,T,C might form a string of important sequence k for eg (ATCGAGC). I might find this sequence k-mer in many lines of this text file called 'r'.

I have to perform following tasks

  1. Identify the position of various k-mer in all lines of text(r) in R (whole set/file)

  2. I have to keep track of positions of k-mer in a particular r.

  3. I have two parameters which are used to compare the k-mers in various r.

  4. If the k-mers in two 'r' satisfy the above parameter comparison I have to update the neighbor set N

If you are interested this is the pseudo code here it is

Given k, ĥ, ȇ
    1.  Make K  by extracting all possible kmers  from Reads
    2.  for all reads r belongs R do
            construct Gk[r] by scanning through r
            end for
    3.  for all k ε K do
                   for all read pairs (r,s) ε Gk × GK
                    if h(r,s) ≥ ĥ  and dk < ȇ h(r,s) then
                         update the N
                    end if
            end for
        end for

       k is k-mer
       K is set of all k
       ĥ minimum overlap distance
       ȇ maximum mismatch tolerance
       N neighbor set
       h(r,s)   overlap length of r and s wrt k
       d(r,s) distance between r and s
like image 897
user1772218 Avatar asked Sep 26 '26 23:09

user1772218


1 Answers

First, this problem looks like 'set similarity' problem. There a bunch of them with various effectivity on mapreduce platforms. Start looking here http://infolab.stanford.edu/~ullman/mmds.html chapter3. But only if your primary task: learn hadoop. If not...

Second, 100MB - is a very small amount of data for hadoop. Definitely, you don't need hadoop at all. Even not more then 2 parallel tasks will be launched (by default hadoop launches 1 task per 64MB). You can implement any similarity algorithm in pure java, and it will works much faster.

like image 114
octo Avatar answered Sep 28 '26 12:09

octo



Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!