Wednesday, June 22, 2011

Meet the Lazyrussian, a hacker that played monkey tricks with Facebook for two-and-half years

Most of you who are reading this post may not have any knowledge as to who The Lazyrussian is. Despite that, most of you may have used his software. I was the same a few months ago. I have never thought very much about Lazyrussian until recently when I visited a homepage of one of his popular technologies and found that he announced that he has ceased developing the software and will no longer host it due to pending lawsuits from Facebook.

He goes by the name Arthur Ariel Sabintsev but he uses to call himself The Lazyrussian or just Lazyrussian in software development circles. He is a twenty-five year-old 'May' baby.  He has a combination of talents. He is an naturally born software developer who holds an undergraduate degree in Biophysics and a graduate degree in Nuclear Physics. He has never done any professional course in software development but has grown a cowboy developer since he was twelve years old.

To his credit, Lazyrussian has developed the following free and open source tools: Buzz It!, Email This!, Email Yourself!, MySpacePAD and FacePAD/PhotoJacker. On 16th June 2011, he announced on his website (http://www.sabintsev.com) that he has foregone his PhD in Physics in pursuit of his software development passion. He also mentioned that he has been hired by Fueled, a New York based software development company where he is honing his programming skills as an iOS/Android developer from 20th June 2011.

Arthur came into my limelight when I was looking for a simple tool that would help me easily download photos from Facebook when I need them. I was looking for a tool that would help me get high resolution photos, because manually downloading them was not giving me the expected quality. One day, I stumbled upon an article that described five technologies that would help me do that. I tried using all of them, but unfortunately none came closer to FacePAD. I finally fell in love with FacePAD.

FacePAD was a simple Firefox extension that would expeditiously download a single photo or an entire photo album off Facebook in few minutes. A Firefox extension is a piece of software, also referred to as an add-on or a plugin, that you can install as part your Firefox Web Browser. Lazyrussian went ahead to release FacePAD on Mozilla's Addons page for public use. On 12 January 2011, Lazyrussian announced that Facebook’s lawyers wrote him stating that he should rebrand FacePAD within 48 hours, or he would be slapped with a lawsuit.  So, the developer enacted his contingency plan and rebranded FacePAD as PhotoJacker.

On a sad note, Lazyrussian announced on PhotoJacker.com on 27th January 2011 that he has ceased from developing the software and has removed from Mozilla because Facebook had sent him a Cease & Desist notice stating that PhotoJacker/FacePAD violates Section 3.2 of their Statement of Rights and Responsibilities. In the notice, Facebook also stated that the developer was opening himself to claims that he is facilitating copyright infringement via the download of the albums of others. The notice also advised him to build something using Facebook APIs rather than working around or contravening Facebook’s limitations, or, make a suggestion to Facebook for functionality he would like to see.

From his two announcements  (about name change and Cease & Desist / Facebook Takedown Notice), and from the copy of Facebook's Cease & Desist notice that I got pa Kanjedza (under the palm tree), it appears that the discussion between the two parties on FacePAD had been around for a while. On the other hand, I personally noticed that Facebook had been frequently changing links to users' photos and photo albums over the past two years. The Lazyrussian kept updating his software till he got this strong warning on 26 January 2011, and finally gave up the fight on grounds that he was a full time student and he really didn't have the time to deal with this issue.

From the word of his mouth, and indeed, from my own opinion, FacePAD was not intended to be of any evil intentions. It developed as an out-of-play alternative way for downloading photos from Facebook. Such application is technically called a hack not a crack. FacePAD was a cool tool that made use of AJAX requests, regular expressions and some built-in Firefox functionalities to download the photos. Facebook was extremely upset that he was able to create this photo-downloading software without using their API. Lazyrussian thinks he was getting punished for being clever. On 16 June 2011, Lazyrussian joined Github and decided to publish FacePAD's source code under copyleft license in an attempt to preserve it.

It was very fascinating to see how a student hacker was making Facebook's life miserable. But Facebook also started as a student hacker's invention at Havard!

There goes The Lazyrussian and his controversial FacePAD!

Monday, March 21, 2011

Indigenous Tweets: The fun side of Tweeting in Your own Language

The most frequently used languages in the world have so far pushed other minority languages down. Most of these frequently used languages happen to be those of the colonial masters. They tend to dominate day-to-day businesses as a result indigenous languages naturally hibernate.

Computer scientists and linguists are trying to revive and sustain those indigenous languages that are surviving. Currently, as a way of reviving and preserving them, we have seen a lot of localization projects sprouting in an effort to push the endangered languages into the computing world.

Chichewa page on IndeginousTweets.com
Localization efforts may not be effective enough unless communities themselves get involved in the process. In an effort to highlight the use of indigenous languages on the Internet, Prof. Kevin Scannell has developed IndigenousTweets.com which puts together statistics of “35 plus” indigenous languages that are being used on Twitter. Launched on St. Patrick's Day 2011, Indigenous Tweets uses data gathered by Scannell's web-crawling application, An Crúbadán, which identifies the details of minority languages being tweeted. According to a post on Indigenous Tweets Blog, the primary aim of indigenoustweets.com is to help build online language communities through Twitter.

Prof. Scannell hopes that the site will aid speakers of indigenous and minority languages to find each other in the vast sea of global languages like English and French that dominate Twitter. Clicking any of the language profiles on the list takes one to a page that lists tweeters in that language with other nice indicators. 

Beyond providing linguistic statistics, I feel Indigenous Tweets provides some new wave of social networking. People will find it more funny to tweet as much as possible in an effort to boot out friends and rank top on their language pages. I hope people will not be ashamed of tweeting in their indigenous languages. In the end, we will have more and more minority languages enjoying the cyber world just as the dominant languages do.

I love Indigenous Tweets from the start. I can see myself ranking low on  the Chichewa page. Now, I am thinking of switching from facebooking to tweeting so that I boot out the top tweeters in my language. Lol!

I hope that you too will enjoy tweeting in your language more than ever!

Thursday, January 06, 2011

Natural Language Processing Tools for Chichewa

Natural language processing (NLP) is an excellent discipline of computer science focusing on developing artificial intelligence systems that are able to interact with human beings in their natural languages. The expert systems so developed try to understand patterns of human languages and process the given data (text or speech) accordingly. The expert tools are very useful in various systems. For example some feature/smart phones are able to read the name of the caller for you as the phone rings. Other applications like word processors are able to read a big document and generate for you a summary document from it. We also have nice applications (e.g. Google Translate) that read text in one  human language and translate it into another target language. All these are products of the field of NLP.

I have been working on big projects for Chichewa, a lingua franca for Malawi (formerly it's national language), Zambia and some parts of Mozambique and Zimbabwe. These systems are still in progress of perfection, but they currently are able to do great stuff at this stage. I am sure that the end of these projects will put Chichewa somewhere as far as NLP is concerned. I would like to share with you my experiences.

ChicMorph: A Morphological Analyzer for Chichewa Verbs
    Chichewa is agglutinative in nature. One word/phrase  is a combination of several sub-words (techinically called morphemes). For example sindibweranso (Lit: I am not coming again) can be broken as follows: si(not)-ndi(I)-bwer(come)-a-nso(again). Notice that the "a" has no literal meaning. It is just a final vowel to complement bwer, the stem of that verb.

    ChicMorph takes raw Chichewa verbs, discovers and isolates the verb constituent morphemes. Some Chichewa verbs are tricky in that their roots also include subwords (morphemes) that are also morphemes on their own. For example, er is an applicaticative morpheme as in gwera (gw-er-a). But it is not  a morpheme in bwera (hence bw-er-a is incorrect, but bwer-a). I have so far improved ChicMorph to evaluate correctly verbs with roots constituting morphemes that are also prefixal or suffixal allomorphs like these ones.
     ChicPOS: Part of Speech Tagger
      From August 2010, I have been working on a Chichewa part of speech tagger and it is doing great. I am hoping to make more breakthroughs in due course. Right now, chicPOS understands all Chichewa parts of speech including punctuations:
      •  Mwana womaliza uja wa a Phiri wabwera kudzagula mchere. (Lit: That last born child to Mr. Phiri has come to buy salt.) => Mwana[NN] womaliza[JJ] uja[DEM] wa[IN] a[HON] Phiri[NNP] wabwera[VB] kudzagula[VB] mchere[NN] .[.]
      Key: DEM => Demonstrative Adjective, HON => Honorific a, IN =>Preposition, JJ => Adjective, NN => Noun, NNP => Proper Noun, POSS => Possessive Adjective, PN => Pronoun, PRP => Personal Pronoun, VB =>  Verb.

      ChicPOS is also able to identify proper nouns within a given phrase. Compare usage of "Talandira" in the following phrases:
      •  Talandira ndalama kuchokera kwa a Chikale. (Lit: We have received money from Mr. Chikale.) => Talandira[VB] ndalama[NN] kuchokera[VB] kwa[ASSOC] a[HON] Chikale[NNP] .[.]
      • Ndamuona Talandira akudutsa apa. (Lit: I have seen Talandira passing by here.) => Ndamuona[VB] Talandira[NNP] akudutsa[VB] apa[DEM] .[.]
      ChicPOS fails to identify proper nouns in some positions, especially when they begin a sentence as in Talandira akudutsa apa. (Lit: Talandira is passing by here). =>  Talandira[VB] akudutsa[VB] apa[DEM] .[.] (compare it with: Akudutsa apa Talandira. (Lit: He/She is passing by here, Talandira) => Akudutsa[VB] apa[DEM] Talandira[NNP] .[.]). Proper nouns are tricky even in "natural/daily conversations" looking at the way names(Proper nouns) are formed in Chichewa. Some proper names originate from verbs/verb phrases (e.g. Talandira => (Lit: We have received), Kalinda-kadye (Lit: It waits to eat)) while others from common nouns (Chipiriro (Patience), Ulemu (Politeness/Respect)). Notice that somehow ChicPOS is also correct  in this special case: Talandira akudutsa apa. =>  Talandira[VB] akudutsa[VB] apa[DEM] .[.] The reason is since Talandira originates from a verb , by just changing the tone of the phrase Talandira akudutsa apa. will translate to We have recieived (something) while he was passing here.. In short, I should say I am still exploring this concept of proper nouns. 

      Right now, I have six thousand Chichewa words (thanks to Prof. Kevin Scannell for compiling the initial wordlist using his An Crubádan). I am in the process of tagging them, and I will be adding some more words. A note on tags, I have tried to preserve popular tags like NN, JJ but for words that I could not find one I prioritized short forms outlined in The Syntax of Chichewa by Prof. Sam Mchombo. Otherwise, I generated my own. I am hoping to create a standardized form for Chichewa (and eventually for other Malawian languages). I am also looking at some similar work in Swahili and Nguni languages.

      AffixGen: Chichewa Verb Generator
        In due course, I also developed a "Chichewa verb generator". It automatically generates 66082 prefixes, 2870 suffixes (using CARP [Causative-Applicative-Reciprocal-Passive] and RCAP suffix combination; in RCAP the reciprocal precedes the other suffixes as in menyanitsa). The suffix extension can take up to three clitics at the moment. For each single verb root, it generates 66082 x 2870 = 189,655,340 possible verb forms. This is awesome because if you have 10 Chichewa verb roots, you are able to generate close to 2 billion Chichewa verbs!! Of course some of them may not be as sensible due to some semantical encodings behind them (compare menya and bwera => akuzimenyanitsa vs akuzibweranitsa.) I am still working on this. I would like to collect all(?????) verb roots (Ha!Ha!Ha! if I can manage) and isolate them accordingly so that such funny combinations do not occur any more, or at least the error rate is reduced drastically. Right now I have 500 verb roots and the system is able to generate 94,827,670,000 (94 Billion) verbs!!!!

        I am using AffixGen output to build plugins for Hunspell spellchecker, and I have so far created two plugins, one for Firefox and another for OpenOffice (It is available online on Openoffice.org website. Of course, the online one is not up to date yet). 

        ChiVisualize: Dynamic Visualization Tool of Chichewa Phrase Structures
          In line with ChicPOS, I am creating a visualization tool for Chichewa phrase structures. This is another great art work that I have ventured into. ChiVisualize text tagged phrases and build a syntax tree as in the following example:
          Mkango[NN] uja[DEM] ukuba[VB] mikanda[NN] yanu[POSS] (That lion is stealing your beads.)
          The syntax tree is interactive and dynamic. You can change the orientation in four directions: top, bottom, left and right. You can also emphasize on a particular level in any of the sub-trees. The system is able to "virtually" simulate the all six Chichewa phrase structures: SVO, SOV, VSO, VOS, OVS and OSV. ChiVisualize uses JavaScript InfoVis Toolkit to create these interactive visualization.

          ChiVisualize is still in its formative stages. Right now, the tagged text is processed into a base phrase structure (as defined by Chomsky's Minimalist Theory/Program) manually and given to ChiVisualize for syntax tree generation. Currently, I am working on an algorithm that will be able to automatically generate a Base Phrase Structure for given tagged text.

          Later on, I will combine ChicPOS, ChicMorph and ChiVisualize into one application. With the new system, one will just be giving it a "normal" Chichewa Phrase and it will be doing all the processing itself. ChicPOS will be generating tagged text and give it to ChiVusualize for visualization. On the other hand, ChicMorph will produce extra morphological constraints that will be displayed when one emphasizes on a certain phrase constituent in a given syntax tree. I will also add a transformational-generative grammar parser that will be able to resolve matching of argument markers to their respective nominals if present in a given phrase. Of course, I am aware of ambiguities resulting from free word ordering and NPs from same classes as depicted in the following: Galimoto ng'ombe yayigunda (galimoto => car, ng'ombe => cow, yayigunda => 'has hit'). (which one hit the other here? FYI: galimoto and ng'ombe fall in the same noun class). One way will be to leave it strictly non-configurational such that the phrase will be illustrated as S = NP + NP + V (or any of its combinations, without a VP). But I'll cross the bridges when I'll come to them :-), :-).

          By the end of everything, I would like to build a head-driven phrase structure grammar checker for Chichewa. This will be useful not only in linguistics, but also in real-world applications like word processing software.

          Thursday, August 12, 2010

          Printing HTML Documents Using Customised CSS and JavaScript

          Recently, I was working on some Ruby on Rails project where users wanted to be able to print nice reports. I should admit that the gem solutions I found that time disappointed me because:
          • Most required a lot of time to understand than I had to make them work.
          • Some solutions required extra gems (and plugins) that I could not get because of usual gem error: ERROR: could not find gem XXX locally or in a repository!
          One interesting fact was that the users had a defined report format which they wanted both screen layouts and printouts to be modelled from! An additional challenge was that the application was developed for use on a mouseless touchscreen computer that always open the browser in fullscreen mode. So, it would be difficult to manually select the Print command on the File menu of the browser. So after making a some Internet research, I got several solutions that I integrated to come up with mine that addresses the users' need. I present the solution here in a tutorial format so that it is easy to follow.
          1. Document Layout
          Let us create a simple-and-easy-to-follow report layout using HTML tags as follows:

          <html>
              <head>
              <title>The Application Name</title>
              <script src="/javascripts/report_printer.js" type="text/javascript"> </script>
              <link href="/stylesheets/report.css" media="screen" rel="stylesheet" type="text/css" />
              </head>
              <body>
              <div id="report">
                  <div id="metaData" onclick="javascript:printContent('document');" style="float: right;">
                      <a> Print Report</a>
                  </div>
                  <div id="document">
                      <div id="reportHeader"> Organization Name </div>
                      <div id="reportSubHeader"> Report Name and/or Description </div>
                      <div id="dataTable">
                          <!-- the data table is placed here -->
                      </div>
                  </div>
              </div>
          </body>
          </html>
          Linked to this file is a standalone CSS file, report.css, that contains CSS definitions for each of the div ids and classes. In this example, the CSS file is placed in stylesheets sub-directory. The metaData div contains a hyper-linked button for printing the report. The metaData div customizes the hyper-link to a button. This div is significant to the whole printing process as we will see in a moment.
          1. Adding Printing Functionality using JavaScript
          With the CSS, everything looks nice on the screen. However, there are two problems:
          1. The Print command on the File menu produces clumsy output.
          2. It would be difficult to manually select the Print command on the File menu of the browser since the application is developed for use on a mouseless touchscreen computer that always open the browser in fullscreen mode.
          In order to print, we create a transitional pop-up window that opens upon clicking print button and closes after the printing action completes, whether successfully or not. All this pop-up window does is initiate an onload action via print_win() function. This is a simple Javascript function that uses the default DOM print()and close() functions to print and close respectively. Here is the print_win() function:
             function print_win(){
                window.print();
                window.close();
             }
          Now we let Javascript create and destroy the pop-up window. This Javascript is invoked upon clicking the print button defined in metaData div. The pop-up window will get its data from the innerHTML of the element whose id is passed as argument to printContent()function. This implies we have to make sure that all data that we would like to be printed is placed in the right div, otherwise any data outside it will not be printed. In our example, the data is placed in the document div.
          We should also remember that since print_win() function belongs to the pop-up window, it will be embedded within the outer Javascript that creates the pop-up window. We place our code in report_printer.js in the javascripts sub-directory. Here is the code that does it all:
               function printContent(id){
               var data = document.getElementById(id).innerHTML;
               var popupWindow = window.open('','printwin',
                    'left=100,top=100,width=400,height=400');
               popupWindow.document.write('<HTML>\n<HEAD>\n');
               popupWindow.document.write('<TITLE></TITLE>\n');
               popupWindow.document.write('<URL></URL>\n');
               popupWindow.document.write('<script>\n');
               popupWindow.document.write('function print_win(){\n');
               popupWindow.document.write('\nwindow.print();\n');
               popupWindow.document.write('\nwindow.close();\n');
               popupWindow.document.write('}\n');
               popupWindow.document.write('<\/script>\n');
               popupWindow.document.write('</HEAD>\n');
               popupWindow.document.write('<BODY onload="print_win()">\n');
               popupWindow.document.write(data);
               popupWindow.document.write('</BODY>\n');
               popupWindow.document.write('</HTML>\n');
               popupWindow.document.close();
            }
          Our application is ready for printing. Try printing your document using the print button.
          1. Improving the Output
          Trying to print the document shows that the data is rightly captured but not well structured as required. Let us add report.css to the pop-up window in print mode using some Javascript code. We modify and add a simple statement in our Javascript code: popupWindow.document.write("<link href='/stylesheets/report.css' media='print' rel='stylesheet' type='text/css' />\n"); We also add a dummy line that formats the screen layout of document on the pop-up window so that it looks better. Notice that the line is similar to the previous one except that we specify the media as screen. Our function now looks like this:

               function printContent(id){
               var data = document.getElementById(id).innerHTML;
               var popupWindow = window.open('','printwin',
                    'left=100,top=100,width=400,height=400');
               popupWindow.document.write('<HTML>\n<HEAD>\n');
               popupWindow.document.write('<TITLE></TITLE>\n');
               popupWindow.document.write('<URL></URL>\n');
               popupWindow.document.write("<link href='/stylesheets/report.css' media='print' rel='stylesheet' type='text/css' />\n");
              popupWindow.document.write("<link href='/stylesheets/report.css' media='screen' rel='stylesheet' type='text/css' />\n");
               popupWindow.document.write('<script>\n');
               popupWindow.document.write('function print_win(){\n');
               popupWindow.document.write('\nwindow.print();\n');
               popupWindow.document.write('\nwindow.close();\n');
               popupWindow.document.write('}\n');
               popupWindow.document.write('<\/script>\n');
               popupWindow.document.write('</HEAD>\n');
               popupWindow.document.write('<BODY onload="print_win()">\n');
               popupWindow.document.write(data);
               popupWindow.document.write('</BODY>\n');
               popupWindow.document.write('</HTML>\n');
               popupWindow.document.close();
            }

          We have finished successfully. Your application should be able to produce documents that are as nice as your original screen versions. You can print directly or to files, whether PDF, PostScript, etc.

          Happy coding!!

          Tuesday, April 27, 2010

          Javascript Functions for Prototypers

          I have been working on a a web-based project for sometime now. Because of the diverse requirements on the system, I was required to have as many interactions with the users' data from a client-side point of view. I could not find inbuilt functions that could solve my problems. So I was compelled to extend the String object so as to suit my needs.

          I should admit that they may not be all that robust to take pride in, but they are able to do what I wanted any way! I believe that there exists someone like me, looking for same (if not similar) functions and probably they are tired with googling. There might also be some who are working on some standardized library and may have been thinking of functions likes these. I guess this will be a starting point for both.

          I have tried to make them very simple (simple in all senses) and straightforward, but if you have questions, keep them flowing. We are here to help each other. Constructive criticisms and suggestions are surely welcome!!

          Programming hint: In each of these functions, You can improve the split functions to make comparisons using regular expression matching

          1. contains() function
          2. /* checks for the presence of a substring in a given string of
            * semi-colon separated substrings.
            * it returns 'true' if found, otherwise it returns 'false'
            *
            * for example :
            * 1. ["programming;in;javascript;is;cool"].contains("javascript") => true
            * 2. ["programming;in;javascript;is;cool"].contains("java") => false
            *
            * TO DO: ADD HANDLING OF 'SPACE' SEPARATED SUBSTRINGS
            *
            */
            String.prototype.contains = function (substring) {

            var array_of_strings = this.split(';');

            if (jQuery.inArray(substring, array_of_strings)>= 0) {
            return true;
            }
            else {
            return false;
            }
            }
          3. capitalize() function
          4. /* capitalizes a given string
            * Author: Edmond Kachale
            * for example :
            * "ProGraMming Is CooL".capitalize() => "Programming is cool"
            */
            String
            .prototype.capitalize = function(){
            var capitalized_string = new Array();

            if((this.length> 0)){
            capitalized_string.push(this[0].toUpperCase());
            capitalized_string.push(this.substring(1,this.length).toLowerCase());

            return
            capitalized_string.join("");
            }
            else{
            return
            this;
            }
            }

          5. titleize() function

          6. This function depends on capitalize() function above

            /* "titleizes" a given string
            * Author: Edmond Kachale
            * for example :
            * "Programming is cool".titleize() => "Programming Is Cool"
            */

            String.prototype.titleize = function(){
            var
            titleized_string = new Array();
            var
            sub_strings = this.split(" ");

            for(i = 0; i <this.length; i++)
            titleized_string.push(sub_strings[i].capitalize());

            return titleized_string.join(" ");
            }

          Enjoy your programming!

          NB: Forgive me for poor formatting. I didn't have enough time to create a custom css file!


          Tuesday, November 24, 2009

          Towards Software Localisation in Malawi

          In the twenty-first century, information has become such a fundamental aspect of each society that its access is a basic human right. This has prompted for convenient access to reliable and up-to-date information for decision making through the use of the Information and Communication Technologies (ICTs). However, most people in Africa fail to access information due to the language barrier as most of ICTs use languages that are of western origin.

          How do we remove the language barrier?

          In order to remove this language barrier, software localization has become one of the best alternatives to make ICT appropriate to a target locale. Localisation of software involves adapting a software product to the linguistic, cultural and technical requirements of a target population.

          Why localisation of software?

          In Malaŵi, Zambia and other Chicheŵa speaking regions, localisation of software is still in its formative stages as there are currently no localised software applications in computer systems.

          What shall we do, men and brethren?
          • The Kalembera Word Processor

          The Kalembera Word Processor

          I started localisation at my undergraduate. In my final year project, I developed a localized lightweight word processor, Kalembera. It has a Chicheŵa interface. You can get more details on the project from http://bkankuzi.blogspot.com/2009/01/undergrad-student-in-malawi-develops.html .

          Of course there are some few things to get fixed and a few features to be added. Right now the word processor does not fully support tables and graphics. In addition, I have developed a spell checker separately. I am yet to add it behind the word processor.
          • The Mozilla Localisation Initiative
          The Mozilla Initiative is an up and coming project. It is one example of free software localization projects that I wish to run. The project will help indigenous Malaŵians (and other Chichewa speaking reagions) access the Internet using a localised web browser.

          I am looking for more contributors. If you are interested please send me an email entitled New Member: Chichewa/Chinyanja Localisation Project (chichewalocaliser(at)gmail(dot)com). This will help us track our members.
          • The OpenOffice Localisation Project
          Talks are under way with members of the OpenOffice.org to help us with Chichewa localization. I would like to have access to the software and add Chichewa words to it. If everything materializes, we will have a fully functional free office suite.

          Conclusion

          There are a lot of things to consider as far as localisation of software is concerned. I am not saying I am the jack of all trades in this respect. I need a lot of friends to join and help in. You do not need to be a programmer. You can contribute to the development of Chichewa equivalents for English terms. If you are too technical, so much the better. Linguists will also be of great significance to these projects as they will help in the editing and revising of the Chichewa terminologies.

          One thing which has to be taken note of is that these projects (except the Kalembera project) will be free and open source. Your contributions will be used and acknowledged using the general and public licenses that are used by the proprietors of these projects.