Showing posts with label research. Show all posts
Showing posts with label research. Show all posts

Thursday, November 11, 2010

PLAGIARISM

What is plagiarism?

The act of plagiarism is defined as “taking someone’s words or ideas as if they were your own”.

In the research community, more often than not, we find papers that cite themselves, authors who copy extracts from their own previous papers, etc. Some of you may call that reasonable plagiarism; and I can’t agree more. Usually, everybody does that. I think I very clearly remember doing that myself, though I might have removed it in later versions. Even better case is while writing a thesis: you do copy sections of your papers straight into your thesis, how much ever your advisor might have warned you against it. Your paper, your thesis, it is probably alright.

When an individual writes a thesis there are a lot of interesting problems, but most irritating of them all seem to be that writing is a lot of hard work. Okay, that is what I heard. But my view is that though it was a lot of hard work (to write my thesis), it was the most satisfying path and one that led to a better understanding of my own work.

I have collaborated with a lot of people in the past on common interests; some that lead to successful papers, some research reports and some that lead to friendships beyond academic life.

When collaborating with people following are the two things I always considered:

1. What is your contribution and what is theirs? This is more important if both (or all) of you are working towards your individual thesis.

2. The line between plagiarism and keeping context of a work is a very thin line. Keeping the thin line visible is as important as making sure that it doesn’t look like a thick line (implying your own work). When ‘n’ collaborators work on similar problems it is equally important to ensure that there isn’t a worrying overlap in content (either methodology or prose) of the theses.

I have tried to come up with a story that explains a case that could come up in anybody's life. An undergrad, say John, collaborated with a certain Peter, who graduated with a degree relatively recently. Peter had ideas that he didn’t have time for during his thesis, and John was looking for an interesting research problem where he could focus his energies towards his thesis. They agreed upon a lot of things and disagreed at some of them; but rather importantly, they got along well to do some fine work that was appreciated by many. Now Peter was the expert in the area they worked together, John used some ideas from his side to be applied in the problem setting that Peter gave.

The work went on well and they did great work; published papers and people talked about them at conferences (even for the wrong reasons at times!), such stuff, you know! But then when it came to writing thesis John did the unthinkable -- just blindly copied (extracted) sections/paragraphs from Peter’s thesis. Though most of the cut-copy-paste happened from Introduction and Related Work sections, I think this qualifies for being called plagiarism, certainly Peter believed so. Unfortunately, Peter didn’t have any knowledge about the content of John’s thesis, and in rather unbelievable circumstances, the thesis went through the thesis review process, and say John defends the thesis. Okay, not so unbelievable as I might sound. Obviously, you aren’t expecting John’s advisor to know the entire area and least of all ‘reading all theses in this area’. Now the same thing applies to the review committee; I’d assume if they pass the thesis for defense it isn’t their mistake since they may not even know the area. Their context of the area is only as limited as the research problem addressed in the thesis at hand and they are just doing their best to justify the contextual correctness of the thesis. If the owner of the thesis fails to know what to write and what not, we can’t expect the reviewers to straighten it either.

Obviously, from the story just described it is reasonable to say that Peter could be pissed at John. Matters became only worse when John said “I didn’t know this was not the right way to do it. I really didn’t know that people take thesis so seriously. Most of all, I didn’t know that our theses are put up on university website, and are publicly accessible”. Whether or not there is a public access to the thesis, whether or not somebody complains about a thesis, it is common knowledge that a copyright can’t be violated, and its infringement can be huge trouble. Today’s 10th grader would know that, and giving such a stupid explanation to Peter didn’t help John anyway.

Off the story. Now let us see why you shouldn’t copy from someone else’s work:

1. Among those who know you, your work and the work you copied from: you are already being considered stupid!

2. You have plans for long term research career in academia or industry? Well, remember that by plagiarizing you are just one step away from getting caught.

a. Industry might be tolerant, but you should know that Academia is a bitch when it comes to these things. Somebody finds out about it and that is it. You are over within the circle.

b. Even worse: Five-ten years from now in the midst of your, say flowery, career someone finds this abyss and guess what; there’s no better insult!

3. Ever wondered how many people are working on plagiarism detection worldwide? Only a bunch of them, but imagine what would happen if they discover that from the entire web your thesis is a good sample to present as a case study for their work.

4. What does it mean to Peter:

a. Well, if he is someone looking up to the academic ladder himself, all the similar academic difficulties could arise for him too. Only, he has no reason to get punished. Alas, it appears there is a reason: working with John in first place.

b. Academia or Industry, wherever the case may be, when this information goes out nobody will believe that Peter had no role in the plagiarism. It is natural to believe “Well they wrote papers together, and shared their theses too!”. This means Peter was also involved in plagiarism, and his career is screwed.

There could be ‘n’ such cases we might be able to study, only if we want to, but none of which is ever complete and we are still only talking about the case of two friends who worked together. It is funny to know in how many ways you could be violated, and without your knowledge. Sometimes, after knowing that such things can happen, it becomes harder to work with anybody inexperienced as a fear creeps in, that overtakes the sheer pleasure of working together on exciting problems. But the message of this story is not that; by maintaining a good decorum on how the thesis will be written and reviewed among the two parties it is trivially easy to achieve moral justice to all the Peters and an intellectual satisfaction to all the Johns.

Nobody suffers!

Moral of the story: see the comments; somebody might have something to say? No?

PS: See the number of labels I have for this post. Clearly shows I need to blog more frequently!

Monday, June 8, 2009

One-liners from Reviews

As I have told before the paper that is now an accepted publication at ACL-IJCNLP 2009 went through 9 review cycles. At each one of those I have received some abusive reviews. At some of them I also received wonderfully positive reviews (with a sugar candy attached to it). At all of them I received neutral reviews (Yes, all of them!).

I thought I could share with you some of the most abusive, some of the most sweet and some stupidly neutral reviews (which won't help an acceptance or rejection or betterment of the paper!). What follows are quotes from the reviews and I hope the reviewers themselves won't read this some day. ;-) It doesn't matter to me at all, they were all anonymous so if they flame me that means they are revealing their identity. :D


Abusive

  • I was, unfortunately, unable to identify the particular contribution that this paper makes to this body of literature.
  • First, I was unable to determine exactly what the authors meant by "query focus" in their discussion.
  • The paper studies the role of query-relevance in multi-document summarization, and makes several findings: (1) ..... (2) ........ (3) .......
    • I find (1) and (2) particularly trivial and not very interesting
    • Finding (3) is perhaps less expected, though still makes quite a bit of sense.
  • I believe that this paper should not be accepted for publication because the experiments described do not support in any way the authors’ aim – which is to determine whether or not query focus affects sentence selection in MDS.
  • The feature suggested is trivial and does not add anything to the set of features already in use in extractive summarisation. The authors ignore decades of research in IR aimed at finding non-obvious matches of query-terms.
  • The observations made in Table 2 are not only trivial and non-surprising, but even absurd.
  • This paper is quite confusing. I was not always entirely sure of the point it was trying to make.
  • Modeling this as a likelihood is interesting, but perhaps complicates the issue somewhat (presumably a simple count of proportions, which can also be tested statistically, would do?)
  • This is not particularly surprising. Some discussion (even speculative) of the implications of this -- i.e. what are the humans doing that the systems are not -- would make this poster much more interesting.
  • I think the fundamental notion behind this paper is sound and interesting, but a lot of the analysis is flawed.
  • Theoretical analyses of data sets used in past work are also interesting, especially if they are thoroughly studied and written up well, but this paper tends to be somewhat hard to read and doesn't (at least to this reviewer) yield an "a ha!" moment at the end.
  • The results seem obvious, and the question does not seem to justify the effort involved in quantifying the differences in query bias between human and automated summaries. Some aspects of the data and the method are questionable.
  • The methodology used in the paper is not wrong, but the outcomes are rather obvious and it is difficult how the findings can prove useful to the research community.
  • One of the conclusions is that computers and humans use different strategies to produce summaries. This is well-known and it is not necessary to use statistics for this. The authors try to suggest that the difference between human summarisation and automatic summarisation is the way a query is handled, but no support for this claim is offered.
  • This paper deals with an interesting issue which is the distinction between query-focused and query-biased summarization. However, it fails for several main reasons:
    • It fails to define the problem and to make the distinction between the query-focused and query-biased summarization. The problem itself is ill-defined.
    • the methods used and the formalization are weak.
    • the justification is irrelevant
    • the results are too lousy and there is no funding at all.

Sweet

  • This is interesting work. You need to stay focused in the conclusions and indicate the significance for automatic text summarisation of feature analysis as a predictor of content.
  • This is an interesting paper. Using the appropriate literature, the paper analyzes the role of query focus in summarization performance on the DUC 2007 task. The analysis is insightful and could become a standard reference, if the presentation of the experiments was clearer and their interpretation less confusing.
  • Sect 6.3: this is the ultimate insight of this paper, please elaborate!
  • Clearly-written paper on an important topic. The gap between human and machine produced summaries is very interesting.
  • The work is interesting, supported with good evidence and is well presented. I support its inclusion in the conference as a short paper.
  • The paper is very well written, and provides a useful comparison between human and machine approaches to query focused automatic summarisation.

Neutral

  • I particularly liked the introduction of an equiprobable summarizer, which allows the impact of query-focus to be observed in a system setting, in addition to observations from the existing data sources.
  • This definitely an interesting and worthwhile topic -- and one that has attracted plenty of attention from researchers working in both the MDS and QA communities.
  • While I commend the authors for noting that their results will not be of use to systems that are performing well, the results seem to me to be too weak in general to be of interest.
  • The investigation reported is interesting, but in some sense it is obvious.
  • The conclusion is interesting, however, not surprising.


I have (now) no complaints about any of the reviews. Some of the abusive reviews were not wrong after all, it is correct from their point-of-view, their context. Of course the one I would want to fight against is the last abusive statement above. If the reviewer doesn't understand the problem then it is his problem partly, he can't have high confidence on his reviews, can he ? Anyway, this post is only to show the variation in human perception about what they read. Everybody has had their judgment. Everybody gave their verdict. But what is the truth ? Whether or not this paper is useful to the research community can be decided a few years from now based on whether or not someone cites it outside our group. I am patient enough to wait. :-)


PS: There is a marginally related discussion on natural language processing blog: How to reduce reviewing overhead?

PS2: There is another recent related discussion on the corpora-list where Adam Kilgarriff raised issues over the current process of reviewing and accused it of wastage of time. Others responded rather critically saying "There is no other simpler way to get author feedback and no known simple way to get away from awful reviewing!".

PS3: Another post on the art of reviewing seems worthy now. May be I will come back with it some day.

Saturday, June 6, 2009

The Journey to ACL

Alright. To start writing on a positive note, I would say "Eureka!!".

The Eureka moment of my life
I was about to go out for dinner that evening. When Surya finds me rather restless while he had some words to share with me.
He said "Your paper has been accepted, right ?".
I thought he was referring to the HLT-NAACL CLIA Workshop paper which got accepted 2 months back.
Then he reiterated "Congratulations! You have done it in ACL".

I closed my eyes in disbelief , opened up again, breathing harder now and asked him not to joke. He said "Check for yourself at the website". I did check it there and was almost crying in happiness. There was a different kind of pain I was addressing suddenly. I was relieved of a lot of pain and that was the reason of this unknown pain. Oh, I was crazy! I must have chanted "ACL....ACL" a few thousand times that night. :D

No doubt, I was happy. :-)

PS: For those who are unaware of ACL, it stands for the Association of Computational Linguistics. The premiere conference of our area of research. There hasn't been a single indian university publication in ACL since 1993. Then, it was Dr Sangal's publication I heard. Last year somebody from IIT-B had a publication, I have heard (not confirmed). And now it was me. :-)

The Midnight oil
This work started from a poster that I saw at IJCNLP last year in Hyderabad. Something struck me at that time in that poster. The poster was by Shilpa Arora (to my memory!). The poster was a little too simple or the explanation wasn't too convincing, that I wasn't impressed with the work. But that is how I thought about a few things at that time. That work is no way related to this my current work, but I somehow still feel that my initial thoughts started because of my disappointment at that poster presentation.

I continuously pondered over the initial ideas but had no one to discuss my points. I had nothing to claim. I just had a few observations by then. All of which would be rubbed off saying "trivial". (Ironically, even one of the ACL reviews mentioned it "Trivial", yet he found it necessary to be included!)

Then started my work independently but I never had the habit of being isolated. I drew upon Suman's habit of staying up late night and used him as a pseudo-Reviewer. Every night I would take him to the white board and explain what I think. Most of the times I repeated myself. This process brought a lot of clarity on my subject. He almost never asked questions but when he did they were usually good. Especially because they came from a person who had no idea what I was going to do with the data at hand.

Suman went away and I submitted the paper to EMNLP 2008, to start with.

End-of-Cycle
To start with I sent the paper to EMNLP 2008. But later on due to the rejects (mostly rightly so!) I kept on sending it to each and every conference that was in my path. I ended up sending the paper to 9 conferences in all. Starting with EMNLP 2008 and ending with EMNLP 2009. Having said that EMNLP 2009 was just a formality to mark the end of cycle.

EMNLP 2008 --> ICON 2008-->ECIR 2009 --> EACL 2009 --> HLT-NAACL 2009 --> SIGIR 2009 --> RANLP 2009 --> ACL-IJCNLP 2009 --> EMNLP 2009.

At ACL-IJCNLP the journey ended peacefully. I had submitted the paper as a single authored short paper and is now an accepted publication.

The role of a mentor
In november 2008, I saw a mail on SIGIR-list. SIGIR 2008 Mentorship program it read. The idea is that senior program committee members could help/mentor junior researchers from unprivileged institutions like ours where there aren't any senior people who could help. This was of some hope to me as I knew that my paper needed major revisions and smoother polishing before I can send to a good place. SIGIR was the right choice. I was handed over to Dr Charles L A Clarke, Assistant Professor at the University of Waterloo. After over 50 conversations including reviews, questions and answers I ended up in a draft that was indeed smooth and soothing to the eye. I submitted to SIGIR poster (which again was suggested by Charlie himself. Thank you!). It got rejected, but that step had changed my own perspective of the whole work. I was thinking from new point of view, I now had new vocabulary to explain the terms. Reviews from SIGIR were equally important they gave me the right foundation as to what I should do before the next submission. I took the reviews seriously and used them thoroughly to extend it to ACL short paper.


Thanking my colleagues!
I never got the chance to thank all my colleagues who helped me mature this paper to the current level or those of them who continuously read each of my innumerable drafts and gave their valuable feedback. I would like to thank Sowmya, Prasad, Swathi, Praneeth, Santosh, Vasu and Chandan for reading my paper and for their inputs (if any!). I would also like to thank some colleagues outside lab Sai Satya, Mahesh Mohan etc who gave their helping hand when in need. Thanks Mahesh for those un-conventional thoughts, I am treasuring them!



PS: The first title I gave to this post was : "From Kondapur to Singapur ". True isn't it ? Why? You'd do good to know that ACL-IJCNLP 2009 is going to be held in Singapore. And by the way, I am going to attend it.

PS2: Considering the events that are happening around me now, there could be more posts on ACL soon. Some more positive publicity to ACL and perhaps negative publicity to some others!

Here is a cache copy of the accepted papers page. (I am hoping that ACL-IJCNLP doesn't have a problem with this.)