Today I read "How does clickthrough data reflect retrieval quality?" by F. Radlinski, M. Kurup, and T. Joachims published at CIKM 2008.
Greg Linden blogged about it ages ago.
Radlinski et al. have access to a search engine that is used by many people every day. In the paper the researchers describe how they successfully frustrated and annoyed users over a period of about 2 months by making it harder to find what they were looking for. The reason why they did this is to study different ways of measuring how unhappy users are with search results.
It's a very interesting paper. And I can't believe they actually did that. The standard approach where I work would be to experiment with measuring improvements as you implement them. Ideally you would have something close to a ground truth (e.g. expert ratings) and then you would try to find which of the numbers you can automagically compute corresponds best. Intentionally annoying users to answer interesting questions is wrong, even if it is very tempting sometimes.
Furthermore, it would have been nice if the researchers would have reported at least one improvement they were able to find using their findings on how to measure the quality of search results.
It would also have been nice to have a bit more of a discussion on the role of expert ratings.
And finally, it would also have been nice if they would have at least briefly discussed how interleaving results does not work if the results are not independent of each other.
Showing posts with label academia vs industry. Show all posts
Showing posts with label academia vs industry. Show all posts
Sunday, 1 March 2009
Subscribe to:
Posts (Atom)