Monday, July 7, 2008

A WWW 2008 paper out of Google, "Video Suggestion and Discovery for YouTube: Taking Random Walks Through the View Graph" (PDF) presents "a novel method based on the analysis of the entire user-video graph to provide personalized video suggestions for users."

Not sure if the same technique was used, but YouTube recently launched ([1] [2]) recommendations.

Update: Ten days later, Googler Sapna Mehta announces video recommendations on Google Video.

These recommendations appear to be quite different than the YouTube recommendations, based on your Google web and search history rather than the YouTube videos you watch and almost certainly using different techniques under the covers.

Interesting to see Google making a push on video recommendations both on YouTube and Google Video.

Please see also a review of the new Google Video recommendation feature on Search Engine Land.

Saturday, July 5, 2008

The press seems to have a pattern reporting on successful technology companies. First, these companies can do no wrong, the heroes of our time, bringing us clever new ways of doing things that promise to dramatically change our lives.

Then, reporters start nipping away at the edges. Maybe the hero has flaws? Is our hero really good? At the first sign of blood, the fangs sink in, and a frenzied attack tears down the hero the press itself had once created.

This happened dramatically when I was at Amazon. In 1999, Jeff Bezos was Time Magazine's Person of the Year. In 2000, we were "Amazon.bomb", no longer a revolution of retailing, but a fraud and a scam.

It probably is easier to draw blood in times of recession, when the business is already strained. The bloom came off Amazon's rose in 2000 during the stock market crash. In this recession of 2008, it looks like Google might have to endure similar treatment.

A few years ago, Google was built up into a hero. "The search engine that could" ([1]) had a "godlike view of what us mere mortals are thinking" ([2]). Google was "one of the most remarkable Internet successes of our time ... powered by the world’s most advanced technology ... [and] in a few short years has revolutionized access to information about everything for everybody everywhere." ([3]). The "all-knowing voice [will] always deliver the precise answer to any question in a fraction of a second." ([4]).

Now, the press is nipping away at our hero. Some question whether Google has become evil ([5] [6]). Some ask if our hero knows too much and question its motives ([7] [8]). There was even a rather goofy piece in the NYT today portraying Sergey Brin and Susan Wojcicki as insensitive tyrants who ignore the day care needs of Googler parents.

So far, the press has failed to draw much blood, but I doubt that will deter them. If the Amazon experience is any guide, this will not end until the beast has been fed. The press has built up their hero. They will now tear it back down. The cycle must complete.

For a more lighthearted view on all of this, don't miss The Onion's story, "Google Announces Plan To Destroy All Information It Can't Index".

Friday, July 4, 2008

Kevin Rose writes that Digg is launching a recommendation engine that "uses your past digging activity to identify what we call Diggers Like You .. [and] suggest stories you might like." Good to see it.

MG Siegler's posts at VentureBeat ([1] [2]) include some early reviews of the feature.

Please see also my previous posts on Digg, especially "Digg struggles with spam" and its discussion of how recommendations could reduce spam by mitigating the winner-takes-all effect.

Wednesday, July 2, 2008

A VLDB 2008 paper out of Yahoo Research, "Scheduling Shared Scans of Large Data Files" (PDF) looks at how "to maximize the overall rate of processing ... by sharing scans of the same file ... [in] Map-Reduce systems [like Hadoop]".

I felt the model used for simulations in the paper was a bit questionable -- seems to me the emphasis should be on newer data being accessed by many jobs simultaneously, most often by smaller jobs -- but the specifics of the solution probably are of less interest than the paper's general discussion of scheduling issues that come up as a large Hadoop cluster is put under load from many users.

Please see also my earlier post, "Hadoop summit notes", especially the problems with scheduling using Hadoop on Demand (which makes a cluster look like multiple virtual clusters) and the idea of sharing intermediate results (such as sorted extracts) from previous jobs on the cluster.
Brady Forrest at O'Reilly Radar points out Amazon's new page recommendation widget in his post, "Amazon's Page Recommender: Foreshadowing A New Web Service?"

Put the widget on your website and, for any page it is on, Amazon can learn what people visit that page, where they go on your site, and then, from those behavior patterns, generate recommendations on where people might want to go from each page.

It could, for example, be used on a news website to generate news recommendations or, on a shopping site, to recommend products.

It is an interesting move by Amazon, a step toward Aggregate Knowledge and others that offer recommendations as a web service.
In his post, "Inside Microsoft's Internet Infrastructure & Its Plans for the Future", Om Malik highlights some interesting "facts about Microsoft-owned data centers":
[Microsoft is] adding 10,000 servers a month

Network backbone ... soon ... [will be] 500 Gigabits.

Data in the near future will soon approach 100s of petabytes.
All the computation you could want and all the data you can eat. Perhaps now it makes more sense why I am at Microsoft?

Please see also my April 2006 post, "Microsoft is building a Google cluster".
Netflix has a feature called Profiles that allows multiple people to use the same Netflix account while keeping their queue and recommendations separate.

Recently, they attempted ([1] [2]) to remove the feature because "too many members found the feature difficult to understand and cumbersome, having to consistently log in and out of the website."

It's an interesting example how difficult it is to allow people to maintain multiple personas on a website. At Amazon, the second biggest complaint about the recommendations, next to recommending items already bought from a different store, was that sharing an account or buying a gift for someone that was not marked as a gift would end up mixing up recommendations for multiple people together.

But, the problem is that there is no easy and convenient way to provide multiple personas. Power users might be able to login and logout of different accounts, but most people don't understand or want to bother with that.

Hard problem. It's not clear to me what the solution might be.
Glinden BlogThe owner of this website is a participant in the Amazon Services LLC Associates Program, an affiliate advertising program designed to provide a means for sites to earn advertising fees by advertising and linking to Amazon properties including, but not limited to, amazon.com, endless.com, myhabit.com, smallparts.com, or amazonwireless.com.