Thursday, May 10, 2007

Paul Kedrosky writes:
At the same as a number of engineering friends of mine are leaving Google -- most common complaint: too big and bureaucratic -- MBAs have declared it to be their employer of choice.
See also some of Paul's past posts ([1] [2] [3]) that speculate that MBA hiring may be a negative indicator for future performance.

See also my previous post, "Google and those TPS reports", that includes thoughts from Googler Chris Sacca and a brief tidbit about my experience with the influx of MBAs at Amazon.com.

Wednesday, May 9, 2007

Laurie Petersen reports on a keynote by Esther Dyson at the Search Insider Summit.

Some excerpts:
Dyson said, "I don't see the quality of search improving very much. Search is like telling a dog, 'Go Fetch,' I want something to 'Go Fetch and Reserve' [as in the right hotel room.]"

What's needed, she said, is switching from a "search and fetch" mentality to a "deliver, act and transact" perspective based on personalization.

The real winner, Dyson said, will be a custom-built tool that understands the nuance of an individual, his or her phrasing, and specific likes and dislikes. This tool will incorporate both domain knowledge and user knowledge.
However, what Esther is suggesting may go well beyond search. It would seem to enter into the realm of software agents.

To "deliver, act, and transact", not only would the computer need to understand intent, but also it may have to come up with complicated plans to satisfy the request. It would have to have a rich understanding of information acquired, combine information from multiple sources, interact with external actors, and deal with uncertainty in information, actions, and intent.

At this point, you are basically talking about building a software robot, a softbot. Building the brains behind a software servant that can deliver, act, and transact is going to be about as hard as building the brains behind a hardware robot servant. Dropping motor skills does not ease the task of building higher brain functions.

A hard problem indeed. And one we are a long, long way from solving.

Wednesday, May 2, 2007

A paper at the upcoming WWW 2007 conference, "Open User Profiles for Adaptive News Systems: Help or Harm?" (PDF), concludes that allowing users to edit profiles used for news personalization can result in worse personalization.

From the paper:
Despite our expectations, our study didn't confirm that the ability to view and edit user profiles of interest in a personalized news system is beneficial to the user. On the contrary, it demonstrated that this ability has to be used with caution.

Our data demonstrated that all objective performance parameters are lower on average for the experimental system. It includes system recommendation performance as well as precision and recall of information collected in the user reports.

Moreover, we found a negative correlation between the system performance for an individual user and the amount of user model changes done by this user. While the performance data vary between users and topics, the general trend is clear – the more changes are done, the larger harm is done to the system recommendation performance.

The results of our study confirmed the controversial results of Waern's study: the ability to change established user profiles typically harms system and user performance.
The paper was not clear on exactly why editing the profiles made the personalization worse, but I would look to what Jason Fry at the Wall Street Journal wrote several months back:
When it comes to describing us as customers and consumers, recommendation engines may do the job better than we would.

In other words, we lie -- and never more effectively than when we're lying to ourselves ... I fancy myself a reader of contemporary literature and history books, but I mostly buy "Star Wars" novels and "Curious George" books for my kid.
As I wrote after seeing Jason's article, "Implicit data like purchases may be noisy, but it also can be more accurate. You may say you want to watch academy award winners, but you really want to watch South Park. You may say you want to read Hemingway, but you really want to read User Friendly."

See also an April 2004 post where I said, "When you rely on people to tell you want their interests are, they (1) usually won't bother, (2) if they do bother, they often provide partial information or even lie, and (3) even if they bother, tell the truth, and provide complete information, they usually fail to update their information over time."

Tuesday, May 1, 2007

I finished reading "Hard Facts, Dangerous Half-Truths, and Total Nonsense" a few weeks ago.

It is a book by Stanford Business School Professors Jeffrey Pfeffer and Robert Sutton arguing for evidence-based management.

Evidence-based management advocates making decisions based on the latest and best knowledge of what actually works. Yes, you might think this would not be controversial, but, sadly, it is in the world of business.

Early on, the authors state their position:
We believe that managers are seduced by far too many half-truths: ideas that are partly right but also partly wrong and that damage careers and companies over and over again.

Yet managers routinely ignore or reject solid evidence that these truisms are flawed.
Why do managers succumb to these half-truths? In part, it is that they get bad advice.
The advice managers get from the vast and ever-expanding supply of business books, articles, gurus, and consultants is remarkably inconsistent.

Consider the following clashing recommendations, drawn directly from popular business books: Hire a charismatic CEO; hire a modest CEO. Embrace complexity theory; strive for simplicity. Become ... strategy-focused; ... strategic planning ... is of little value.

Consultants and others who sell ideas and techniques are always rewarded for getting work, only sometimes rewarded for doing good work, and hardly ever rewarded for whether their advice actually enhances performance ... If a client company's problems are only partially solved, that leads to more work for the consulting firm.

The senior executive of a human resources consulting firm, for example, told us that because pay-for-performance programs almost never work that well, you usually get asked back again and again to repair the programs your clients bought from you.
But, it is not all just bad advice.
When the late Peter Drucker was asked why managers fall for bad advice and fail to use sound evidence, he didn't mince words: "Thinking is very hard work. And management fashions are a wonderful substitute for thinking."
There is a lot to like in the book, but I found this tidbit on building a culture that promotes learning to be particularly good:
A series of studies by Columbia University's Carol Dweck shows .... that when people believe they are born with natural and unchangeable smarts ... [they] learn less over time. They don't bother to keep learning new things.

People who believe that intelligence is malleable keep getting smarter and more skilled ... and are willing to do new things.

These findings also mean that if you believe that only 10 percent or 20 percent of your people can ever be top performers, and use forced rankings to communicate such expectations in your company, then only those anointed few will probably achieve superior performance.

[Managers should] treat talent as something almost everyone can earn, not that just a few people own.

Having people who know the limits of their knowledge, who ask for help when they need it, and are tenacious about teaching and helping colleagues is probably more important for making constant improvements in an organization.
The authors are no fans of forced rank or pay-for-performance, and they spend a fair amount of time explaining why:
A renowned (but declining) high-technology firm [used] a forced-ranking system ... [where] managers were required to rank 20 percent of employees as A players, 70 percent as Bs, and 10 percent as Cs ... They gave the lion's share of rewards to As, modest rewards to Bs, and fired the Cs.

But in an anonymous poll, the firm's top 100 or so executives were asked which company practices made it difficult to turn knowledge into action. The stacking system was voted the worst culprit.

A survey of more than 200 human resource professionals ... reported that forced ranking resulted in lower productivity, inequity and skepticism, negative effects on employee engagement, reduced collaboration, and damage to moral and mistrust in leadership.

People are more likely to ... see themselves more positively than others see them [and] believe they are above average or not recognize their lack of competence .... People who receive a smaller reward than they expect routinely resent the organization.

A 2004 survey ... of 350 companies showed that "83 percent of organizations believe their pay-for-performance programs are only somewhat successful or not successful at accomplishing their goals.
Similarly, an experiment at HP with "13 different pay programs" in the 1990s found that the "local managers who enthusiastically initiated these pay-for-performance programs ran into difficulties in implementation and maintenance" and soon wanted to abandon them. The experiments "did find that pay motivated performance" but that "the costs weren't worth the benefits" in terms of the "lost trust", "damaged employee commitment", "shift of focus away from the work and toward pay", "infighting about pay", and overhead of managing the programs.

Though HP learned the right lesson from their experiment at the time, this story does not have a happy ending. Years later, Carla Fiorina "forced the system throughout HP, disregarding the evidence gathered by the company itself." The authors' judgement is scathing: "CEO whim, belief, ego, and ideology, rather than evidence, direct what too many companies do and how they do it."

And on this note, Pfeffer and Sutton write that the "belief that leaders ought to be in control is a dangerous half-truth" because "leaders make mistakes -- all people do" and, with "few or no checks or balances", there are no way to correct the inevitable errors.

It's a good book, more grounded than most of the business fluff around these days, and worth reading if you are a manager or play one during the day.

Also, if you haven't read it already, I would strongly recommend Pfeffer's earlier book, "The Human Equation", which I briefly discussed in an old April 2004 post. The older book is tighter and more focused on management and HR practices than the newer book; the newer book is more of an assault on and lamentation of the state of management.

See also my earlier posts, "The problem with forced rank" and "Microsoft drops forced rank, increases perks".
Googlers Neil Daswani and Michael Stoppelman wrote a paper, "The Anatomy of Clickbot.A" (PDF), on a clever click fraud attack against Google using a botnet.

From the paper:
This paper presents a detailed case study of the Clickbot.A botnet. The botnet consisted of over 100,000 machines ... [and] was built to conduct a low-noise click fraud attack against syndicated search engines.

Some computers [in the botnet] were infected by downloading a known trojan horse. The trojan horse disguised itself as a game .... It does not slow down a machine or adversely affect a machine’s performance too much. As such, users have no incentive to disinfect their machines of such a bot.

Several tens of thousands of IP addresses of machines infected with Clickbot.A were obtained. An analysis of the IP addresses revealed that they were globally distributed ... The IP addresses also exhibited strong correlation with email spam blacklists, implying that infected machines may have also been participating in email spam botnets as well.

Conducting a botnet-based click fraud attack directly against a top-tier search engine might generate noticable anomalies in click patterns, Clickbot.A attempted to avoid detection by employing a low-noise attack against syndicated search engines.
This is a pretty impressive attempt. Since the attacker disguises itself as a large number of independent users, this type of click fraud would be a challenge to detect.

I have to wonder whether there could be other, more clever attackers using similar methods that are slipping by undetected. For example, using a larger botnet would make the fraud more difficult to detect. It does appear that a several tens of thousands of IP addresses is not huge for a botnet; one was discovered back in 2005 that had 1.5M machines.

I also suspect an attacker could also make the fraud pattern more challenging to find by mimicking normal search and ad click patterns most of the time, especially on machines that are otherwise idle, which otherwise would stand out by doing nothing but fraudulent clicks.

As the paper says, this type of botnet-based click fraud is only likely to increase. Security researchers like Neil and Michael have their work cut out for them.
Brad Fitzpatrick from LiveJournal gave an April 2007 talk, "LiveJournal: Behind the Scenes", on scaling LiveJournal to their considerable traffic.

Slides (PDF) are available and have plenty of juicy details.

One thing I wondered after reviewing the slides -- and this is nothing but a crazy thought exercise on my part -- is whether they might be able to simplify the LiveJournal architecture (the arrows pointing everywhere shown on slide 4).

In particular, there seems to be a heavier emphasis on caching layers and read-only databases than I would expect. I tend to prefer aggressive partitioning, enough partitioning to get each database working set small enough to easily fit in memory.

I know little about LiveJournal's particular data characteristics, but I wonder if aggressive partitioning in the database layer might yield performance as high as a caching layer without the complexity of managing the cache consistency. Databases with data sets that fit in memory can be as fast as in-memory caching layers.

Likewise, I wonder if there would be benefit from dropping the read-only databases in favor of partitioned databases. With partitioned databases, the databases may be able to fit the data they do have entirely in memory; read-only replicas may still be hitting disk if the data is large.

Hitting disk is the ultimate performance killer. Developers often try to avoid database accesses because their database accesses hit disk and are dog slow. But, if you can make your database accesses not hit disk, then they can be blazingly fast, so fast that separate layers even might become unnecessary.

Again, wild speculation on my part. Everything I have seen indicates that Brad knows exactly what he is doing. Still, I cannot help myself from wondering, is there any way to get rid of all those layers?

[Slides found via Sergey Chernyshev]

Update: Another version of Brad's talk with some updates and changes.
Greg Sterling at Search Engine Land, in his article "iGoogle, Personalized Search and You", quotes Google VP Marissa Mayer as saying:
[Personalization is] one of the biggest relevance advances in the past few years.

Personalization doesn't affect all results, but when it does it makes results dramatically better.
See also excerpts from a few recent articles by Search Engine Land columnist Gord Hotchkiss in my earlier post, "Personalization, Google, and discovery".
Glinden BlogThe owner of this website is a participant in the Amazon Services LLC Associates Program, an affiliate advertising program designed to provide a means for sites to earn advertising fees by advertising and linking to Amazon properties including, but not limited to, amazon.com, endless.com, myhabit.com, smallparts.com, or amazonwireless.com.