Sebastian's Pamphlets

If you've read my articles somewhere on the Internet, expect something different here.

MOVED TO SEBASTIANS-PAMPHLETS.COM

Please click the link above to read actual posts, this archive will disappear soon!

Stay tuned...

Friday, June 08, 2007

Blogger abuses rel-nofollow due to ignorance

I had planned a full upgrade of this blog to the newest blogger version this weekend. The one and only reason to do the upgrade was the idea that I perhaps could disable the auto-nofollow functionality in the comments. Well, what I found was a way to dofollow the author's link by editing the <dl id='comments-block'> block, but I couldn't figure out how to disable the auto-nofollow in embedded links.

Considering the hassles of converting all the template hacks into the new format, and the risk of most probably losing the ability to edit code my way, I decided to stick with the old template. It just makes no sense for me to dofollow the author's link, when a comment author's links within the content get nofollow'ed automatically. Andy Beard and others will hate me now, so let me explain why I don't move this blog to my own domain using a not that insane software like WordPress.
  • I own respectively author on various WordPress blogs. Google's time to index for posts and updates from this blogspot thingy is 2-3 hours (Web search, not blog search). My Wordpress blogs, even with higher PageRank, suffer from a way longer time to index.
  • I can't afford the time to convert and redirect 150 posts to another blog.
  • I hope that Google/Blogger can implement reasonable change requests (most probably that's just wishful thinking).
That said, WordPress is a way better software than Blogger. I'll have to move this blog if Blogger is not able to fulfill at least my basic needs. I'll explain below why I think that Blogger lacks any understanding of the rel-nofollow semantics. In fact, they throw nofollow crap on everything they get a hand on. It seems to me that they won't stop jeopardizing the integrity of the Blogosphere (at least where they control the linkage) until they get bashed really hard by a Googler who understands what rel-nofollow is all about. I nominate Matt Cutts, who invented and evolved it, and who does not tolerate BS.

So here is my wishlist. I want (regardless of the template type!)
  • A checkbox "apply rel=nofollow to comment author links"
  • A checkbox "apply rel=nofollow to links within comment text"
  • To edit comments, for example to nofollow links myself, or to remove offensive language
  • A checkbox "apply rel=nofollow to links to label/search pages"
  • A checkbox "apply a robots meta tag 'noindex,follow' to label/search pages"
  • A checkbox "apply rel=nofollow to links to archive pages"
  • A checkbox "apply a robots meta tag 'noindex,follow' to archive pages"
  • A checkbox "apply rel=nofollow to backlink listings"
As for the comments functionality, I'd understand when these options get disabled when comment moderation is set to off.

And here are the nofollow-bullshit examples.
  • When comment moderation and captchas are activated, why are comment author links as well as links within the comments nofollow'ed? Does blogger think their bloggers are minor retards? I mean, when I approve a comment, then I do vouch for it. But wait! I can't edit the comment, so a low-life link might slip through. Ok, then let me edit the comments.

  • When I've submitted a comment, the link to the post is nofollowed. Nofollow insane II.This page belongs to the blog, so why the fudge does Blogger nofollow navigational links? And if it makes sense for a weird reason not understandable by a simple webmaster like me, why is the link to the blog's main page as well as the link to the post one line below not nofollow'ed? Linking to the same URL with and without rel-nofollow on the same page deserves a bullshit award.

  • Nofollow insane III. (dashboard)On my dashbord Blogger features a few blogs as "Blogs Of Note", all links nofollow'ed. These are blogs recommended by the Blogger crew. That means they have reviewed them and the links are clearly editorial content. They're proud of it: "we've done a pretty good job of publishing a new one each day". Blogger's very own Blogs Of Note blog does not nofollow the links, and that's correct.

    So why the heck are these recommended blogs nofollow'ed on the dashboard? Nofollow insane III. (blogspot)

  • Blogger inserted robots meta tags "nofollow,noindex" on each and every blog hosted outside the controlled blogspot.com domain earlier this year.

  • Blogger inserted robots meta tags "nofollow,noindex" on Google blogs a few days ago.


If Blogger's recommendation "Check google.com. (Also good for searching.)" is a honest one, why don't they invest a few minutes to educate themselves on rel-nofollow? I mean, it's a Google-block/avoid-indexing/ranking-thingy they use to prevent Google.com users from finding valuable contents hosted on their own domains. And they annoy me. And they insult their users. They shouldn't do that. That's not smart. That's not Google-ish.

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Monday, June 04, 2007

Google nofollow's itself

Awesome. Nofollow-insane at its best. Check the source of Google's Webmaster Blog. In HEAD you'll find an insane meta tag:
<meta name="ROBOTS" content="NOINDEX,NOFOLLOW" />

Well, that's one of many examples. Read the support forums. Another case of Google nofollow'ing herself: Google fun

Matt thought that all teams understood the syntax and semantics of rel-nofollow. It seems to me that's not the case. I really can't blame Googlers applying rel-nofollow or even nofollow/noindex meta tags to everything they get a hand on. It is not understandable. It's not useable. It's misleading. It's confusing. It should get buried asap.

Hat tip to John (JLH's post).

Update 1: A friendly Googler just told me that a Blogger glitch (pertaining only Google blogs) inserted the crawler-unfriendly meta element, it should be solved soon. I thought this bug was fixed months ago ... if page.isPrivate == true by mistake then insert "<meta content='NOINDEX,NOFOLLOW' name='ROBOTS' />" ... (made up)

Update 2: The 'noindex,nofollow' robots meta tag is gone now, and the Webmaster Central Blog got a neat new logo:
Google Webmaster Central Blog - Offic'ial news on crawling and indexing sites for the Google index (I'd add ALT and TITLE text: alt="Google Webmaster Central Blog - Official news on crawling and indexing sites for the Google index" title="Official news on crawling and indexing sites for the Google index")

Labels: , , , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Wednesday, May 02, 2007

Yahoo! search going to torture Webmasters

According to Danny Yahoo! supports a multi-class nonsense called robots-nocontent tag. CRAP ALERT!

Can you senseless and cruel folks at Yahoo!-search imagine how many of my clients who'd like to use that feature have copied and pasted their pages? Do you've a clue how many sites out there don't make use of SSI, PHP or ASP includes, and how many sites never heard of dynamic content delivery, respectively how many sites can't use proper content delivery techniques because they've to deal with legacy systems and ancient business processes? Did you ask how common templated Web design is, and I mean the weird static variant, where a new page gets build from a randomly selected source page saved as new-page.html?

It's great that you came out with a bastardized copy of Google's somewhat hapless (in the sense of cluttering structured code) section targeting, because we dreadfully need that functionality across all engines. And I admit that your approach is a little better than AdSense section targeting because you don't mark payload by paydirt in comments. But why the heck did you design it that crappy? The unthoughtful draft of a microformat from what you've "stolen" that unfortunate idea didn't become a standard for very good reasons. Because it's crap. Assigning multiple class names to markup elements for the sole purpose of setting crawler directives is as crappy as inline style assignments.

Well, due to my zero-bullshit tolerance I'm somewhat upset, so I repeat: Yahoo's robots-nocontent class name is crap by design. Don't use it, boycott it, because if you make use of it you'll change gazillions of files for each and every proprietary syntax supported by a single search engine in the future. When the united search geeks can agree on flawed standards like rel-nofollow, they should be able to talk about a sensible evolvement of robots.txt.

There's a way easier solution, which doesn't require editing tons of source files, that is standardizing CSS-like syntax to assign crawler directives to existing classes and DOM-IDs. For example extent robots.txt syntax like:

A.advertising { rel: nofollow; } /* devalue aff links */

DIV.hMenu, TD#bNav { content:noindex; rel:nofollow; } /* make site wide links unsearchable */


Unsupported robots.txt syntax doesn't harm, proprietary attempts do harm!

Dear search engines, get together and define something useful, before each of you comes out with different half-baked workarounds like section targeting or robots-nocontent class values. Thanks!

Labels: , , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Friday, April 27, 2007

How Google & Yahoo handle the link condom

Loren Baker over at SEJ got a few official statements on use and abuse of the rel-nofollow microformat by the major players: How Google, Yahoo & Ask treat NoFollow'ed links. Great job, thanks!

Ask doesn't "officially" support nofollow, whatever that means. Loren didn't ask MSN, probably because he didn't expect that they've even noticed that they officially support nofollow since 2005, same procedure with sitemaps by the way. Yahoo implemented it along the specs, and Google stepped way over the line the norm sets. So here is the difference:

1. Do you follow a nofollow'ed link?
Google: No (longer)
Yahoo: Yes

2. Do you index the linked page following a nofollow'ed link?
Google: Obsolete, see 1.
Yahoo: Yes

3. Does your ranking algos factor in reputation, anchor/alt/title text or whichever link love sourced from a nofollow'ed link?
Google: Obsolete, see 1.
Yahoo: No

4. Do you show nofollow'ed links in reverse citation results?
Google: Yes (in link: searches by accident, in Webmaster Central if the source page didn't make it into the supplemental index)
Yahoo: Yes (Site Explorer)

Q&A#4 is made up but accurate. I think it's safe to assume that MSN handles the link condom like Yahoo. (Update: As Loren clarifies in the comments, he asked MSN search but they didn't answer in a timely fashion.)

And here's a remarkable statement from Google's search evangelist Adam Lasnik, who may like nofollow or not:
On a related note, though, and echoing Matt’s earlier sentiments ... we hope and expect that more and more sites — including Wikipedia — will adopt a less-absolute approach to no-follow ... expiring no-follows, not applying no-follows to trusted contributors, and so on.
Bravo!


Related link: rel="nofollow" Google, Yahoo and MSN

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Thursday, April 12, 2007

In need of a "Web-Robot Directives Standard"

The Robots Exclusion Protocol from 1994 gets used and abused, best described by Lisa Barone citing Dan Crow from Google: "everyone uses it but everyone uses a different version of it". De facto we've a Robots Exclusion Standard covering crawler directives in robots.txt and robots meta tags as well, said Dan Crow. Besides non-standardized directives like "Allow:", Google's Sitemaps Protocol adds more inclusion to the mix, now even closely bundled with robots.txt. There are more ways to put crawler directives. Unstructured (in the sense of independence from markup elements) like with Google's section targeting, on link level applying the commonly disliked rel-nofollow microformat or XFN, and related thoughts on block level directives.

All in all that's a pretty confusing conglomerate of inclusion and exclusion, utilizing many formats respectively markup elements, and lots of places to put crawler directives. Not really the sort of norm the webmaster community can successfully work with. No wonder that over 75,000 robots.txt files have pictures in them, that less than 35 percent of servers have a robots.txt file, that the average robots.txt file is 23 characters ("User-agent: * Disallow:"), gazillions of Web pages carry useless and unsupported meta tags like "revisit-after" ... for more funny stats and valuable information see Lisa's robots.txt summit coverage (SES NY 2007), also covered by Tamar (read both!).

How to structure a "Web-Robot Directives Standard"?

To handle redundancies as well as cascading directives properly, we need a clear and understandable chain of command. The following is just a first idea off the top of my head, and likely gets updated soon:

  • Robots.txt

    1. Disallows directories, files/file types, and URI fragments like query string variables/values by user agent.
    2. Allows sub-directories, file names and URI fragments to refine Disallow statements.
    3. Gives general directives like crawl-frequency or volume per day and maybe even folders, and restricts crawling in particular time frames.
    4. References general XML sitemaps accessible to all user agents, and specific XML sitemaps addressing particular user agents as well.
    5. Sets site-level directives like "noodp" or "noydir".
    6. Predefines page-level instructions like "nofollow", "nosnippet" or "noarchive" by directory, document type or URL fragments.
    7. Predefines block-level respectively element-level conditions like "noindex" or "nofollow" on class names or DOM-IDs by markup element. For example "DIV.hMenu,TD#bNav 'noindex,nofollow'" could instruct crawlers to ignore the horizontal menu as well as navigation at the very bottom on all pages.
    8. Predefines attribute-level conditions like "nofollow" on A elements. For example "A.advertising REL 'nofollow'" could tell crawlers to ignore links in ads, or "P#tos > A 'nofollow'" could instruct spiders to ignore links in TOS excerpts found on every page in a P element with the DOM-ID "tos".
    • XML Sitemaps

      1. Since robots.txt deals with inclusion now, why not add an optional URL specific "action" element allowing directives like "nocache" or "nofollow"? Also a "delete" directive to get outdated pages removed from search indexes would make sound sense.
      2. To make XML sitemap data reusable, and to allow centralized maintenance of page meta data, a couple of new optional URL elements like "title", "description", "document type", "language", "charset", "parent" and so on would be a neat addition. This way it would be possible to visualize XML sitemaps as native (and even hierarchical) site maps.
      Robots.txt exclusions overrule URLs listed for inclusion in XML sitemaps.
      • Meta Tags

      • Page meta data overrule directives and information provided in robots.txt and XML sitemaps. Empty contents in meta tags suppress directives and values given in upper levels. Non-existent meta tags implicitly apply data and instructions from upper levels. The same goes for everything below.
        • Body Sections

        • Unstructured parenthesizing of parts of code certainly is undoable with XMLish documents, but may be a pragmatic procedure to deal with legacy code. Paydirt in HTML comments may be allowed to mark payload for contextual advertising purposes, but it's hard to standardize. Lets leave that for proprietary usage.
          • Body Elements

          • Implementing a new attribute for messages to machines should be avoided for several good reasons. Classes are additive, so multiple values can be specified for most elements. That would allow to put standarized directives as class names, for example class="menu robots-noindex googlebot-nofollow slurp-index-follow" where the first class addresses CSS. Such inline robot directives come with the same disadvantages as inline style assignments and open a can of worms so to say. Using classes and DOM-IDs just as a reference to user agent specific instructions given in robots.txt is surely the preferable procedure.
            • Element Attributes

            • More or less this level is a playground for microformats utilizing the A element's REV and REL attributes.

Besides the common values "nofollow", "noindex", "noarchive"/"nocache" etc. and their omissible positive defaults "follow" and "index" etc., we'd need a couple more, for example "unapproved", "untrusted", "ignore" or "skip" and so on. There's a lot of work to do.

In terms of of complexity, a mechanism as outlined above should be as easy to use as CSS in combination with client sided scripting for visualization purposes.

However, whatever better ideas are out there, we need a widely accepted "Web-Robot Directives Standard" as soon as possible.



Tags: ()

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Saturday, February 17, 2007

Google going to revamp the rel=nofollow microformat?

I've asked Adam Lasnik, Google's search evangelist:

Adam, what is Google's take on extending the nofollow functionality by working out a microformat that covers the existing mechanism w/o being that unclear and confusing, and which takes care of similar needs like section targeting on element level and qualified votes as well?

and he answered

Sebastian, nothing's set in stone. Stuff is likely to evolve :)

That's an elating signal, thank you Adam. And it leads to a bunch of questions.

Will Google continue to cook nofollow in its secret sauce, revealing morphed semantics (affiliate links), unpopular areas of application (paid links) and changed functionality (no longer fetching the linked resource) every now and then? From my interpretation of Google's ongoing move to candidness I guess not.

Will Google gather a couple search companies to work out a new standard? I hope not, it would be a mistake not to involve content providers, webmasters, publishers, CMS vendors, even SEOs and opinion makers again.

Will Google ask for input? Will the process of defining a standard for micro crawler directives be an open and public discussion? Are we talking about an extended microformat, limited to the A element's rel and rev attributes, or does Google think of a broader approach covering for example section targeting and other crawler directives in class attributes on block level too? Will a new or more powerful interfere other norms like , , , or drafts like the not yet that comprehensive microformat (also badly named because it covers inclusion too)? By the way, the links above lead you to interesting thoughts on reach, functionality and implementation of an extended norm replacing nofollow, and I, like many of you, have a couple more ideas and concepts in mind.

I take Adam's tidbit as call for participation. Dear no-to-nofollow-sayers and nofollow-supporters out there, join the crowd at the white board! Throw in your thoughts, concepts, wishes and ideas.

In the meantime make use of this catalogue of do-follow plugins.



Tags: ()

Labels: , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Wednesday, January 31, 2007

Dear search engines, please bury the rel=nofollow-fiasko

The misuse of the rel=nofollow initiative is getting out of control. Invented to fight comment spam, nowadays it is applied to commercial links, biased editorial links, navigational links, links to worst enemies (funny example: Matt Cutts links to a SEO-Blackhat with rel=nofollow) and whatever else. Gazillions of publishers and site owners add it to their links for the wrong reasons, simply because they don't understand its intention, its mechanism, and especially not the ongoing morphing of its semantics. Even professional webmasters and search engine experts have a hard time to follow the nofollow-beast semantically. As more its initial usage gets diluted, as more folks suspect search engines cook their secret sauce with indigestibly nofollow-ingredients.

Not only rel=nofollow wasn't able to stop blog-spam-bots, it came with a build-in flaw: confusion.

Good news is that currently the nofollow-debate gets stoked again. Threadwatch hosts a thread titled Nofollow's Historical Changes and Associated Hypocrisy, folks are ranting on the questionable Wikipedia decision to nofollow all outbound links, Google video folks manipulated the PageRank algo by plastering most of their links with rel=nofollow by mistake, and even Yahoo's top gun Jeremy Zawodny is not that happy with the nofollow-debacle for a while now.

Say NO to NOFOLLOW - copyright jlh-design.comI say that it is possible to replace the unsuccessful nofollow-mechanism with an understandable and reasonable functionality to allow search engine crawler directives on link level. It can be done although there are shitloads of rel=nofollow links out there. Here is why, and how:

The value "nofollow" in the link's REL attribute creates misunderstandings, recently even in the inventor's company, because it is, hmmm, hapless.

In fact, back then it meant "passnoreputation" and nothing more. That is search engines shall follow those links, and they shall index the destination page, and they shall show those links in reversed citation results. They just must not pass any reputation or topical relevancy with that link.

There were micro formats better suitable to achieve the goal, for example Technorati's votelinks, but unfortunately the united search geeks have chosen a value adapted from the robots exclusion standard, which is plain misleading because it has absolutely nothing to do with its (intended) core functionality.

I can think of cases where a real nofollow-directive for spiders on link level makes perfect sense. It could tell the spider not to fetch a particular link destination, even if the page's robots tag says "follow", for example printer friendly pages. I'd use an "ignore this link" directive for example in crawlable horizontal popup menus to avoid theme dilution when every page of a section (or site) links to every other page. Actually, there is more need for spider directives on HTML element level, not only in links, for example to tag templated and/or navigational page areas like with Google's section targeting.

There is nothing wrong with a mechanism to neutralize links in user input. Just the value "nofollow" in the type-of-forward-relationship attribute is not suitable to label unchecked or not (yet) trusted links. If it is really necessary to adopt a well known value from the robots exclusion standard (and don't misunderstand me, reusing familiar terms in the right context is a good idea in general), the "noindex" value would have been be a better choice (although not perfect). "Noindex" describes way better what happens in a SE ranking algo: it doesn't index (in its technical meaning) a vote for the target. Period.

It is not too late to replace the rel=nofollow-fiasco with a better solution which could take care of some similar use cases too. Folks at Technorati, the W3C and whereever have done the initial work already, so it's just a tiny task left: extending an existing norm to enable a reasonable granularity of crawler directives on link level, or better for HTML elements at all. Rel=nofollow would get deprecated, replaced by suitable and standardized values, and for a couple years the engines could interpret rel=nofollow in its primordial meaning.

Since the rel=nofollow thingy exists, it has confused gazillions of non-geeky site owners, publishers and editors on the net. Last year I've got a new client who added rel=nofollow to all his internal links because he saw nofollowed links on a popular and well ranked site in his industry and thought rel=nofollow could perhaps improve his own rankings. That's just one example of many where I've seen intended as well as mistakenly misuse of the way too geeky nofollow-value. As Jill Whalen points out to Matt Cutts, that's just the beginning of net-wide nofollow-insane.

Ok, we've learned that the "nofollow" value is a notional monster, so can we please have it removed from the search engine algos in favour of a well thought out solution, preferably asap? Thanks.



Tags: ()

Labels: , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Thursday, January 19, 2006

Yahoo's handling of the link condom

Folks are wondering why nofollow-links are shown in Yahoo's backlink searches, site explorer results etc., and I'm wondering why the heck they're wondering.

First, that's not a new thing, the link condom has nothing to do with the ability to locate backlinks, so Yahoo always listed castrated citations and votes in link: and linkdomain: searches.

Second, there is absolutely nothing wrong with Yahoo's handling of rel=nofollow links. The value "nofollow" of the REL attribute creates misunderstandings, because it is, hmmm, hapless.

In fact, it means "passnoreputation" and nothing more. That is search engines shall follow those links, and they shall index the destination page, and they shall show those links in reversed citation results.

There were micro formats better suitable to achieve the goal, for example Technorati's votelinks, but unfortunately the search geeks have chosen a value adapted from the robots exclusion standard, which is plain misleading because it has absolutely nothing to do with its functionality.

So, since we now know that the "nofollow" value is a notional monster, can we please have it removed from the search engine algos asap? Thanks.

Tags: ()

Labels: ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Thursday, December 29, 2005

Is the spam condom efficient and ethical?

Jim Boykin from WeBuildPages raises a few very good questions in his 2-part-essay on link condoms in blog comments. Jim finally asks "Is the rel=nofollow our friend or our enemy?" and I've no definite answer.

If Blogger would allow me to opt out of the comment condom thingy I would do it with this blog. When I don't delete a comment containing a link, then the poster has something to say, and an embedded link doesn't deserve castration regardless whether I agree or not. Well, perhaps I'd unlink overdone URL drops in some cases.

If I would run a popular blog, I'd like a white-list approach best. That is every link in comments gets sterilized by default and all posts are pre-moderated, captchas in place. Trusted users could post instantly without link condom, and I could pull the condom from particular comments. I'm not aware of any blog software handling it this way, unfortunately.

Is the spam condom efficient? Nope. Comment moderation, captchas, spam filters, perhaps even registering users is enough to prevent a blog from comment spam. Also, many blogs run outdated, never updated pre-nofollow software, that is savvy spammers can still inject crappy links at enough places to keep it profitable.

Is the spam condom ethical? Nope. At least not when the blogger can't opt out. Not every comment is spam. Comments add content to a blog. Why penalize the content vendors?

Tags: without

Labels: ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->