Sebastian's Pamphlets

If you've read my articles somewhere on the Internet, expect something different here.

MOVED TO SEBASTIANS-PAMPHLETS.COM

Please click the link above to read actual posts, this archive will disappear soon!

Stay tuned...

Wednesday, August 15, 2007

Google's 5 sure-fire steps to safer indexing

Nofollow plagueAre you wondering why Gray Hat Search Engine News (GHN) is so quiet recently? One reason may be that I've borrowed their Google savvy spy. I've sent him to Mountain View again to learn more about Google's nofollow strategy.

He returned with a copy of Google's recently revised mission statement, discovered in the wastebasket of a conference room near office 211 in building 43. Read the shocking and unbelievable head note printed in bold letters:

Google's mission is to condomize the world's information and make it universally uncrawlable and useless.

Read and reread it, then some weird facts begin to make sense. Now you'll understand why:

  1. The rel-nofollow plague was designed to maximize collateral damage by devaluing all hyperlinked votes by honest users of nearly all platforms you're using everyday, for example Twitter, Wikipedia, corporate blogs, GoogleGroups ... ostensibly to nullify the efforts of a few spammers.
  2. Nobody bothers to comment on your nofollow'ed blog.
  3. Google invented the supplemental index (to store scraped resources suffering from too many condomized links) and why it grows faster than the main index.
  4. Google installed the Bigdaddy infrastructure (to prevent Ms. Googlebot from following nofollow'ed links).
  5. Google switched to BlitzCrawling (to list timely contents for a moment whilst fat resources from large archives get buried in the supplemental index). RIP deep crawler and freshbot.

Seriously, the deep crawler isn't defunct, it's called supplemental crawler nowadays, and the freshbot is still alive as Feedfetcher.




Disclaimer: All these hard facts were gathered by torturing sources close to Google, robbery and other unfair methods. If anyone bothers to debunk all that as bad joke, one question still remains: Why does Google next to nothing to stop the nofollow plague? I mean, ongoing mass abuse of rel-nofollow is obviously counterproductive with regard to their real mission.

Labels: , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Wednesday, August 08, 2007

Google manifested the axe on reciprocal link exchanges

Yesterday Fantomaster via Threadwatcher pointed me to this page of Google's Webmaster help system. The cache was a few days old and didn't show a difference, I don't archive each and every change of the guidelines, so I asked and a friendly and helpful Googler told me that this item was around for a while now. Today this page made it on Sphinn and probably a few other Webmaster hangouts too.

So what the heck is the scandal all about? When you ask Google for help on "link exchange", the help machine rattles for a second, sighs, coughs, clears its throat and then yells out the answer in bold letters: "Link schemes", bah!

Ok, we already knew what Google thinks about artificial linkage: "Don't participate in link schemes designed to increase your site's ranking or PageRank". Honestly, what is the intent when I suggest that you link to me and concurrently I link to you? Yup, it means I boost your PageRank and you boost mine, also we chose some nice anchor text and that makes the link deal perfect. In the eyes of Google even such a tiny deal is a link scheme, because both links weren't put up for users but for search engines.

Pre-Google this kind of link deal was business as usual and considered natural, but frankly back then the links were exchanged for traffic and not for search engine love. We can rant and argue as much as we want, that will not revert the changed character of link swaps nor Google's take on manipulative links.

Consequently Google has devalued artificial reciprocal links for ages. Pretty much simplified these links nullify each other in Google's search index. That goes for tiny sins. Folks raising the concept onto larger link networks got caught too but penalized or even banned for link farming.

Obviously all kinds of link swaps are easy to detect algorithmically, even triangular link deals, three way link exchanges and whatnot. I called that plain vanilla link 'swindles', but only just recently Google has caught up with a scalable solution and seems to detect and penalize most if not all variants covering the whole search index, thanks to the search quality folks in Dublin and Zurich even overseas in whatever languages.

The knowledge that the days of free link trading are numbered was out for years before the exodus. Artificial reciprocal links as well as other linkage considered link spam by Google was and is a pet peeve of Matt's team. Google sent lots of warnings, and many sane SEOs and Webmasters heard their traffic master's voice and acted accordingly. Successful link trading just went underground leaving the great unwashed alone with their obsession about exchanging reciprocal links in the public.

Also old news is, that Google does not penalize reciprocal links in general. Google almost never penalizes a pattern or a technique. Instead they try to figure out the Webmaster's intent and judge case by case based on their findings. And yes, that's doable with algos, perhaps sometimes with a little help from humans to compile the seed, but we don't know how perfect the algo is when it comes to evaluations of intent. Natural reciprocal links are perfectly fine with Google. That applies to well maintained blogrolls too, despite the often reciprocal character of these links. Reading the link schemes page completely should make that clear.

Google defines link scheme as "[...] Link exchange and reciprocal links schemes ('Link to me and I'll link to you.') [...]". The "I link to you and vice versa" part literally addresses link trading of any kind, not a situation where I link to your compelling contents because I like a particular page, and you return the favour later on because you find my stuff somewhat useful. As Perkiset puts it "linking is now supposed to be like that well known sex act, '68? - or, you do me and I'll owe you one'" and there is truth in this analogy. Sometimes a favor will not be returned. That's the way the cookie crumbles when you're keen on Google traffic.


The fact that Google openly said that link exchange schemes designed "exclusively for the sake of cross-linking" of any kind violate their guidelines indicates that first they were sure to have invented the catchall algo, and second that they felt safe to launch it without too much collateral damage. Not everybody agrees, I quote Fantomaster's critique not only because I like his inimitably parlance:
This is essentially a theological debate: Attempting to determine any given action's (and by inference: actor's) "intention" (as in "sinning") is always bound to open a can of worms or two.

It will always have to work by conjecture, however plausible, which makes it a fundamentally tacky, unreliable and arbitrary process.

The delusion that such a task, error prone as it is even when you set the most intelligent and well informed human experts to it (vide e.g. criminal law where "intention" can make all the difference between an indictment for second or first degree murder...) can be handled definitively by mechanistic computer algorithms is arguably the most scary aspect of this inane orgy of technological hubris and naivety the likes of Google are pressing onto us.
I've seen some collateral damage already, but pragmatic Webmasters will find --respectively have found long ago-- their way to build inbound links under Google's regime.

And here is the context of Google's definition link exchanges = link schemes which makes clear that not each and every reciprocal link is evil:
[…] However, some webmasters engage in link exchange schemes and build partner pages exclusively for the sake of cross-linking, disregarding the quality of the links, the sources, and the long-term impact it will have on their sites. This is in violation of Google’s webmaster guidelines and can negatively impact your site’s ranking in search results. Examples of link schemes can include:

• Links intended to manipulate PageRank
• Links to web spammers or bad neighborhoods on the web
• Link exchange and reciprocal links schemes ('Link to me and I'll link to you.')
• Buying or selling links [...]
Again, please read the whole page.


Bear in mind that all this is Internet history, it just boiled up yesterday as the help page was discovered.


Related article: Eric Ward on reciprocal links, why they do good, and where they do bad.

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Tuesday, August 07, 2007

NOPREVIEW - The missing X-Robots-Tag

Google provides previews of non-HTML resources listed on their SERPs:View PDF as HTML document
These "view as text" and "view as HTML" links are pretty useful when you for example want to scan a PDF document before you clutter your machine's RAM with 30 megs of useless digital rights management (aka Adobe Reader). You can view contents even when the corresponding application is not installed, Google's transformed previews should not stuff your maiden box with unwanted malware, etcetera. However, under some circumstances it would make sound sense to have a NOPREVIEW X-Robots-Tag, but unfortunately Google forgot to introduce it yet.

Google is rightfully proud of their capability to transform various file formats to readable HTML or plain text: Adobe Portable Document Format (pdf), Adobe PostScript (ps), Lotus 1-2-3 (wk1, wk2, wk3, wk4, wk5, wki, wks, wku), Lotus WordPro (lwp), MacWrite (mw), Microsoft Excel (xls), Microsoft PowerPoint (ppt), Microsoft Word (doc), Microsoft Works (wks, wps, wdb), Microsoft Write (wri), Rich Text Format (rtf), Shockwave Flash (swf), of course Text (ans, txt) plus a couple of "unrecognized" file types like XML. New formats are added from time to time.

According to Adam Lasnik currently there is no way for Webmasters to tell Google not to include the "View as HTML" option. You can try to fool Google's converters by messing up the non-HTML resource in a way that a sane parser can't interpret it. Actually, when you search a few minutes you'll find e.g. PDF files without the preview links on Google's SERPs. I wouldn't consider this attempt a bullet-proof nor future-proof tactic though, because Google is pretty intent on improving their conversion/interpretation process.

I like the previews not only because sometimes they allow me to read documents behind a login screen. That's a loophole Google should close as soon as possible. When for example PDF documents or Excel sheets are crawlable but not viewable for searchers (at least not with the second click) that's plain annoying both for the site as well as for the search engine user.

With HTML documents the Webmaster can apply a NOARCHIVE crawler directive to prevent non paying visitors from lurking via Google's cached page copies. Thanks to the newish REP header tags one can do that with non-HTML resources too, but neither NOARCHIVE nor NOSNIPPET etch away the "view-as HTML" link.

<speculation>Is the lack of a NOPREVIEW crawler directive just an oversight, or is it stuck in the pipeline because Google is working on supplemental components and concepts? Google's yet inconsistent handling of subscription content comes to mind as an ideal playground for such a robots directive in combination with a policy change.</speculation>

Anyways, there is a need for a NOPREVIEW robots tag, so why not implement it now? Thanks in advance.

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Tuesday, July 31, 2007

Handling Google's neat X-Robots-Tag - Sending REP header tags with PHP

It's a bad habit to tell the bad news first, and I'm guilty of that. Yesterday I linked to Dan Crow telling Google that the unavailable_after tag is useless IMHO. So todays post is about a great thing: REP header tags aka X-Robots-Tags, unfortunately mentioned as second news somewhat concealed in Google's announcement.

The REP is not only a theatre, it stands for Robots Exclusion Protocol (robots.txt and robots meta tag). Everything you can shove into a robots meta tag on a HTML page can now be delivered in the HTTP header for any file type:
  • INDEX|NOINDEX - Tells whether the page may be indexed or not
  • FOLLOW|NOFOLLOW - Tells whether crawlers may follow links provided on the page or not
  • ALL|NONE - ALL = INDEX, FOLLOW (default), NONE = NOINDEX, NOFOLLOW
  • NOODP - tells search engines not to use page titles and descriptions from the ODP on their SERPs.
  • NOYDIR - tells Yahoo! search not to use page titles and descriptions from the Yahoo! directory on the SERPs.
  • NOARCHIVE - Google specific, used to prevent archiving (cached page copy)
  • NOSNIPPET - Prevents Google from displaying text snippets for your page on the SERPs
  • UNAVAILABLE_AFTER: RFC 850 formatted timestamp - Removes an URL from Google's search index a day after the given date/time

So how can you serve X-Robots-Tags in the HTTP header of PDF files for example? Here is one possible procedure to explain the basics, just adapt it for your needs:

Rewrite all requests of PDF documents to a PHP script knowing wich files must be served with REP header tags. You could do an external redirect too, but this may confuse things. Put this code in your root's .htaccess:

RewriteEngine On
RewriteBase /pdf
RewriteRule ^(.*)\.pdf$ serve_pdf.php

In /pdf you store some PDF documents and serve_pdf.php:

...
$requestUri = $_SERVER['REQUEST_URI'];
...
if (stristr($requestUri, "my.pdf")) {
header('X-Robots-Tag: index, noarchive, nosnippet', TRUE);
header('Content-type: application/pdf', TRUE);
readfile('my.pdf');
exit;
}
...

This setup routes all requests of *.pdf files to /pdf/serve_pdf.php which outputs something like this header when a user agent asks for /pdf/my.pdf:

Date: Tue, 31 Jul 2007 21:41:38 GMT
Server: Apache/1.3.37 (Unix) PHP/4.4.4
X-Powered-By: PHP/4.4.4
X-Robots-Tag: index, noarchive, nosnippet
Connection: close
Transfer-Encoding: chunked
Content-Type: application/pdf

You can do that with all kind of file types. Have fun and say thanks to Google :)

Labels: , , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Monday, July 30, 2007

Unavailable_After is totally and utterly useless

I've a lot of respect for Dan Crow, but I'm struggling with my understanding, or possible support, of the unavailable_after tag. I don't want to put my reputation for bashing such initiatives from search engines at risk, so sit back and grab your popcorn, here comes the roasting:

As a Webmaster, I did not find a single scenario where I could or even would use it. That's because I'm a greedy traffic whore. A bazillion other Webmasters are greedy too. So how the heck is Google going to sell the newish tag to the greedy masses?

Ok, from a search engine's perspective unavailable_after makes sound sense. Outdated pages bind resources, annoy searchers, and in a row of useless crap the next bad thing after an outdated page is intentional Webspam.

So convincing the great unwashed to put that thingy on their pages inviting friends and family to granny's birthday party on 25-Aug-2007 15:00:00 EST would improve search quality. Not that family blog owners care about new meta tags, RFC 850-ish date formats, or search engine algos rarely understanding that the announced party is history on Aug/26/2007. Besides there may be painful aftermaths worth submitting a desperate call for aspirins the day after in the comments, what would be news of the day after expiration. Kinda dilemma, isn't it?

Seriously, unless CMS vendors support the new tag, tiny sites and clique blogs aren't Google's target audience. This initiative addresses large sites which are responsible for a huge amount of outdated contents in Google's search index.

So what is the large site Webmaster's advantage of using the unavailable_after tag? A loss of search engine traffic. A loss of link juice gained by the expired page. And so on. Losses of any kind are not that helpful when it comes to an overdue raise nor in salary negotiations. Hence the Webmaster asks for the sack when s/he implements Google's traffic terminator.

Who cares about Google's search quality problems when it leads to traffic losses? Nobody. Caring Webmasters do the right thing anyway. And they don't need no more useless meta tags like unavailable_after. "We don't need no stinking metas" from "Another Brick in the Wall Part Web 2.0" expresses my thoughts perfectly.

So what separates the caring Webmaster from the 'ruthless traffic junky' who Google wants to implement the unavailable_after tag? The traffic junkie lets his stuff expire without telling Google about it's state, is happy that frustrated searchers click the URL from the SERPs even years after the event, and enjoys the earnings from tons of ads placed above the content minutes after the party was over. Dear Google, you can't convince this guy.

[It seems this is a post about repetitive "so whats". And I came to the point before the 4th paragraph ... wow, that's new ... and I've put a message in the title which is not even meant as link bait. Keep on reading.]


So what does the caring Webmaster do without the newish unavailable_after tag? Business as usual. Examples:


Say I run a news site where the free contents go to the subscription area after a while. I'd closely watch which search terms generate traffic, write a search engine optimized summary containing those keywords, put that on the sales pitch, and move the original article to the archives accessible to subscribers only. It's not my fault that the engines think they point to the original article after the move. When they recrawl and reindex the page my traffic will increase because my summary fits their needs more perfectly.

Say I run an auction site. Unfortunately particular auctions expire, but I'm sure that the offered products will return to my site. Hence I don't close the page, but I search my database for similar offerings and promote them under a H3 heading like "[product] (stuffed keywords) is hot" /H3 P buy [product] here: /P followed by a list of identical products for sale or similar auctions.

Say I run a poll expiring in two weeks. With Google's newish near real time indexing that's enough time to collect keywords from my stats, so the textual summary under the poll's results will attract the engines as well as visitors when the poll is closed. Also, many visitors will follow the links to related respectively new polls.


From Google's POV there's nothing wrong with my examples, because the visitor gets what s/he was searching for, and I didn't cheat. Now tell me, why should I give up these valuable sources of nicely targeted search engine traffic just to make Google happy? Rather I'd make my employer happy. Dear Google, you didn't convince me.



Update: Tanner Christensen posted a remarkable comment at Sphinn:
I'm sure there is some really great potential for the tag. It's just none of us have a need for it right now.

Take, for example, when you buy your car without a cup holder. You didn't think you would use it. But then, one day, you find yourself driving home with three cups of fruit punch and no cup holders. Doh!

I say we wait it out for a while before we really jump on any conclusions about the tag.
John Andrews was the first to report an evil use of unavailable_after.

Also, Dan Crow from Google announced a pretty neat thing in the same post: With the X-Robots-Tag you can now apply crawler directives valid in robots meta tags to non-HTML documents like PDF files or images.

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Saturday, July 28, 2007

Analyzing search engine rankings by human traffic

Recently I've discussed ranking checkers at several places, and I'm quite astonished that folks still see some value in ranking reports. Frankly, ranking reports are --in most cases-- a useless waste of paper and/or disk space. That does not mean that SERP positions per keyword phrase aren't interesting. They're just useless without context, that is traffic data. Converting traffic pays the bills, not sole rankings. The truth is in your traffic data.

That said, I'd like to outline a method to get a particular useful information out of raw traffic data: underestimated search terms. That's not a new attempt, and perhaps you have the reports already, but maybe you don't look at the information which is somewhat hidden in stats ordered by success, not failure. And you should be --respective employ-- a programmer to implement it.


The first step is gathering data. Create a database table to record all hits, then in a footer include or so, when the complete page got outputted already, write all data you have in that table. All data means URL, timestamp, and variables like referrer, user agent, IP, language and so on. Be a data rat, log everything you can get hold of. With dynamic sites it's easy to add page title, (product) IDs etcetera, with static sites write a tool to capture these attributes separately.

For performance reasons it makes sense to work with a raw data table, which has just a primary key, to log the requests, and normalized working tables which have lots of indexes to allow aggregations, ad hoc queries, and fast reports from different perspectives. Also think of regular purging the raw log table and historization. While transferring raw log data to the working tables in low traffic hours or on another machine you can calculate interesting attributes and add data from other sources which were not available to the logging process.

You'll need that traffic data collector anyway for a gazillion of purposes where your analytics software fails, is not precise enough, or just can't deliver a particular evaluation perspective. It's a prerequisite for the method discussed here, but don't build a monster sized cannon to chase a fly. You can gather search engine referrer data from logfiles too.


For example an interesting information is on which SERP a user clicked a link pointing to your site. Simplified you need three attributes in your working tables to store this info: search engine, search term, and SERP number. You can extract these values from the HTTP_REFERER.

http://www.google.com/search?q=keyword1+keyword2~
&ie=utf-8&oe=utf-8&aq=t&rls=org.mozilla:en-US:official&client=firefox-a

1. "google" in the server name tells you the search engine.
2. The "q" variable's value tells you the search term "keyword1 keyword2".
3. The lack of a "start" variable tells you that the result was placed on the first SERP. The lack of a "num" variable lets you assume that the user got 10 results per SERP, so it's quite safe to say that you rank in the top 10 for this term. Actually, the number of results per page is not always extractable from the URL because it's pulled from a cookie usually, but not so many surfers change their preferences (e.g. less than 0.5% surf with 100 results according to JohnMu and my data as well). If you've got a "num" value then add 1 and divide the result by 10 to make the data comparable. If that's not precise enough you'll spot it afterwards, and you can always recalculate SERP numbers from the canned referrer.

http://www.google.co.uk/search?q=keyword1+keyword2~
&hl=en&start=10&sa=N

1. and 2. as above.
3. The "start" variable's value 10 tells you that you got a hit from the second SERP. When start=10 and there is no "num" variable, most probably the searcher got 10 results per page.

http://www.google.es/search?q=keyword1+keyword2~
&rls=com.microsoft:*&ie=UTF-8&oe=UTF-8&startIndex=~
&startPage=1

1. and 2. as above.
3. The empty "startIndex" variable and startPage=1 are useless, but the lack of "start" and "num" tells you that you've got a hit from the 1st spanish SERP.

http://www.google.ca/search?q=keyword1+keyword2~
&hl=en&rls=GGGL,GGGL:2006-30,GGGL:en&start=20~
&num=20&sa=N

1. and 2. as above.
3. num=20 tells you that the searcher views 20 results per page, and start=20 indicates the second SERP, so you rank between #21 and #40, thus the (averaged) SERP# is 3.5 (provided SERP# is not an integer in your database).

You got the idea, here is a cheat sheet and official documentation on Google's URL parameters. Analyze the URLs in your referrer logs and call them with cookies off what disables your personal search preferences, then play with the values. Do that with other search engines too.


Now a subset of your traffic data has a value in "search engine". Aggregate tuples where search engine is not NULL, then select the results for example where SERP number is lower or equal 3.99 (respectively 4), ordered by SERP number ascending, hits descending and keyword phrase, break by search engine. (Why sorted by traffic descending? You have a report of your best performing keywords already.)

The result is a list of search terms you rank for on the first 4 SERPs, beginning with keywords you've probably not optimized for. At least you didn't optimize the snippet to improve CTR, so your ranking doesn't generate a reasonable amount of traffic. Before you study the report, throw away your site owner hat and try to think like a consumer. Sometimes those make use of a vocabulary you didn't think of before.

Research promising keywords, and decide whether you want to push, bury or ignore them. Why bury? Well, in some cases you just don't want to rank for a particular search term, [your product sucks] being just one example. If the ranking is fine, the search term smells somewhat lucrative, and just the snippet sucks in a particular search query's context, enhance your SERP listing.

Every once in a while you'll discover a search term making a killing for your competitors whilst you never spotted it because your stats package reports only the best 500 monthly referrers or so. Also, you'll get the most out of your rankings by optimizing their SERP CTRs.


Be crative, over time your traffic database becomes more and more valuable, allowing other unconventional and/or site specific reports which off-the-shelf analytics software usually does not deliver. Most probably your competitors use standard analytics software, individually developed algos and reports can make a difference. That does not mean you should throw away your analytics software to reinvent the wheel. However, once you're used to self developed analytic tools you'll think of more interesting methods not only to analyse and monitor rankings by human traffic than you can implement in this century ;)


Bear in mind that the method outlined above does not and cannot replace serious keyword research.


Another --very popular-- approach to get this info would be automated ranking checks mashed up with hits by keyword phrase. Unfortunately, Google and other engines do not permit automated queries for the purpose of ranking checks, and this method works with preselected keywords, that means you don't find (all) search terms created by users. Even when you compile your ranking checker's keyword lists via various keyword research tools, you'll still miss out on some interesting keywords in your seed list.


Related thoughts: Why regular and automated ranking checks are necessary when you operate seasonal sites by Donna

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Friday, July 27, 2007

Rediscover Google's free ranking checker!

Nowadays we're searching via toolbar, personalized homepage, or in the browser address bar by typing in "google" to get the search box, typing in a search query using "I feel lucky" functionality, or -my favorite- typing in google.com/search?q=free+pizza+service+nearby.

Old fashioned, uncluttered and nevertheless sexy user interfaces are forgotten, and pretty much disliked due to the lack of nifty rounded corners. Luckily Google still maintains them. Look at this beautiful SERP:
Google's free ranking checker
It's free of personalized search, wonderful uncluttered because the snippets appear as tooltip only, results are nicely numbered from 1 to 1,000 on just 10 awesome fast loading pages, and when I've visited my URLs before I spot my purple rankings quickly.

http://google.com/ie?num=100&q=keyword1+keyword2 is an ideal free ranking checker. It supports &filter=0 and other URL parameters, so it's a perfect tool when I need to lookup particular search terms.

Mass ranking checks are totally and utterly useless, at least for the average site, and penalized by Google. Well, I can think of ways to semi-automate a couple queries, but honestly, I almost never need that. Providing fully automated ranking reports to clients gave SEO services a more or less well deserved snake oil reputation, because nice rankings for preselected keywords may be great ego food, but they don't pay the bills. I admit that with some setups automated mass ranking checks make sense, but those are off-topic here.

By the way, Google's query stats are a pretty useful resource too.

Labels: , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Wednesday, July 25, 2007

Blogger to rule search engine visibility?

Via Google's Webmaster Forum I found this curiosity:
http://www.stockweb.blogspot.com/robots.txt
User-agent: *
Disallow: /search
Disallow: /
A standard robots.txt at *.blogspot.com looks different:
User-agent: *
Disallow: /search
Sitemap: http://*.blogspot.com/feeds/posts/default?orderby=updated

According to the blogger the blog is not private, what would explain the crawler blocking:
It is a public blog. In the past it had a standard robots.txt, but 10 days ago it changed to "Disallow: /"

Copyscape thinks that the blog in question shares a fair amount of content with other Web pages. So does blog search:
http://stockweb.blogspot.com/2007/07/ukraine-stock-index-pfts-gained-97-ytd.html
has a duplicate, posted by the same author, at
http://business-house.net/nokia-nok-gains-from-n-series-smart-phones/,
http://stockweb.blogspot.com/2007/07/prague-energy-exchange-starts-trading.html
is reprinted at
http://business-house.net/prague-energy-exchange-starts-trading-tomorrow/
and so on. Probably a further investigation would reveal more duplicated contents.

It's understandable that Blogger is not interested in wasting Google's resources by letting Ms. Googlebot crawl the same contents from different sources. But why do they block other search engines too? And why do they block the source (the posts reprinted at business-house.net state "Originally posted at [blogspot URL]")?

Is this really censorship, or just a software glitch, or is it all the blogger's fault?


Update 07/26/2007: The robots.txt reverted to standard contents for unknown reasons. However, with a shabby link neigborhood as expressed in the blog's footer I doubt the crawlers will enjoy their visits. At least the indexers will consider this sort of spider fodder nauseous.

Labels: , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Monday, July 16, 2007

Getting the most out of Google's 404 stats

The 404 reports in Google's Webmaster Central panel are great to debug your site, but they contain URLs generated by invalid --respectively truncated-- URL drops or typos of other Webmasters too. Are you sick of wasting the link love from invalid inbound links, just because you lack a suitable procedure to 301-redirect all these 404 errors to canonical URLs?

Your pain ends here. At least when you're on a *ix server running Apache with PHP 4+ or 5+ and .htaccess enabled. (If you suffer from IIS go search another hobby.)

I've developed a tool which grabs all 404 requests, letting you map a canonical URL to each 404 error. The tool captures and records 404s, and you can add invalid URLs from Google's 404-reports, if these aren't recorded (yet) from requests by Ms. Googlebot.

It's kinda layer between your standard 404 handling and your error page. If a request results in a 404 error, your .htaccess calls the tool instead of the error page. If you've assigned a canonical URL to an invalid URL, the tool 301-redirects the request to the canonical URL. Otherwise it sends a 404 header and outputs your standard 404 error page. Google's 404-probe requests during the Webmaster Tools verification procedure are unredirectable (is this a word?).

Besides 1:1 mappings of invalid URLs to canonical URLs you can assign keywords to canonical URLs. For example you can define that all invalid requests go to /fruit when the requested URI or the HTTP referrer (usually a SERP) contain the strings "apple", "orange", "banana" or "strawberry". If there's no persistent mapping, these requests get 302-redirected to the guessed canonical URL, thus you should view the redirect log frequently to find invalid URLs which deserve a persistent 301-redirect.

Next there are tons of bogus requests from spambots searching for exploits or whatever, or hotlinkers, resulting in 404 errors, where it makes no sense to maintain URL mappings. Just update an ignore list to make sure those get 301-redirected to example.com/goFuckYourself or a cruel and scary image hosted on your domain or a free host of your choice.

Everything not matching a persistent redirect rule or an expression ends up in a 404 response, as before, but logged so that you can define a mapping to a canonical URL. Also, you can use this tool when you plan to change (a lot of) URLs, it can 301-redirect the old URL to the new one without adding those to your .htaccess file.

I've tested this tool for a while on a couple of smaller sites and I think it can get trained to run smoothly without too many edits once the ignore lists etcetera are up to date, that is matching the site's requisites. A couple of friends got the script and they will provide useful input. Thanks! If you'd like to join the BETA test drop me a message.

Disclaimer: All data get stored in flat files. With large sites we'd need to change that to a database. The UI sucks, I mean it's usable but it comes with the browser's default fonts and all that. IOW the current version is still in the stage of "proof of concept". But it works just fine ;)

Labels: , , , , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Wednesday, July 11, 2007

Google helps those who help themselves

And if that's not enough to survive on Google's SERPs, try Google's Webmaster Forum where you can study Adam Lasnik's FAQ which covers even questions the Webmaster Help Center provides no comprehensive answer for (yet), and where Googlers working in Google's Search Quality, Webspam, and Webmaster Central teams hang out. Google dumps all sorts of questioners to the forum, where a crowd of hardcore volunteers (aka regulars as Google calls them) invests a lot of time to help out Webmasters and site owners facing problems with the almighty Google.

Despite the sporadic posts by Googlers, the backbone of Google's Webmaster support channel is this crew of regulars from all around the globe. Google monitors the forum for input and trends, and intervenes when the periodic scandal escalates every once in a while. Apropos scandal ... although the list of top posters mentions a few of the regulars, bear in mind that trolls come with a disgusting high posting cadency. Fortunately, currently the signal drowns the noise (again), and I appreciate very much that the Googlers participate more and more.

Some of the regulars like seo101 don't reveal their URLs and stay anonymous. So here is an incomplete list of folks giving good advice:If I've missed anyone, please drop me a line (I stole the list above from JLH and Red Cardinal, so it's all their fault!).

So when you're a Webmaster or site owner, don't hesitate to post your Google related question (but read the FAQ before posting, and search for your topics), chances are one of these regulars or even a Googler offers assistance. Otherwise when you're questionless carrying a swag of valuable answers, join the group and share your knowledge. Finally, when you're a Googler, donate the sites linked above a boost on the SERPs ;)



Micro-meme started by John Honeck, supported by Richard Hearne, Bert Vierstra ...

Labels: , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Thursday, July 05, 2007

Why eBay and Wikipedia rule Google's SERPs

It's hard to find an obscure search query like [artificial link] which doesn't deliver eBay spam or a Wikipedia stub within the first few results at Google. Although both Wikipedia and eBay are large sites, the Web is huge, so two that different sites shouldn't dominate the SERPs for that many topics. Hence it's safe to say that many nicely ranked search results at Googledia, pulled from eBaydia, are plain artificial positioned non-results.

Curious why my beloved search engine fails so badly, I borrowed a Google-savvy spy from GHN and sent him to Mountain View to uncover the eBaydia ranking secrets. He came back with lots of pay-dirt scraped from DVDs in the safe of building 43. Before I sold Google's ranking algo to Ask (the price Yahoo! and MSN offered was laughable), I figured out why Googledia prefers eBaydia from comments in the source code. Here is the unbelievable story of a miserable failure:

When Yahoo! launched Mindset, Larry Page and Sergey Brin threw chairs out of anger because Google wasn't able to accomplish such a simple task. The engineers, eager to fulfill their founder's wishes asap, tried to integrate mindset-functionality without changing Google's fascinating simple search interface (that means without a shopping/research slider). Personalized search still lived in the labs, but provided a somewhat suitable API (mega beta): scanSearchersBrainForContext([search query]). Not knowing that this function of personalized search polls a nano-bugging-device (pre alpha) which Google had not yet released nor implemented into any searcher's brain at this time, they made use of that piece of experimental code to evaluate the search query's context. Since the method always returned "false", though they had to deliver results quickly, they made up some return values to test their algo tweaks:

/* debug - praying S&L don't throw more chairs */
if (scanSearchersBrainForContext($searchQuery) === false) then {
$contextShopping = "%ebay%";
$contextResearch = "%wikipedia%";
$context = both($contextShopping, $contextResearch);
}
else {[pretty complex algo])


This worked fine and found its way into the ranking algo under time pressure. The result is that with each and every search query where a page from eBay and/or Wikipedia is in the raw result set, those get a ranking boost. Sergey was happy because eBay is generally listed on page #1, and Larry likes the Wikipedia results on the first SERP. Tell me why the heck should the engineers comment out these made up return values? No engineer on this planet likes flying chairs, especially not in his office.


PS: Some SEOs push Wikipedia stubs too.

Labels: , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Who is responsible for the paid link mess?

Look at this graph showing the number of [buy link] searches since 2004:

Interestingly this search term starts out in September or October 2004, and shows a quite stable trend until the recent paid links debate started.

Who or what caused SEOs to massively buy links since 2004?
  • The Playboy interview with Google cofounders Larry Page and Sergey Brin just before Google was about to go public?
  • Google's IPO?
  • Rumors that Google ran out of index space and therefore might restrict the number of doorway pages in the search index?
  • Nick Wilson preparing the launch of Threadwatch?
  • AdWords and Overture no longer running gambling ads?
  • The Internet Advancement scandal?
  • Google's shortage of beer at the SES Google dance?
  • A couple UK based SEOs invented bought organic rankings?

Seriously, buying links for rankings was an established practice way before 2004. If you know the answer, or if you've a somewhat plausible theory, leave it in the comments. I'm really curious. Thanks.

Labels: , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Wednesday, July 04, 2007

Google assists SERP Click-Through Optimization

Big Mama Google in her ongoing campaign to keep her search index clean assists Webmasters with reports allowing click-trough optimization of a dozen or so pages per Web site. Google launched these reports a while ago, but most Webmasters didn't make the best use of them. Now that Vanessa has revealed her SEO secrets, lets discuss why and how Google helps increasing, improving, and targeting search engine traffic.

Google is not interested in gazillions of pages which rank high for (obscure) search terms but don't get clicked from the SERPs. This clutter tortures the crawler and indexer, and it wastes expensive resources the query engine could use to deliver better results to the searchers.

Unfortunately, legions of clueless SEOs work hard to increase mount clutter by providing their clients with weekly ranking reports, what leads to even more pages which rank for (potentially money making) search phrases but appear on the SERPs with such crappy titles and snippets that not even a searcher coming with an IQ slightly below a slice of bread clicks them.

High rankings don't pay the bills, converting traffic from SERPs on the other hand does. A nicely ranking page is an asset, which in most cases just needs a few minor tweaks to attract search engine users (mount clutter contains machine generated cookie-cutter pages too, but that's a completely other story).

For example unattended pages gaining their SERP position from anchor text of links pointing to them often have a crappy click through rate (CTR). Say you've a page about a particular aspect of green widgets, which applies to widgets of all colors. For some reason folks preferring red widgets like your piece and link to it with "red widgets" as anchor text. The page will rank fine for [red widgets], but since "red widgets" is not mentioned on the page this keyword phrase doesn't appear on the SERP's snippets, not to speak of the linked title. Search engine users seeking for information on red widgets don't click the link about green widgets, although it might be the best matching search result.

So here is the click-thru optimization process based on Google's query stats (it doesn't work with brand new sites nor more or less unindexed sites, because the data provided in Google's Webmaster Tools are available, reliable and quite accurate for somewhat established sites only):

Login, choose a site and go to query stats. In an ideal world you'll see two tables of rather identical keyword lists (all examples made up).
Top search queriesAvg.
Pos.
Top SERP clicksAvg.
Pos.
1. web site design51. web site design4
2. google consulting42. seo consulting5
3. seo consulting33. google consulting2
4. web site structures24. internal links3
5. internal linkage15. web site structure3
6. crawlability36. crawlability5

The "Top search queries" table on the left shows positions for search phrases on the SERPs, regardless whether these pages got clicks or not. The "Top search query clicks" table on the right shows which search terms got clicked most, and where the landing pages were positioned on their SERPs. If good keywords appear in the left table but not in the right one, you've CTR optimization potentials.

The "average top position" might differ from todays SERPs, and it might differ for particular keywords even if those appear in the same line in both tables. Positioning fluctuation depends on a couple of factors. First, the position is recorded at the run time of each search query during the last 7 days, and within seven days a page can jump up and down on the SERPs. Second, positioning on for example UK SERPs can differ from US SERPs, so an average 3rd position may be a utterly useless value, when a page ranks #1 in the UK and gets a fair amount of traffic from UK SERPs, but ranks #8 on US SERPs and searchers don't click it because the page is about a local event near Loch Nowhere in the highlands. Hence refine the reports by selecting your target markets in "location", and if necessary "search type" too. Third, if these stats are generated based on very few searches and even fewer click throughs, they are totally and utterly useless for optimization purposes.

Lets say you've got a site with a fair amount of Google search engine traffic, the next step is identifying the landing pages involved (you get only 20 search queries, so the report covers only a fraction of your site's pages). Pull these data from your referrer stats, or extract SERP referrers from your logs to create a crosstab of search terms from Google's reports per landing page. Although the click data are from Google's SERPs, it might make sense to do this job with a broader scope, that is including referrers from all major search engines.

Now perform the searches for your 20 keyword phrases (just click on the keywords on the report) to check how your pages look at the SERPs. If particular landing pages trigger search results for more than one search term, extract them all. Then load your landing page, and view its source. Read your page first rendered in your browser, then check out semantic hints in the source code, for example ALT or TITLE text and stuff like that. Look at the anchor text of incoming links (you can use link stats and anchor text stats from Google, We Build Pages Tools, ...) and other ranking factors to understand why Google thinks this page is a good match for the search term. For each page, let the information sink before you change anything.


If the page is not exactly a traffic generator for other targeted keywords, you can optimize it with regard to a better CTR for the keyword(s) it ranks for. Basically that means use the keyword(s) naturally on all page areas where it makes sense, and provide each occurence with a context which hopefully makes it into the SERP's snippet.

Make up a few natural sentences a searcher might have in mind when searching for your keyword(s). Write them down. Order them by their ability to fit the current page text in a natural way. Bear in mind that with personalized search Google could have scanned the searcher's brain to add different contexts to the search query, so don't concentrate too much on the keyword phrase alone, but on short sentences containing both the keyword(s), respectively their synonyms, and a sensible context as well.

There is no magic number like "use the keywords 5 times to get a #3 spot" or "7 occurences of a keyword gain you a #1 ranking". Optimal keyword density is a myth, so just apply common sense by not annoying human readers. One readable sentence containing the keyword(s) might suffice. Also, emphasizing keywords (EM/I, STRONG/B, eye catching colors ...) makes sense because it helps catching the attention of scanning visitors, but don't over-emphasize because that looks crappy. The same goes for H2/H3/... headings. Structure your copy, but don't write in headlines. When you emphasize a word or phrase in (bold) red, then don't do that consistently but only in the most important sentence(s) of your page, and better only on the first visible screen of a longer page.

Work in your keyword+context laden sentences, but -again!- do it in a natural way. You're writing for humans, not for algos which at this point already know what your page is all about and rank it properly. If your fine tuning gains you a better ranking that's fine, but the goal is catching the attention of searchers reading (in most cases just skimming) your page title and a machine generated snippet on a search result page. Convince the algo to use your inserted sentence(s) in the snippet, not keyword lists from navigation elements or so.

Write a sensible summary of the page's content, not more than 200-250 characters, and put that into the description meta tag. Do not copy the first paragraph or other text from the page. Write the summary from scratch instead, and mention the targeted keyword(s). The first paragraph on the page can exceed the length of the meta description to deliver an overview of the page's message, and it should provide the same information, preferably in the first sentence, but don't make it longish.

Check the TITLE tag in HEAD: when it is truncated on the SERP then shorten it so that the keyword becomes visible, perhaps move the keyword(s) to the beginning, or create a neat page title around the keyword(s). Do title changes very carefully, because the title is an important ranking factor and your changes could result in a ranking drop. Some CMSs change the URL without notice on changes of the title text, and you certainly don't want to touch the URL at this point.

Make sure that the page title appears on the page too. Putting the TITLE tag's content (or a slight variation) in a H1 element in BODY cannot hurt. If you for some weird reasons don't use H-elements, then at least format it prominently (bold, different color but not red, bigger font size ...).


If the page performs nice with a couple money terms and just has a crappy CTR for a particular keyword it ranks for, you can just add a link pointing to a (new) page optimized for that keyword(s), with the keyword(s) in the anchor text, preferably embedded in a readable sentence within the content (long enough to fill two lines under the linked title on the SERP), to improve the snippet. Adding a (prominent) link to a related topic should not impact rankings for other keywords too much, but the keywords submitted by searchers should appear in the snippet a short while after the next crawl. In such cases better don't change the title, at least not now. If the page gained its ranking solely from anchor text of inbound links, putting the search term on the page can give it a nice boost.


Make sure you get an alert when Ms. Googlebot fetches the changed pages, and check out the SERPs and Google's click stats a few days later. After a while you'll get a pretty good idea of how Google creates snippets, and which snippets perform best on the SERPs. Repeat until success.


Related post: Google Quality Scores for Natural Search Optimization by Chris Silver Smith

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Monday, June 25, 2007

Playing with Google Translate (still beta)

I use translation tools quite often, so after reading Google's Udi Manber - Search is a Hard Problem I just had to look at Google Translate again.

Under Text and Web it offers the somewhat rough translations available from the toolbar and links on SERPs. Usually, I use that feature only with languages I don't speak to get an idea of the rough meaning, because the offered translation is, well, rough. Here's an example. Translating "Don't make a fool of yourself" to German gives "einen Dummkopf nicht von selbst bilden". That means "not forming a dullard of its own volition" but Google's reverse translation "a fool automatically do not educate" is even funnier.

Coming with at least rudimentary practices in foreign languages really helps reading Google's automated translations. Quite often the translation is just not understandable without knowledge of the other language's grammar and distinctiveness. For example my french is a bit rusty, so translating Le Monde to english leads to understandable text I can read way faster than the original. Italian to English is another story (my italian skills should be considered "just enough for tourists"), for example the frontpage of la Repubblica is, partly due to the summarizing language, hard to read in Google's english translation. Translated articles on the other hand are rather understandable.

By the way, the quality of translated news, technical writings or academic papers is much better than rough translations of everyday language, so better don't try to get any sense out of translated forum posts and stuff like that. Probably that's caused by the lack of trusted translations of these sources which are necessary to train Google's algos.

Google Translate fails miserably sometimes. Although arabic-english is labelled "BETA", it cannot translate even a single word from the most important source of news in arabic, Al Jazeera - it just delivers a copy of the arabic home page. Ok, that's a joke, all the arabic text is provided on images. Translations of Al Jazeera's articles are terrific, way better than any automated translation from or to european languages I've seen, ever. Comparing Google's translation of the Beijing Review to the english edition makes no sense due to sync issues, but the automated translation looks great, even the headlines make sense (semantically, not in their meanings - but what do I know, I'm not a stalinistic commie killing and jailing dissidents practicing human rights like the freedom of speech).


On the second tab Google translates search results, that's a neat way to research resources in other languages. You can submit a question in english, Google translates it on the fly to the other language, queries the search index with the translated search term and delivers a bilingual search result page, english in the left column and the foreign language on the right side. I don't like that the page titles are truncated, also the snippets are way too short to make sense in most cases. However, it is darn useful. Let's test how Google translates her own pamphlets:

A search in english for [Google Webmaster guidelines] on german pages delivers understandable results. The second search result, "Der Ankauf von Links mit der Absicht, die Rangfolge einer Website zu verbessern, ist ein Verstoß gegen die Richtlinien für Webmaster von Google", gets translated to "The purchase from left with the intention of improving the order of rank of a Website is an offence against the guidelines for Web master of Google". Here it comes straight from the horse's mouth: Google's very own Webmasters must not sell links on the left sidebar of pages on Google.com. I'm not a Webmaster at Google, so in my book that means I can remove the crappy nofollow from tons of links as long as I move them to the left sidebar. (Seriously, the german noun for "link" is "Verbindung" respectively "Verweis", which both have tons of other meanings besides "hyperlink", so everybody in Germany uses "Link" and the plural "Links", but "links" means "left" and Google's translator ignores capitalization as well as anglicisms. The german translation of "Google's guidelines for Webmasters" as "Richtlinien für Webmaster von Google" is quite hapless by the way. It should read "Googles Richtlinien für Webmaster" because "Webmaster von Google" really means "Webmasters of Google" which is (in German) a synonym for "Google's [own] Webmasters".)

An extended search like [Google quality guidelines hidden links] for all sorts of terms from the guidelines like "hidden text", "cloaking", "doorway page" (BTW why is the page type described as "doorway page" in reality a "hallway page", and why doesn't explain Google the characteristics of deceitfully doorway pages, and why doesn't Google explain that most (not machine generated) doorway pages are perfectly legit landing pages?), "sneaky redirects" and many more did not deliver a single page from google.de on the first SERP. No wonder that german Internet marketers are the worst spammers on earth when Google doesn't tell them what particular techniques they should avoid. Hint for Riona: to improve findability consider adding these tags untranslated to all versions of the help system in foreign languages. Hint for Matt: please admit that not each and every doorway page is violating Google's guidelines. A well done and compelling doorway page just highlights a particular topic, hence from a Webmaster's as well as from a search engine's perspective that's perfectly legit "relevance bait" (I can resist to call it spider fodder because it really ain't that in particular).

Ok, back to the topic.


I really fell in love with the recently added third tab Dictionary. This tool beats the pants off Babylon and other word translators when it comes to lookups of single words, but it lacks the reverse functionality provided by these tools, that is the translations of phrases. And it's Web based, so (for example) a middle mouse click on a word or phrase in any application except of my Web browser with Google's toolbar enabled doesn't show the translation. Actually, the quality of one-word lookups is terrific, and when you know how to search you get phrases too. Just play and get familar with it, then when you've at least a rudimentary understanding of the other language you'll often get the desired results.

Well, not always. Submitting "schlagen" ("beat") in German-English mode when I search for a phrase like "beats the pants off something" leads to "outmatch" ("übertreffen, (aus dem Felde) schlagen") as best match. In reverse (English-German) "outmatch" is translated to "übertreffen, (aus dem Felde) schlagen" without alternative or supplemental results, but "beat" has tons of german results, unfortunately without "beats the pants off something".

I admit that's unfair, according to the specs the dictionary thingy is not able to translate phrases (yet). The one-word translations are awesome, I just couldn't resist to max it out with my tries to translate phrases. Hopefully Google renames "Dictionary" to "Words" and adds a tab "Phrases" soon.

Labels: , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Monday, June 18, 2007

Erol ships patch fixing deindexing of online stores by Google

If you run an Erol driven store and you suffer from a loss of Google traffic, or you just want to make sure that your store's content presentation is more compliant to Google's guidelines, then patch your Erol software (*ix hosts / Apache only). For a history of this patch and more information click here.

Tip: Save your /.htaccess file before you publish the store. If it contains statements not related to Erol, then add the code shipped with this patch manually to your local copy of .htaccess and the .htaccess file in the Web host's root directory. If you can't see the (new) .htaccess file in your FTP client, then add "-a" to the external file mask. If your FTP client transfers .htaccess in binary mode, then add ".htaccess" to the list of ASCII files in the settings. If you upload .htaccess in binary mode, it may not exactly do what you expect it to accomplish.

I don't know when/if Erol will ship a patch for IIS. (As a side note, I can't imagine one single reason why hosting an online store under Windows could make sense. OTOH there are many reasons to avoid hosting of anything keen on search engine traffic on a Windows box.)

Labels: , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Thursday, June 14, 2007

Which Sebastian Foss is a spammer?

Obviously pissed by my post Fraud from the desk of Sebastian Foss, Sebastian Foss sent this email to Smart-IT-Consulting.com:
Remove your insults from your blog about my products and sites... as you may know promote-biz.net is not registered to my name or my company.. just look it up in some whois service. This is some spammer who took my software and is now selling it on his spammer websites. Im only selling my programs under their original .com domains and you did not receive any email from me since im only using doube-optin lists.

You may not know it - but insulting persons and spreading lies is under penalty.

Sebastian Foss
Sebastian Foss e-trinity Marketing Inc.
sebastian@etrinity-mail.com
Well, that's my personal blog, and I've a professional opinion about the software Sebastian Foss sells, more on that later. It's public knowledge that spammers do register domains under several entities to obfuscate their activities. I'm not a fed, and I'm not willing to track down each and every multiple respectively virtual personality of a spammer, so I admit that there's at least a slight possibility that the Sebastian Foss spamming my inbox from promote-biz.net is not the Sebastian Foss who wrote and sells the software promoted by the email spammer Sebastian Foss. Since I still receive email spam from the desk of Sebastian Foss at promote-biz.net, I think there's no doubt that this Sebastian Foss is a spammer. Well, Sebastian Foss himself calls him a spammer, and so do I. Confused? So am I. I'll update my other post to reflect that.

Now that we've covered the legal stuff, lets look at the software from the desk of Sebastian Foss.

  • Blog Blaster claims to submit "ads" to 2,000,000 sites. Translation: Blog Blaster automatically submits promotional comments to 2 million blogs. The common description of this kind of "advertising" is comment spam.
    Sebastian Foss tells us that "Blog Blaster will automatically create thousands of links to your website - which will rank your website in a top 10 position!". The common description of this link building technique is link spam.
    The sales pitch signed by Sebastian Foss explains "I used it [Blog Blaster] to promote my other website called ezinebroadcast.com and Blog Blaster produced thousands of links to ezinebroadcast.com - resulting in a #1 position in Google for the term "ezine advertising service". So I understand that Sebastian Foss admits that he is a comment spammer and a link spammer.
    I'd like to see the written permissions of 2,000,000 bloggers allowing Sebastian Foss and his customers to spam their blogs: "Advertising using Blog Blaster is 100% SPAM FREE advertising! You will never be accused of spamming. Your ads are submitted to blogs whose owners have agreed to receive your ads." Laughable, and obviously a lie. Did Sebastian Foss remember that "spreading lies is under penalty"? Take care, Sebastian Foss!

  • Feed Blaster with a very similar sales pitch aims to create the term feed spam. Also, it seems that FeedBlaster™ is a registered trademark of DigitalGrit Inc. And I don't think that Microsoft, Sun and IBM are happy to spot their logos on Sebastian Foss' site e-trinity Internetmarketing GmbH

  • The Money License System aka Google Cash Machine seems to slip through a legal loophole. May be it's not explicit illegal to sell software build to to trick Google Adwords respectively AdSense or ClickBank, but using it will result in account terminations and AFAIK legal actions too.

  • Instant Booster claims to spam search engines, and it does, according to many reports. The common term applied to those techniques is web spam.

All these domains (and there are countless more sites selling similar scams from the desk of Sebastian Foss) are registered by Sebastian Foss respectively his companies e-trinity Internetmarketing GmbH or e-trinity Marketing Inc.

He's in the business of newsgroup spam, search engine spam, comment spam ... probably there's no target left out. Searching for Sebastian Foss scam and similar search terms leads to tons of rip-off reports.

He's even too lazy to rephrase his sales pitches, click a few of the links provided above, then search for quoted phrases you saw on every sales pitch to get the big picture. All that may be legal in Germany, I couldn't care less, but it's not legit. Creating and selling software for the sole purpose of spamming makes the software vendor a spammer. And he's proud of it. He openly admits that he uses his software to spam blogs, search engines, newsgroups and whatever. He may make use of affiliates and virtual entities who send out the email spam, perhaps he got screwed by a chinese copycat selling his software via email spam, but is that relevant when the product itself is spammy?

What do you think, is every instance of Sebastian Foss a spammer? Feel free to vote in the comments.


Update 08/01/2007 Here is the next email from the desk of Sebastian Foss:
Hi,
thanks for the changes on your blog entry - however like i mentioned if you look up the domains which were advertised in the spam mails you will notice that they are not registered to me or my company. You can also see that visiting the sites you will see some guy took my products and is selling them for a lower price on his own websites where he is also copying all of my graphic files. The german police told me that they are receiving spam from your forms and that it goes directly to their trash... however please remove your entries about me from your blog - There is no sense in me selling my own products for a lower price on some cheap, stolen websites - if that would make sense then why do i have my own .com domains for my products ? I just want to make clear that im not sending out any spam mails - please get back to me.

Thanks,
Sebastian

Sebastian Foss
e-trinity Internetmarketing GmbH
sebastian@etrinity-mail.com

It deserves just a short reply:

It makes perfect sense to have an offshore clone in China selling the same outdated and pretty much questionable stuff a little cheaper. This clone can do that because first there's next to no costs like taxes and so on, and second he does it per spamming my inbox on a daily base, hence probably he sells a lot of the 'borrowed' stuff. Whether or not the multiple Sebastian Fosses are the same natural person is not my problem. I claim nothing but leave it up to you dear reader's speculation, common sense, and probability calculation.

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Friday, June 08, 2007

Blogger abuses rel-nofollow due to ignorance

I had planned a full upgrade of this blog to the newest blogger version this weekend. The one and only reason to do the upgrade was the idea that I perhaps could disable the auto-nofollow functionality in the comments. Well, what I found was a way to dofollow the author's link by editing the <dl id='comments-block'> block, but I couldn't figure out how to disable the auto-nofollow in embedded links.

Considering the hassles of converting all the template hacks into the new format, and the risk of most probably losing the ability to edit code my way, I decided to stick with the old template. It just makes no sense for me to dofollow the author's link, when a comment author's links within the content get nofollow'ed automatically. Andy Beard and others will hate me now, so let me explain why I don't move this blog to my own domain using a not that insane software like WordPress.
  • I own respectively author on various WordPress blogs. Google's time to index for posts and updates from this blogspot thingy is 2-3 hours (Web search, not blog search). My Wordpress blogs, even with higher PageRank, suffer from a way longer time to index.
  • I can't afford the time to convert and redirect 150 posts to another blog.
  • I hope that Google/Blogger can implement reasonable change requests (most probably that's just wishful thinking).
That said, WordPress is a way better software than Blogger. I'll have to move this blog if Blogger is not able to fulfill at least my basic needs. I'll explain below why I think that Blogger lacks any understanding of the rel-nofollow semantics. In fact, they throw nofollow crap on everything they get a hand on. It seems to me that they won't stop jeopardizing the integrity of the Blogosphere (at least where they control the linkage) until they get bashed really hard by a Googler who understands what rel-nofollow is all about. I nominate Matt Cutts, who invented and evolved it, and who does not tolerate BS.

So here is my wishlist. I want (regardless of the template type!)
  • A checkbox "apply rel=nofollow to comment author links"
  • A checkbox "apply rel=nofollow to links within comment text"
  • To edit comments, for example to nofollow links myself, or to remove offensive language
  • A checkbox "apply rel=nofollow to links to label/search pages"
  • A checkbox "apply a robots meta tag 'noindex,follow' to label/search pages"
  • A checkbox "apply rel=nofollow to links to archive pages"
  • A checkbox "apply a robots meta tag 'noindex,follow' to archive pages"
  • A checkbox "apply rel=nofollow to backlink listings"
As for the comments functionality, I'd understand when these options get disabled when comment moderation is set to off.

And here are the nofollow-bullshit examples.
  • When comment moderation and captchas are activated, why are comment author links as well as links within the comments nofollow'ed? Does blogger think their bloggers are minor retards? I mean, when I approve a comment, then I do vouch for it. But wait! I can't edit the comment, so a low-life link might slip through. Ok, then let me edit the comments.

  • When I've submitted a comment, the link to the post is nofollowed. Nofollow insane II.This page belongs to the blog, so why the fudge does Blogger nofollow navigational links? And if it makes sense for a weird reason not understandable by a simple webmaster like me, why is the link to the blog's main page as well as the link to the post one line below not nofollow'ed? Linking to the same URL with and without rel-nofollow on the same page deserves a bullshit award.

  • Nofollow insane III. (dashboard)On my dashbord Blogger features a few blogs as "Blogs Of Note", all links nofollow'ed. These are blogs recommended by the Blogger crew. That means they have reviewed them and the links are clearly editorial content. They're proud of it: "we've done a pretty good job of publishing a new one each day". Blogger's very own Blogs Of Note blog does not nofollow the links, and that's correct.

    So why the heck are these recommended blogs nofollow'ed on the dashboard? Nofollow insane III. (blogspot)

  • Blogger inserted robots meta tags "nofollow,noindex" on each and every blog hosted outside the controlled blogspot.com domain earlier this year.

  • Blogger inserted robots meta tags "nofollow,noindex" on Google blogs a few days ago.


If Blogger's recommendation "Check google.com. (Also good for searching.)" is a honest one, why don't they invest a few minutes to educate themselves on rel-nofollow? I mean, it's a Google-block/avoid-indexing/ranking-thingy they use to prevent Google.com users from finding valuable contents hosted on their own domains. And they annoy me. And they insult their users. They shouldn't do that. That's not smart. That's not Google-ish.

Labels: , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Google to kill the power of links

Well, a few types of links will survive and don't do evil in Google's search index ;)    I've updated my first take on Google's updated guidelines stating paid links and reciprocal links are evil. Well, regardless whether one likes or dislikes this policy, it's already factored in - case closed by Google. There are so many ways to generate natural links ...

The official call for paid-link reports is pretty much disliked across the boards:
Google is Now The Morality Police on the Internet
Google's Ideal Webmaster: Snitch, Rake It In And Don't Deliver
Other sites can hurt your ranking
Google's Updated Webmaster Guidelines Addresses Linking Practices
Google clarifies its stance on links

More information, and discussion of paid/exchanged links in my pamphlets:
Matt Cutts and Adam Lasnik define "paid link"
Where is the precise definition of a paid link?
Full disclosure of paid links
Revise your linkage
Link monkey business is not worth a whoop
Is buying and selling links risky? (02/2006)

Labels: , , , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Thursday, June 07, 2007

Danny Sullivan did not strip for Matt Cutts

Nope, this is not recycled news. I'm not referring to Matt asking Danny to strip off his business suit, although the video is really funny. I want to comment on something Matt didn't say recently, but promised to do soon (again).

Danny Sullivan stripped perfectly legit code from Search Engine Land because he was accused to be a spammer, although the CSS code in question is in no way deceitful.

StandardZilla slams poor Tamar just reporting a WebProWorld thread, but does an excellent job in explaining why image replacement is not search engine spam but a sound thing to do. Google's recently updated guidelines need to tell more clearly that optimizing for particular user agents is not considered deceitful cloaking per se. This would prevent Danny from stripping (code) not for Matt or Google but for lurid assclowns producing canards.

Labels: , , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->

Tuesday, June 05, 2007

Google enhances the quality guidelines

Maybe todays update of Google's quality guidelines is the first phase of the Webmaster help system revamp project. I know there's more to come, Google has great plans for the help center. So don't miss out on the opportunity to tell Google's Webmaster Central team what you'd like to have added or changed. Only 14 replies to this call for input is an evidence of incapacity, shame on the Webmasters community.

I haven't had the time to write a full-blown review of the updates, so here are just a few remarks from a Webmaster's perspective. Scroll down to Quality guidelines - specific guidelines to view the updates, that means click the links to the new (sometimes overlapping) detail pages.

As always, the guidelines outline best practices of Web development, refer to common sense, and don't encourage over-interpretations (not that those are avoidable, nor utterly useless). Now providing Webmasters with more explanatory directives, detailed definitions and even examples in the "Don'ts" section is very much appreciated. Look at the over five years old first version of this document before you bitch ;)

Avoid hidden text or hidden links
The new help page on hidden text and links is descriptive and comes with examples, well done. What I miss is a hint with regard to CSS menus and other content which is hidden until the user performs a particular action. Google states "Text (such as excessive keywords) can be hidden in several ways, including [...] Using CSS to hide text". The same goes for links by the way. I wish they would add something in the lines of "... Using CSS to hide text in a way that a user can't visualize it by a common action like moving the mouse over a pointer to a hidden element, or clicking a text link or descriptive widget or icon". The hint at the bottom "If you do find hidden text or links on your site, either remove them or, if they are relevant for your site's visitors, make them easily viewable" comes close to this but lacks an example.

Susan Moskwa from Google clarifies what one can hide with CSS, and what sorts of CSS hidden stuff is considered a violation of the guidelines, in the Google forum on June/11/2007:
If your intent in hiding text is to deceive the search engines, we frown on that; if your intent is purely to improve the visual user experience (e.g. by replacing some text with a fancier image of that same text), you don't need to worry. Of course, as with many techniques, there are shades of gray between "this is clearly deceptive and wrong" and "this is perfectly acceptable". Matt [Cutts] did say that hiding text moves you a step further towards the gray area. But if you're running a perfectly legitimate site, you don't need to worry about it. If, on the other hand, your site already exhibits a bunch of other semi-shady techniques, hidden text starts to look like one more item on that list. [...] As the Guidelines say, focus on intent. If you're using CSS techniques purely to improve your users' experience and/or accessibility, you shouldn't need to worry. One good way to keep it on the up-and-up (if you're replacing text w/ images) is to make sure the text you're hiding is being replaced by an image with the exact same text.


Don't use cloaking or sneaky redirects
This sentence in bold red blinking uppercase letters should be pinned 5 pixels below the heading: "When examining [...] your site to ensure your site adheres to our guidelines, consider the intent" (emphasis mine). There are so many perfectly legit ways to do the content presentation, that it is impossible to assign particular techniques to good versus bad intent, nor vice versa.

I think this page leads to misinterpretations. The major point of confusion is, that Google argues completely from a search engine's perspective and dosn't write for the targeted audience, that is Webmasters and Web developers. Instead of all the talk about users vs. search engines, it should distinguish plain user agents (crawlers, text browsers, JavaScript disabled ...) from enhanced user agents (JS/AJAX enabled, installed and activated plug-ins ...). Don't get me wrong, this page gives the right advice, but the good advice is somewhat obfuscated in phrases like "Rather, you should consider visitors to your site who are unable to view these elements as well".

For example "Serving a page of HTML text to search engines, while showing a page of images or Flash to users [is considered deceptive cloaking]" puts down a gazillion of legit sites which serve the same contents in different formats (and often under different URLs) depending on the ability of the current user agent to render particular stuff like Flash, and a bazillion of perfectly legit AJAX driven sites which provide crawlers and text browsers with a somewhat static structure of HTML pages, too.

"Serving different content to search engines than to users [is considered deceptive cloaking]" puts it better, because in reverse that reads "Feel free to serve identical contents under different URLs and in different formats to users and search engines. Just make sure that you accurately detect the capabilities of the user agent before you decide to alter a requested plain HTML page into a fancy conglomerate of flashing widgets with sound and other good vibrations, respectively vice versa".

Don't send automated queries to Google
This page doesn't provide much more information than the paragraph on the main page, but there's not that much to explain: don't use WebPosition Gold™. Period.

Don't load pages with irrelevant keywords
Tells why keyword stuffing is not a bright idea, nothing to note.

Don't create multiple pages, subdomains, or domains with substantially duplicate content
This detail page is a must read. It starts with a to the point definition "Duplicate content generally refers to substantive blocks of content within or across domains that either completely match other content or are appreciably similar", followed by a ton of good tips and valuable information. And fortunately it expresses that there's no such thing as a general duplicate content penalty.

Don't create pages that install viruses, trojans, or other badware
Describes Google's service in partnership with StopBADware.org, highlighting the quickest procedure to get Google's malware warning removed.

Avoid "doorway" pages created just for search engines, or other "cookie cutter" approaches such as affiliate programs with little or no original content
The info on doorway pages is just a paragraph on the "cloaking and sneaky redirect" page. I miss a few tips on how one can identify unintentional doorway pages created by just bad design, without any deceptive intent. Also, I think a few sentences on thin SERP-like pages would be helpful in this context.

"Little or no original content" targets thin affiliate sites, again doorway pages, auto-generated content, and scraped content. It becomes clear that Google does not love MFA sites.

If your site participates in an affiliate program, make sure that your site adds value. Provide unique and relevant content that gives users a reason to visit your site first
The link points to the "Little or no original content" page mentioned above.


"Buying links in order to improve a site’s ranking is in violation of Google's webmaster guidelines and can negatively impact a site's ranking in search results. [...] Google works hard to ensure that it fully discounts links intended to manipulate search engine results, such link exchanges and purchased links."

Basically that means: if you purchase a link, then make dead sure it's castrated or Google will take away the ability to pass link love from the page (or even site) linking out for green. Or don't get caught respectively denunciated by competitors (I doubt that's a surefire tactic for the average Webmaster).

Note that in the second sentence quoted above Google states officially that link exchanges for the sole purpose of manipulating search engines are a waste of time and resources. That means reciprocal links of particular types nullify each other, and site links might have lost their power too. <speculation>Google may find it funny to increase the toolbar PageRank of pages involved in all sorts of link swap campaigns, but the real PageRank will remain untouched.</speculation>

There's much confusion with regard to "paid link penalties". To the best of my knowledge the link's destination will not be penalized, but the paid link(s) will not (or no longer) increase its reputation, so that in case the link's intention got reported or discovered ex-post its rankings may suffer. Penalizing the link buyer would not make much sense, and Googlers are known as pragmatic folks, hence I doubt there is such a penalty. <speculation>Possibly Google has a flag applied to known link purchasers (sites as well as webmasters), which --if it exists-- might result in more scrupulous judgements of other optimization techniques.</speculation>


What I really like is that the Googlers in charge honestly tried to write for their audience, that is Webmasters and Web developers, not (only) search geeks. Hence the news is that Google really cares. Since the revamp is a funded project, I guess the few paragraphs where the guidelines are still mysterious (for the great unwashed), or even potentially misleading, will get an update soon. I can't wait for the next phase of this project.


Vanessa Fox creates buzz at SMX today, so I'll update this post when (if?) she blogs about the updates later on (update: Vanessa's post). Perhaps Matt Cutts will comment the updated quality guidelines at the SMX conference today, look for Barry's writeup at Search Engine Land, and SEO Roundtable as well as the Bruce Clay blog for coverage of the SMX Penalty Box Summit. Marketing Pilgrim covered this session too. This post at Search Engine Journal provides related info, and more quotes from Matt. Just one SMX tidbit: according to Matt they're going to change the name of the re-inclusion request to something like a reconsideration request.

Labels: , , , , , , ,

Share this post at StumbleUpon
Stumble It!
    Share this post at del.icio.us
Post it to
del.icio.us
 


-->