Analyzing and Disavowing Spammy Back links in GSC by Claude Cowork


It’s been years since I had worked on any SEO related tasks as our dedicated team handles it. I had been busy with other stuff like Crypto, Agriculture, B2B, Travelling and my sincere apologies to my readers who had been expecting frequent content updates. Recently I was interested in the AI space but then after trying so many AI models, I gave up because of the repetition, predictable hollow output, hallucinations until I tried Claude. Initially was using Claude to analyze business data but hooked up with it after witnessing the accuracy, speed and performance.

Last week while I was casually browsing through the GSC (Google Search Console), I discovered several issues and thought why not use Claude to do a basic analysis on the issues as the data was huge and manually looking into every aspect of it would be a nightmare. So I decided to work with a couple of GSC issues initially starting with the backlinks which was very weird as there were so many spammy links from raw IP’s which our team confirmed that they never built.

Task-1 Analyzing and Disavowing Spammy Back links in GSC by Claude Cowork

Approximate Time with CoWork: 30 mins | Without CoWork 6-8 hours.
Time saved ~ 7 hours. Zero Errors.

GSC Backlink Data

backlinks


As you see the backlinks doesn’t make any sense and we had no idea on how we got it. So the next step will be to get rid of these links and disavow these links. 

Downloading Backlinks & Analyze it with Claude cowork.

In order to disavow the spammy links we need to analyze all the domains which is 750+ and manually going through every domain will take hours and hours. Here comes Claude cowork. I downloaded the unique domain list and created a project named SIM-GSC in Claude cowork. I then fed the csv file containing 750+ domains to claude cowork with the following prompt to analyze the domains

Cowork Prompt to analyze linking domains from GSC

“I just downloaded and attached the linking domains from GSC, but see lot of ip address, can you identify the worthy domains here and provide me a csv report on what is needed and what isn't so that i can disavow the domains”

Claude cowork thought for a while and after running 5 tasks provided this output.

529 → Disavow. Three groups make up nearly all of it:

~490 bare IP addresses (e.g. 13.251.156.60) — these are never real referring domains; they're bot/crawler artifacts GSC picks up, and Google explicitly says to disavow at domain level when this happens en masse.

 4 classic link-farm/PBN domains with wildly disproportionate link counts vs. target pages

~35 more spam/PBN domains, either on throwaway TLDs (.forum, .garden, .wtf, .live, .casa, .tattoo, .xyz, .link) or with word-salad names

 82 → Keep. Real platforms your links legitimately live on

12 → Review manually. Plausible-sounding domains I couldn't confidently classify either way from the name alone

So now I have to analyze only 12 domains which made my life easier with only around 12k tokens and a few minutes.  Well now that I have manually analyzed the 12 domains and filtered a couple of them, I need to create a disavow file to be updated on the GSC.

Claude Created the Disavow File and also provided instructions on how to upload it with the following guidance

Given the volume of bot/IP noise and obvious PBN spam here, that threshold looks met. If you upload a new file later, it fully replaces the old one — it doesn't merge

Also, disavowing doesn't remove the links from GSC's report — they'll keep showing up as referring domains.

Here is the screenshot of Disavow file that Claude cowork created.

disavow file


Task-2 Analyzing Crawled but not Indexed News URL’s in GSC by Claude Cowork

Approximate Time with CoWork: 45 mins | Without CoWork 7-9 hours.
Time saved ~ 8 hours.

After disavowing the bad backlinks, the task was to identify the URL’s which were crawled by Google bots , yet not indexed by Google. This is a serious issue as after hours and hours of research and creating these articles, why is it not indexed despite being crawled. If it is not indexed it is not to be shown on searches. This needed a real in-depth analysis on why these pages are not indexed and what to do next. I took only the news section of SIM for this task to start with.

The crawled but not indexed pages where filtered for the string “news” and as you see below there are 315 pages which are crawled but not indexed. So I exported all the URL’s to do an in-depth analysis of the URL/Content. The CSV file of URL’s was fed to claude cowork as input.

crawled-not-indexed


I used the following prompt to clean up the list to analyze the links.

“I have exported the crawled but not indexed URL’s and I want you to remove the URL’s of widgets, RSS feeds, News Aggregator pages, Pagination and have only the urls with article content”

Cowork after working for a few minutes and consuming a couple of thousand tokens was able to provide with 147 genuine content URLs from 315 URL’s and removing 168. Now that we have a clean list of content URL’s I tested Cowork with the following prompt to analyze the first 30 URL’s to see how it is performing.

“Okay now can you analyze the first 30 links and suggest what went wrong and why it is not in index after crawling”

Cowork gave a decent summary of what may be the reasons that these pages are not indexed.
-Thin, templated content competing with itself.
-Stale, superseded topics.
-A few pages that are thin by design, not by neglect.
-One structural gap worth fixing regardless of content

It also provided a list of URL’s falling in these categories and also asked if I wanted to do the remaining 117 pages. When I gave the go ahead signal it analyzed the remaining URL’s and provided the following summary.

Soft-404s (26 pages) — 301 individually, not to the homepage. For each one, check if there's a live article covering the same or a closely related topic (a lot of these are old SEO-news slugs that were likely renamed or rewritten).

AMP duplicates (13 pages) — 301 to the canonical non-AMP article.

Case-duplicate URLs (10 pages, 5 pairs) — 301 the non-canonical case variant to whichever version you want indexed.

Parameterized/UTM duplicates (9 pages) — leave live, add a self-referencing canonical tag pointing to the clean URL.

Thin utility pages (7 pages) — noindex, follow, not 301

The redirect-loop bugs (2 pages) — fix the loop itself,

The 70 pages with no technical defect — don't touch their URLs at all. Where several old posts cover the same evergreen topic (multiple Panda/Penguin recaps, multiple "why outsource digital marketing" posts)

 

I asked to provide the results in a language where our developer can understand and make the changes. Claude went o to provide a clean .md file with instructions for the developer on what to do.

claude-gsc-vscode

 


Category :

Technology

Tags :

claude, gsc, cowork

About Guna Nadar

Guna Nadar

Know Guna Nadar

I mostly write technical aspects and not much into creative writing. For the past decade I worked along with top notch SEO & Internet Marketing professionals which naturally lured me into the world of Search Engines. When I am not writing I read from comics to philosophy.Antiques, Fishing, hunting are my passions. Currently I am working on Google Penalty protection and .... more info about the author