
Internet Marketing and AI blog

Now let me be the very first to note, I am in no way an expert in neural networks nor Natural Language Processing. I do however, take the time to read a lot of technical documents, including the documents outlining how BERT was trained. A bit dry to be sure, but brilliant. So take what follows in that context. There are people who know more than I about the mechanics but I like to think I know how it impacts SEO a bit more than they might. I'll list some other articles I find to be great reading on the topic at the end of this post. They're highly recommended. So let's cut to the chase ...the most exciting thing about BERT is that suddenly, overnight, every SEO is now an expert on neural networks and natural language processing.
— Ryan Jones (@RyanJones) October 25, 2019
BERT is a fundamental change in how natural language processing works.
It stands for Bidirectional Encoder Representations from Transformers.
It's the B that is at the heart of the breakthrough that makes it state-of-the-art, but as we will get into and as Rani Horev importantly points out, it is more accurately considered non-directional than bidirectional. His article is linked to below.
Let's consider the fundamental issue with natural language processing. The methods used to this point (and even in many implementations and tasks still in use for that matter) operate from left to right or right to left. That is to say, they only understand one thing based on what has been encountered before it. For our purposes here we will consider the left-to-right structure given that it's how you're reading this article.
If we take for example the sentences:
When Dave looked across the room he got mad. Trevor was eating his ice cream and he didn't like it.When the system hits the word "his" in the second sentence it will have a choice between Dave and Trevor as to who that refers to. When it then hits the word "he" it will hit the same options. Using non-BERT Natural Language Processing (NLP) systems (ELMo for example) it assigns probabilities to these values as to their level of likelihood in being correct. The systems worked pretty well, they were accurate up to 75.1% of the time. Not bad right? Until you imagine a world where you only understood the context of a sentence 75.1% of the time.
This doesn't illustrate it perfectly and covers more than we're discussing right now but what we're considering is that each of these elements is viewed independently rather than being reliant on what came before it. This is not to say that positional context is not considered, but often to the contrary - becomes more emphasized. If the context of a word is made clearer by what comes after it, or by words both before and after it, this is now understood where previously is was not.
The input element "dog" for example, is known to be located one position to the right of "my". This would produce the connection "my dog" or "the dog who belongs to me".
But instead of relying on position, the system takes advantage of the full range of neural matching, including a far stronger capacity to consider an element's various meanings and definitions in isolation, then take that information into the context of its position relative to other elements.
When we think about this from a single statement as shown above it we can imagine many glaring holes the system would encounter in building a model to predict how language is arranged on a new phrase.
That's why BERT was trained with the BooksCorpus and English Wikipedia giving it a total of over 3 billion words, formed into sentences of various reading levels, meanings and contexts.
Yes ... apparently a revolution in natural language processing is, in this layperson's opinion, "smart". Just a titch of an understatement.
Basically, the new system understands the direction of the traveler and is not actually giving a boost to the sites that fit that, but rather filtering out or devaluing sites that do not.
Yes, this may seem like semantics and yes I may sound like a Googler, BUT it's an important difference and is what leads Google to the accurate statement:
You can't optimize for BERT, so ...There's nothing to optimize for with BERT, nor anything for anyone to be rethinking. The fundamentals of us seeking to reward great content remain unchanged.
— Danny Sullivan (@dannysullivan) October 28, 2019
Remember: Optimizing for BERT is optimizing for a filter, not an algorithm.The difference is substantial. In the case where we find ourselves optimizing for a filter, the goal is not to optimize but to clarify. Let's look at the example above, we want to rank for a query related to citizens of the Brazil wondering if they need a visa to visit the US. The question we need to ask ourselves is: how do we clarify our content? I would: