A blog for fans of Bananagrams, word games, puzzles, and amazing things

Monday, November 7, 2011

How Scrabble dictionaries are made

A web site called Word Buff has an interview with Darryl Francis on the making of the British Scrabble tournament word lists. From his self-description, Francis sounds like a cool guy with interests in wordplay and language. He writes articles for Word Ways, a magazine of recreational linguistics (which I can recommend if you enjoy wordplay). And he cites Martin Gardner's Scientific American columns as a major influence.

Darryl Francis and Allan Simmons are the Dictionary Committee for WESPA, the World English Scrabble Players Association. They've basically been in charge of the British Scrabble word list since its inception.

The British Scrabble tournament word list (previously called SOWPODS and now apparently "CSW" as an abbreviation for Collins Official Scrabble Words) is formed by taking all the words in the most current American Scrabble tournament word list and adding in any valid words from the Chambers Dictionary and the Collins English Dictionary. Francis's interpretation of what constitutes a valid word is given in his response to a question about whether he would ever exclude words that satisfy all the rules:
Let's go back to a group of dictionary entries I mentioned earlier - the internet domain names for countries. There's around 200 of these, running from AC (Ascension Island) to ZW (Zimbabwe). They appear to satisfy the criteria for acceptability of words.

They're not dictionary-listed with an initial capital letter, nor a hyphen nor an apostrophe. They're not marked as abbreviations, they're not marked as foreign. On what basis should they not be allowed as two-letter words?

My answer to this question is that a) these two-letter abbreviations (called "country code top-level domains") are proper nouns, and b) they are in fact abbreviations, whether the dictionary says so or not.

Francis goes on to say:
Yet to allow a sudden influx of two-letter words, most of which are unpronounceable and not recognisable to the man in the street, would be to upset the fine balance that already exists with two-letter words.

Two-letter words are so key to the game that to double their number overnight would almost certainly provoke an outcry from Scrabble players - and probably the media, too.

I could portray this as a question of how to balance strict rule-following with common sense. Ultimately the Dictionary Committee chose not to include all those country codes, so they do use some common sense in their decisions. And they do have to make many difficult judgment calls. But it seems like they only rejected these country codes because there are so many of them and because Scrabble players would be upset by their inclusion.

To me, this demonstrates the subtle biases that have crept into the system to make official Scrabble dictionaries (unsurprisingly) give tournament Scrabble players what they want. As I understand it, what a plurality of them want is a word list that retains the words they have spent so much time memorizing, while occasionally adding handfuls of new words that increase Scrabble scores and make the game easier and more fun for them.

And this is perfectly fine, so long as these Scrabble word lists aren't misappropriated as authoritative sources for other games...


Saturday, October 22, 2011

The 27th Letter of the Alphabet

Rogues, to speak thus irreverently of the alphabet, I shall live to see you glad to serve old Q — to curl the wig of great S — adjust the dot of little i — stand behind the chair of X. Y. Z. — wear the livery of Etcetera — and ride behind the sulky of And-by-itself-and.

From Act I of Charles Lamb's Mr. H

If you were a schoolchild in the 19th century, the alphabet that you learned would have had 27 letters: all 26 letters of our current alphabet, plus the ampersand symbol.
ABCDE
FGHIJ
KLMNO
PQRST
UVWXY
Z &

Due to the awkwardness of ending a recitation of the alphabet with "W X Y Z and", it was traditional to instead say "W X Y and Z, and per se and", where per se, Latin for "by itself", means that &, standing by itself, represents "and". (Words with one-letter spellings, like A or I, were often orally spelt as "A per se" or "A per se A".) It was this process of alphabet recitation, and hurried enunciations of "and per se and" which spawned a variety of names for the & symbol which ultimately converged on "ampersand". & had been part of the alphabet going back to the days of Old English.

Other than standing for the conjunction "and", the ampersand also sometimes appears in the abbreviation &c, for et cetera. This is due to the origins of the & symbol in the first century A.D., when the Romans would write et (Latin for "and") in cursive in a run-together fashion which became a stand-alone written symbol.

So why do we no longer consider & to be part of the alphabet?

The leading theory is that it's because of that alphabet song, the one that goes
A B C D E F G
H I J K LMNOP
Q R S, T U V
W X, Y and Z.
Now I know my A B Cs.
Next time won't you sing with me?
Many incorrectly believe that this is based on a tune by Mozart. While Mozart wrote variations on this theme at the age of 25 [see Köchel listing K. 265], the original melody that inspired him was a French folk song called "Ah! vous dirai-je, Maman" which eventually served as the music for Twinkle, Twinkle, Little Star. In 1835, the alphabet song was copyrighted under the name "The A.B.C., a German air with variations for the flute with an easy accompaniment for the piano forte", so it does seem like Mozart was responsible for popularizing the melody.

It turns out that that song has influence beyond the ousting of the ampersand. Historically, it has been mainly in the U.S. that Z has been pronounced zee; pretty much everywhere else they say zed. But a quick look at the rhyming scheme of the alphabet song, shows that the zee pronunciation works better. And apparently a lot of children who learn English outside of the U.S. are still exposed to this alphabet song through American children's programming, like Sesame Street. Teachers in England reportedly have to correct kindergarteners who enter school singing the alphabet song in an American accent, right down to the zee.

I leave you with this quote from Steven Wright:
Why is the alphabet in that order? Is it because of that song? The guy who wrote that song wrote everything.

Sunday, October 9, 2011

PAX, the Omegathon, and novelty in video games

I love books and documentaries that examine quirky subcultures. The book Word Freak provides a fascinating look inside the world of tournament Scrabble. Murderball was a great film about the players of wheelchair rugby. My favorite quirky documentary though is The King of Kong: A Fistful of Quarters which dramatizes the competition between two players for the high score in the classic arcade game Donkey Kong.

I've just found an amazing article online called PAX Primer which is the perfect introduction to the quirky subculture that is the Penny Arcade Expo. It covers the origins of the Expo (in case you ever wanted to know how a web comic can spawn a convention dedicated to video games and board games), the growth of video games, and their transition from fringe to mainstream culture.

Earlier this year, I posted about PAX because Bananagrams was an event in the PAX East Omegathon. It turns out that the Omegathon organizers decided to feature Bananagrams in the west coast PAX Omegathon as well. The article says that in the convention program, Bananagrams is described as "like Scrabble, only not boring and for old people".

It later goes on to discuss a few computer games which I might opine are "like video games, only not boring and for old people". The new wave of computer games does not suck you into endless repetition.

Portal is a game where you solve puzzles by shooting two holes on different walls, ceilings, or other surfaces in your environment. These "portals" are connected (as if by a wormhole), so whatever goes in one, comes out the other, with the same momentum. I recently started playing this game and can not get enough of it.

Braid is an even stranger game in which the player gets to control the flow of time. The selling points of the game are listed on the game's web site:
  • Every puzzle in Braid is unique. There is no filler.
  • Braid treats your time and attention as precious.
  • Braid does everything it can to give you a mind-expanding experience.
Braid's programmer, Jonathan Blow, self-financed the game as he coded it over three years as a statement about how video games could and should be different.

Braid does not look like any other computer game. The artwork is great. It was done by the artist behind the surreal web comic A Lesson Is Learned But The Damage Is Irreversible. Braid also does not sound like any other computer game. Its atmospheric music helped to win me over.

With the success of these games, even more ambitious games are in the works, on topics such as non-Euclidean geometry (Antichamber) and four-dimensional space (Miegakure).


In a world where video games have become mainstream, it makes sense for a niche to develop for games that emphasize originality. I am glad that quirky subcultures exist to sustain this kind of bold experimentation.

Tuesday, September 20, 2011

Why words are the lengths they are

Some words are long and others are short. What determines how long a particular word should be? If you look at some long words (like "serendipity", "pandemonium", and "hypothesis") and some short words ("my", "in", and "of"), you might come to the conclusion that short words are short because they are used frequently while long words can afford to be long because they come up rarely. This idea was first proposed by a Harvard linguist named George Zipf in 1936.

Researchers at the MIT Department of Brain and Cognitive Sciences took a fresh look at this question and came up with a new theory. They present their results in a paper titled (spoilers!) "Word lengths are optimized for efficient communication".

How much information is conveyed by a word? Consider the sentence that starts "After I got home, I walked the...". If I finish the sentence as "I walked the dog", the extra word "dog" doesn't convey much information because it's probably one of the words your brain was expecting. More surprising would have been "I walked the cat" or "I walked the bulldozer", "I walked the quasar" or "I walked the plank". It is the amount of surprise that researchers are equating with the information contained in a word. Consequently the information content of a word depends on the context that it appears in. [For those who want a more quantitative explanation, the information contribution from a particular context (like, "I walked the...") is -log(p), where p is the probability that the word appears at the end of that phrase and where log() is the natural logarithm function. To get the total information for a word like "dog", you just sum -p log(p) over all the contexts that "dog" appears in.]

Ideally what the researchers would have liked to examine is the relationship between how long it takes to say words and how much information they convey, but it was easier (and, they argue, an adequate approximation) to use the number of letters in a word in place of its utterance duration. But later, they went back and ran the same tests (for a few languages) using number of syllables instead of number of letters, and the results were the same.

To calculate the relationship between word length and frequency, the researchers used the same N-gram data set that Google used in its N-gram viewer. This figure from the paper summarizes their findings:
The plot on the left shows word length versus word use frequency, with frequency decreasing from left to right. (Here the data has been divided into large groups of words ("bins") and the average lengths and frequency have been used.) For the first few points (high-frequency words like "the"), the slope of the line is strong, but then it quickly flattens out, indicating that for low-frequency words, the frequency of the word doesn't change the length very much.

The plot on the right shows average word length versus the information content of the word. Here, the line starts off jagged but then becomes strongly-sloped and very straight. This tells us that how much information a word carries is indeed a good predictor of how long the word will be.

The researchers also cite other work that has shown that, when speaking, people will speak more information-dense syllables more slowly than less information-dense syllables. (If you've ever listened to the synthesized voice of something like a GPS, you'll be familiar with the jerkiness of the pronunciation that sounds like it is speaking some syllables too slowly and others too quickly.)

It would seem that a corollary to this principle is that as a word becomes more common (or more precisely, loses information density), it experiences a linguistic force, pushing it toward a shorter form. This shortening process is called phonetic erosion. Examples of the resulting shortenings (also called clippings) are "refrigerator" becoming "fridge", "going to" becoming "gonna", and "cabriolet" being completely replaced by "cab". Here are a few other terms that have evolved much shorter forms:
  • advertisement → ad
  • caravan → van
  • examination → exam
  • gasoline → gas
  • gymnasium → gym
  • influenza → flu
  • public house → pub
So, essentially, the researchers found that the old idea that word length is based mainly on frequency of word usage (short words are used often while long words are used rarely) does a poor job of explaining why words are the lengths they are. The amount of information in a word (averaged over the various contexts that it is used in) is a far better predictor for how long the word will be. The only exception to this is the 5% to 20% of words that are the least informative (generally short, high-frequency words like "the" and "and").

This result holds, not just for English, but also for the other ten languages that they examined (Czech, Dutch, French, German, Italian, Polish, Portuguese, Romanian, Spanish, and Swedish).

The basic idea that I take away from this work is that there is some maximum rate that our brains can understand incoming speech, and that our speech patterns reformulate what we are saying to evenly distribute information over time. It makes me wonder whether pausing for effect is taking advantage of this fact. Similarly, when I say a word slowly to emphasize it, maybe I am just slowing it down to suggest that it contains a lot of information.

Epilogue: In case you were wondering, the actual ending to the sentence that started "After I got home, I walked the..." was "...tightrope.".


Saturday, August 6, 2011

The Bananagrammer Equation

A warning to regular readers: This post is not about games nor about words. It is about math and bananas.

Recently the "Batman Equation" has been memetically propagating around the Internet.

The equation represents the outline of the Batman logo. It is apparently the work of a user on Reddit.

I liked it enough to try to make my own. There are a few tricks to this process. First break the shape up into curves that you can easily write equations for, of the form f(x,y)=0.
Then, to make the curves stop at the desired end points, add in terms like the ones you see under the square roots. They evaluate to either 1 or -1, depending upon the grid position; when this value is negative, the square root is no longer real, and the plotting program will not plot anything. Finally, multiply all the equations together, and you get one big long equation:

(This is a cleaned-up and slightly approximated version of the equation I used for plotting.)

The final plot looks like this:

If I stay there can be no party. I must be out there in the night, staying vigilant. Wherever a party needs to be saved, I'm there. Wherever there are words that need anagramming, I'm there. But sometimes I'm not because I'm out there in the night staying vigilant, watching, lurking, running, jumping, hurdling, sleeping. No, I can't sleep. You sleep. I'm awake. I don't sleep. I don't blink. Am I a bird? No. I'm a banana. I am Bananagrammer. Or am I? Yes, I am Bananagrammer. [applies chapstick]

It is remarkable how well a single ellipse traces out the outer edge of a banana silhouette. I checked a couple of other bananas, and they also have this property. I finally went to a grocery store and sifted through all their bananas to find the least elliptical one I could:


From the red sample points along the edge, I found that even this banana was almost well-approximated by an ellipse.


It is the first three data points that make this an exception to Bananagrammer's First Law of Bananas:
The outer edge of the longitudinal section of a banana follows an elliptical path, with the banana's stem being roughly on the end of the ellipse's long axis.

The question to ask at this point is, "Why are bananas shaped the way they are?". The simple answer is that when a bunch of bananas start growing on a tree, they are initially pointing more down than up. As they become larger, they curve up toward the sun. A banana's exact shape will therefore depend on where it is with respect to its neighbors.

A full explanation of why bananas are so elliptical will require more investigation. People who want to give me research grants are welcome to do so. Actually, everybody is welcome to do so. To everyone else, tune in next week. Same Banana-time, same Banana-channel!

Friday, July 22, 2011

August 2011 Bananagrams events

A couple of notable Bananagrams-related events are scheduled to take place next month.

1) Bananagrams is sponsoring the August 13th instance of WaterFire, a spectacular event that takes place along the river in Providence, Rhode Island. One hundred fires blaze along two-thirds of a mile of the river, illuminating the art and performances that accompany the festivities. WaterFire happens several times each summer, but on August 13th, there will be special Bananagrams-related events.

waterfire.org/bananas is the official page for the Bananaganza. Also available is a schedule for the evening's events.

Providence is home to Brown University and a strong arts scene. A graduate of Brown, Barnaby Evans, created the WaterFire concept and has been running it since 1994. Evans was a friend of Abe Nathanson (the inventor of Bananagrams) and wrote a tribute to Abe.

It's very cool that Bananagrams is sponsoring this event. The Bananagrams components of the evening have not yet been revealed. I think they're going to be surprises, but this idea has been cooking for over a year, so I expect it will be a great event. If you are in the area, I recommend checking it out.

Also, if you happen to go to Providence, keep an eye out for the Bananagrams headquarters sign while driving around:

(This photo was sent in by a personal acquaintance and Bananagrams fan. In case you were wondering, that is not a sign for the "Bananagrams Archives Gallery". I believe the "Archives Gallery" is a separate business in the same building.)


2) On August 14th, large-scale Bananagrams will be played in Prospect Park in Brooklyn. They are going to use 1-foot square tiles made of Masonite (a type of processed wood, sometimes used for house siding and interior doors). The game will look something like this:

Further details on the event are here.


I suppose very large-scale Bananagrams would be played by moving around human-sized tiles, like "human chess" (those games of chess where people act as the chess pieces) except it would be much faster. Played with a full set of tiles, you'd need 144 people. Watching people run around and try to figure out where to stand to form words while other players are peeling off the bunch would be awesome. Human Bananagrams is really something that has to be played.

Thursday, July 21, 2011

Hebrew Bananagrams


As posted in the (now dearly departed) Bananagrammer forum, the Hebrew version of Bananagrams is now out.
The Israeli distributor has a web site dedicated to this version of the game, http://www.bananagrams.co.il which is in Hebrew. There is also an English translation of the site.

If there were an award for the language that Bananagrams would play most differently in, Hebrew would be a contender. Hebrew doesn't have an alphabet; it has an abjad - an alphabet without vowels. In written Hebrew, vowels may optionally be indicated by a system of diacritic marks placed above or below the consonants. In the image above, the tiles are all Hebrew consonants. My suspicion is that this will make Bananagrams matches in Hebrew markedly faster than in other languages.

UPDATE: This article gives more information on how the Hebrew translation of Bananagrams came about:
For the Israeli version, the Dalfens [the family that owns and runs the Israeli distributor] did not only translate the game’s instructions and special terms − such as “split,” dump,” and “peel” − but also consulted a Hebrew linguist regarding the frequency of each letter. [...] the allocation of letters would vary depending on whether Biblical or Modern Hebrew is used, according to the linguist. The Dalfens opted to allocate letters corresponding to the spoken language.