21 January 2013

anaphoric clogs

in the past week i have seen two different signs, in two restrooms separated by several hundred miles, one typed and one handwritten, that both said:

Do not put paper towels in the toilet!!!
It clogs them up.
wait, what? that should be the other way around: They clog it up. the paper towels are the clog-ers, not the clog-ees. how did such a complete mix-up escape from the pens of two people who are presumably competent speakers of English, if not native speakers?

it requires some charitable grammatical analysis, but i think there could be a few contributing factors. first, it is probably intended as a propositional anaphor, referring to the act of paper-towel-putting. one way to render the intended meaning is It clogs it up, which is a little confusing, since there are two occurrences of it—one propositional and one referential. this is probably the moment during the composition of this sentence when panic set in, and—with a gentle semantic push—the second anaphor became them. despite the unambiguously singular antecedent the toilet in the previous sentence, what the sign-writer means to convey is that paper-towel-putting leads to toilet-clogging in general. that generic reading involves (potentially) multiple toilets, and hence them.

the misfortune lies in the fact that it clogs them up is such a short sentence that the reader can latch on to them before they've even decided that it should be a propositional anaphor. in that case, both get interpreted referentially, with the strange consequence that paper towels are being clogged up by toilets. that would be quite the plumbing problem.

21 May 2012

what's in a language's name?

a few weeks ago we held the 7th edition of the Semantics of Semantics of Under-Represented Languages in the Americas conference at Cornell. we had a great bunch of speakers, working on a wide range of languages. besides the variety, i was very interested in the way that languages were represented in the program. issues of identity and linguistic and cultural preservation are of major concern for most native communities. as a result, there's a trend towards using speakers' preferred name for their language. frequently this is an endonym — the name of the language in the language itself.


i decided to compile a list of all the languages represented in talks and posters and see how far this trend has gone. i was a little surprised how few endonyms there were, although some of them jump off the page. in the list below, the name on the left is how the language appeared in the program; on the right is an alternate name. endonyms are italicized.

YucatecMàaya t'àan
TseltalBatźil Ḱop
NavajoDiné bizaad
BlackfootSiksika
Karitiana------
CheyenneTsėhesenėstsestotse
Tlingit------
KtunaxaKootenai
NsyilxcenColville-Okanagan
Ka'aporUrubú
MapudungunMapuche
Mi'kmaqLnuismk
Yauyos Quechua           ------
InuktitutEskimo
Skarù:rę'Tuscarora
Q'anjob'al------
Guaraníavañe'ẽ
NɬeʔkepmxcínThompson
SaanichSENĆOŦEN
Mbyá------
Nez PerceNiimiipuutímt
K'ichee'Qatzijob'al

aside from Blackfoot and Nez Perce, the SULA linguists didn't use any names that clearly came from Euro-Americans naming other peoples or languages. (i'm fairly sure that in these two cases, the Euro names have been claimed by the associated groups, and are considered standard and non-deprecating.) there are some such names that seem over the line for 21st century use. when constructing a LING 101 phonology problem set once, we had to turn to Wikipedia to find a less awful-sounding name for Swampy Cree; the endonym is Omaškêkowak.

nevertheless, just shy of half of the presenters used the endonym for their language of study. some of them are on beyond tongue twisters for native English speakers (i'm looking at you, Nɬeʔkepmxcín). some seem downright otherworldly. i'm considering using a selection of these in an exercise for my freshman seminar on constructed languages. if i give my students an unannotated list of language names, and ask which are natural languages (not telling them that they all are), i wonder which ones they will peg as constructed? i feel like Q'aanjob'al may trick them along the lines of space opera, and the all caps neo-orthography of SENĆOŦEN and the very non-English consonant clusters of Lnuismk could be a trap too.

keep in mind that this is not limited to Native American languages. much more familiar languages have strange endonyms that could trip up even well-educated English speakers. some examples would be Euskara, Suomi, and Cymraeg. (those are Basque, Finnish, and Welsh, respectively. all European languages, all of which i didn't know the endonyms for until after i became a linguist). the question is whether they would just be unfamiliar, or actually make people believe they were constructed. i'm not teaching my course again until next spring, but when i do i'll report back with results. if anyone else wants to try something similar in the meantime, i'd be interested in hearing how it goes.

11 January 2012

'only' peeving on the comics

i read all of my daily comic strips online now.  one of the serious downsides to this is that gocomics.com has a comments thread(!) on every single strip that they post.  they don't generate the same type of bottom-dwelling stuff as youtube comments, but they are some of the most mirthless places on the internet.  nothing is worse than going all Van Hœt on something that's just supposed to be harmless fun.  so i wasn't surprised, but still baffled when i saw this response to a Frazz comic a few weeks ago:


misplaced a modifier? what? i wasn't even familiar with this language peeve. yesterday i was catching up on my RSS backlog, and found a post about "The elusive 'misplaced only'" on Jan Freeman's blog Throw Grammar From The Train. it details how this peeve works: basically people swear up and down about a relationship between linear order and scope involving only, despite the fact that English doesn't work that way.

so the "only fetishists" would like Caulfield to put only immediately preceding heartburn, because they are blinded by dogma and can't see that it modifies the entire VP. all in spite of the fact that putting it there makes the sentence actually sound worse. in other contexts it would sound worser and worser. compare the following constructed sentences, also using heartburn just for fun.
my uncle went to the ER yesterday because he thought he was having a heart attack…

but it turned out he only had heartburn.
?but it turned out he had only heartburn.
moving only makes the sentence sound worse.  put it in the progressive and it becomes even more terrible: he was having only heartburn??  not modern English. so no, our only conclusion here is that neither Caulfield nor Jef Mallett misplaced a modifier.  he put it exactly where it's supposed to go.

09 January 2012

"un reality" and unreality

today, the following bit of Italian headlinese came across the tubes to my RSS reader:

Napoli, presentato Vargas: "Mi sembra di essere in un reality"

the post is about a new acquisition of the soccer team in Naples, and his reaction to arriving in the city. the quote in the headline appears to be "I seem to be in a reality." this would be a decidedly odd thing to say in English — some sort of metaphysical claim.

but actually, it's just the result of a creative borrowing from English. if we were talking about reality vs. fantasy, there's no doubt that the headline would have used the word realtà or verità. (note that i have absolutely no idea what Vargas actually said; the quote is given in a different form later in the article, although still including the word reality, and it may be translated from his native Spanish.)

so what is un reality? it's a reality TV show. wordreference.com even has a separate Italian to English entry for reality indicating this. Italian has a propensity for this type of clip-and-borrow process, often taking just the attributive piece of a phrase or compound, and they turn up very frequently in headlinese, where space is at a premium. (another famous example is Italian basket for English basketball, which is more common than the native pallacanestro.)

perhaps more interesting than the morphological process here, though, is the semantic shift. Vargas' use of un reality clearly indicates that the content of un reality is anything but reality! replace reality with sogno 'dream' or fantasia 'fantasy' and the sentence means basically the same thing. of course, the blame for making a compositional phrase that can easily shift to mean the opposite probably falls more on English here, but Italian helps to obfuscate the process. i'm sure if we start using a reality to mean a fantasy in English, peevers will tell us that it's just another sign that 2012 is certainly the end of days. i guess we'll just have to wait and see what the realtà turns out to be.

23 December 2011

"grammar" as catch-all

yesterday, Mignon Fogarty (aka Grammar Girl) posted a link to "The 20 Most Controversial Rules in the Grammar World" on Google+. sufficiently baited, i read through it. as i did, i noticed that "The Grammar World" is a very vast place, and may in fact encompass several galaxies.

i went through the list a second time to categorize each of these alleged "grammar" points in terms of what linguistic realm they fall under. (several fail to qualify for any linguistic subfield.)  here were the results.

  1. the Oxford Comma
    orthography / style
  2. the pronunciation of "controversial"
    morpho-phonologycalling it that may be generous.
  3. double negatives
    syntax
    they give a lousy example and fairly tone-deaf comments, but it's definitely a syntactic issue.
  4. "irregardless"
    morphology
  5. ending sentences with prepositions
    syntax
    while writing this post, i noticed that when you create a link in the new blogger interface it asks you "To what URL should this link go?" at least it's reassuring that my blog software is a robot, not a native speaker of English.
  6. "hanged" vs. "hung"
    morphology
  7. "like" as a conjunction
    wait…what?
    hold on, this is a double problem. anyone criticizing someone for using like as a conjunction should first ask themselves whether they know what a conjunction is. this peeve requires it to be anything that takes a clausal complement. the built in Mac OS X dictionary lists such uses of like under a subheading "conjunction", and does the same for when (which also baffled me: apparently Saturday is the day when I get my hair done contains a "relative adverb" while     I loved math when I was in school contains a "conjunction". this is completely backwards terminology, since it's clearly the latter that's modifying something verbal, the VP [loved math].)  however,  the entry for "after" has no erroneous label as conjunction, despite the fact that it too can clearly take clausal complements.
    so if this one was controversial, it's more likely because it's poorly defined, rather than having zealots on two sides of a clearly drawn line.
  8. "good" vs. "well"
    morphology
    again, generously. this is a fight over the meaning and distribution of lexical items.
  9. text/internet speak
    vague orthographical morass
    not particularly grammar-y, and a grab-bag unto itself. this should have failed the criteria for inclusion on the list by not being a rule.
  10. starting sentences with "however"
    syntax / style
  11. starting sentences with "but" or "and"
    syntax / style
    should've been 10b.
  12. gender-neutral pronouns
    morphology
    totally misses the point by not even mentioning singular they. in an article about controversy, we're failing to teach the controversy.
  13. split infinitives
    syntax
  14. passive voice
    syntax
  15. punctuation inside quotation marks
    orthography
  16. possessive apostrophes on words that end in 's'
    orthography
    side note: my students this past semester were unduly concerned with this. apparently this is a failing offense in certain US high schools now.
  17. "e-mail" vs. "email"
    very specific piece of orthography
    all the rest up to this point were at least generalizations that had to be applied to individual circumstances.
  18. universal grammar rules
    what is this i don't even.
    we are informed that Noam Chomsky is an "influential linguist", but otherwise…i mean sure, there's some controversy over whether UG exists, and plenty over what it contains, but this description of it says so little. and if you were thinking that this doesn't seem like a "grammar rule" on par with the rest of the list, just wait for the next one.
  19. the fact that there are different kinds of dashes
    typography
    GAH. orthography, ok. punctuation, at the fringe of orthography, maybe. EM AND EN DASHES ARE GRAMMAR? about as much as camera ISO settings are visual cognition, or audio file formats are hearing. this wins the award for shameless list-padding.
  20. "who" vs. "whom"
    morphosyntax
    infamous. covered previously. (now with dead Google Video link!)
the final tally, giving benefit of the doubt: 13 items that could actually be considered on the spectrum of "grammar" from phonology to pragmatics; 4 orthographic points that tangentially bear on the written encoding of language; 2 ill-defined bits of nonsense; and 1 complaint about the geometry of making writing look pretty.  i give it a C- and suggest it repeat its course on the definition of grammar.

30 April 2010

Ignite Ithaca talk

this is storming the internets thanks to all my friends promoting it. they will probably drive more traffic than me posting this here, but nevertheless, here it is: my Ignite Ithaca talk entitled "Why nobody ever taught you how to write good (and what you can do about it)".

01 February 2010

newsflash: language peeves potentially irritable

coming across the twitter tubes this morning (via @jillianp), this story out of New Zealand: "Research could dismay English language purists". in other news, "Water is wet", "Vegetarians not so keen on meat", etc. etc.


i shouldn't mock the bit of news that prompted this piece: a USD $400K+ grant to do a massive morphological survey of English. this could be incredibly useful. it's just the "context" that the writer put it in. gems like:
It is the first time the morphology of the English language has been looked at in this depth since rules were first laid out in the 19th century.
because before the 19th century, there were no rules! sheer and utter chaos! it's a miracle people could even form words.

but this brings up an interesting question: why did the (mostly inaccurate) grammar texts of the 19th century become sacrosanct to the so-called "purists"? and there's no doubt that they've taken on a mystical value, because they have the ability to trump basic logical reasoning. present a "purist" with two options—150 years of hearsay based on something initially wrong, or the collaborative research of renowned language experts—and they'll chose the former every time. and it can't just be anti-institutional, "down with the man!" sentiment; if that were the case, they should have rejected the prescriptivist poppycock (to borrow Geoff Pullum's term) in the first place.

oh well, we all know that the surest way to go insane is to argue with irrational people. so instead, i'll just pretend that the headline on the story was "Research could be pivotal for English linguists" and go about my day.

19 January 2010

give and take: math and linguistics

this is a response to the excellent post "Why Linguists Should Study Math" over at The Lousy Linguist, which i found via fellow Cornellian @nmashton on twitter. i was going to just write a comment there, but i realized that it would probably become rather long.


first of all, let me say that i am in absolute agreement with the sentiments put forth by Chris in his post. in fact, i'm going to be auditing the brand new, never-before-offered Statistics for Linguists course this semester. but i think that one major point needs to be added.

simply: there is a grave asymmetry in linguists learning math versus mathematicians (or statisticians, or computer scientists, etc.) learning linguistics.

here's the scenario. you're a grad student in linguistics. this means that you went to high school once, and probably were rather good at most of your subjects, or you wouldn't be a student anymore. in high school, they made you learn math. if you were really good at it, you made it through single-variable calculus; if not, probably trig. even if you didn't like it and haven't touched math since, you should have a decent sense of How Math Works, in case you need to pick it up again.

but the converse just isn't true. i've audited the NLP course at Cornell, which is taught by an excellent professor in the CS department who has a very solid grounding in theoretical linguistics. but that almost doesn't matter given the fact that there are zero prerequisites for the course. that's right, no LING101, no nothing. the demographics of the ~80-person lecture break down roughly as 70 CS undergrads, 9 linguistics undergrads, and 1 lonely linguistics grad student.

so what's the big problem? they'll learn as they go, right? learning by doing is the best way, no? wrong. as has been shown time after time on Language Log and elsewhere for this and other fields (law, education, etc.), these would-be NLPers have a complex against linguistics. i think they recognize that they're uninformed on the finer points of linguistic theory, but because "hell, i speak a language!" they don't think they need any more expertise to solve complex linguistic problems. throw more code at it, throw more servers at it, we can brute force our way through. i've watched them re-invent the wheel, and it's a square wheel with an off-center axis. and they're not looking to refine its design, or ask those crazy round-wheeler linguists what they've got cooking in their lab. instead they're trying to make titanium and carbon-fiber square wheels, thinking that will improve things. the mantra is to strive for good enough rather than (i concede, unattainable) perfection.

i think that linguists are more and more cognizant of the need for mathematical training. and for those who just aren't math types, they're willing to go find fellow linguists who are, or even statisticians and computer scientists outside their departments to collaborate with. but nobody comes knocking on the linguistics department door. it's open, guys, and seriously, you could stand to visit. we won't bite.

07 December 2009

seek and ye shall not find

morphological revelations on my morning comb through twitter and facebook statuses:

whoa. on the other hand, this isn't entirely unexpected. morphology tends to be entropic, that is, it favors simplicity and regularity and minimal expression, and moves in that direction over time. this doesn't mean the language apocalypse is upon us any more than the heat death of the universe, as predicted by physical entropy, is. just like physical entropy, language entropy can be locally reduced by other factors, particularly token frequency. that is to say—in the broadest terms—speakers are likelier to hang on to irregular forms of words that are used all the time, and tend to regularize words that aren't as common.

that brings me to my "whoa" moment. i just hadn't realized that 'seek' was possibly on the cusp of regularization. so the question is, how does 'seek'/'sought' stack up to other verbs with past tense forms in -ought? to get a comprehensive list, i turned to a reverse dictionary, which yielded just five non-compound -ought pasts: bought, fought, thought, brought, and our test case, sought. next to test their frequencies i headed to wordcount.org, a nifty visualization of frequency in the British National Corpus. admittedly the BNC might not give the most precise results for predicting the tendencies of young speakers in Michigan, but should be accurate enough. here are their ranks (not token counts; smaller numbers indicate higher frequency):

buy/bought: 785/1129
fight/fought: 1484/3204
think/thought: 102/152*
bring/brought: 631/461
seek/sought: 1875/1895

the data reveals that i perhaps shouldn't be as surprised as i was. 'seek' is the least frequent of the five verbs, although strangely 'fought' is the least frequent past tense form. i starred 'thought' since its frequency is probably affected considerably by use of the noun 'thought'. also of note is the fact that 'bring' is the only item whose past tense is more frequent than the base form; this is due to the fact that 'bring' requires a progressive present tense ("I bring the wine" ≠ "I am bringing the wine" but rather "I (habitually) bring the wine"). despite—or perhaps owing in part to—its frequency, 'bring' is subject to taking on a different irregular pattern, 'bring'/'brang'/'brung' in many children's speech and some adult dialects.

anyhow, to wrap this up, it looks like 'sought' might well be the best candidate of these forms to undergo regularization, even if i hadn't expected it before. the only other form that might do the same is 'fought'-->'fighted', but i think that would be even more surprising...i'm actually wondering why its frequency turned out to be so low in the BNC.

a postscript: although i certainly have 'sought' as the past tense of 'seek' in its basic sense "to look for", 'seeked' is also in my lexicon. it's the past tense of the relatively new lexical item 'seek' "to move rapidly through a video or audio clip". 'sought' is terrible as its past tense:

i seeked ahead 2 minutes to skip the commercials.
*i sought ahead 2 minutes to skip the commercials.

this kind of regularization is a common symptom of generating a new, distinct lexical entry from an existing form, cf. the classic case bad/worse/worst vs. bad/badder/baddest.

[UPDATE] regarding 'wrought', which is very low frequency, and i (rightly) eliminated from consideration as not being a productive past form. i commented the following on the ongoing facebook thread that prompted this all:
'wrought' is a strange case...it's actually the old past participle of 'work' (e.g. "wrought iron" = "worked iron" ≠ "wreaked iron"), and the historical past tense of 'wreak' is regular 'wreaked'. they got conflated because both 'work' and 'wreak' were used in the "____ havoc" idiom. since 'wrought' is almost never used outside the idiom any more, it probably doesn't fit into the regularization question here.

27 August 2009

surprisal for dogs

today's Frazz comic:

i don't know if anybody has actually done research on dogs' abilities to learn frequency-based patterns (although we had cottontop tamarins not so long ago). and unlike the grammarpattern-sensitive monkeys, Mario didn't even wait to confirm the probability-based prediction, he just went for it.

07 February 2009

AND??? and i hate you, congress.

why, why do i do things like try to figure out what has been going on in the Senate regarding the scazillion-dollar stimulus plan? it only a) gets my blood pressure up and b) confirms that our elected representatives are morons, or at least have stared at legislative doubletalk for so long that their judgements about English have been seriously compromised. exhibit 1: Senate Amendment 309, introduced by Thomas Coburn (R-OK)

At the appropriate place, insert the following:

SEC. __. LIMIT ON FUNDS.

None of the amounts appropriated or otherwise made available by this Act may be used for any casino or other gambling establishment, aquarium, zoo, golf course, swimming pool, stadium, community park, museum, theater, art center, and highway beautification project.

take another look. "…museum, theater, art center and highway beautification project"??? that is one hell of a project. in fact, i'm pretty sure you won't be finding any such mega-conglomerate initiative anywhere in the original bill. and they passed this amendment. what a waste of time. idiots.

of course, having someone proofread the damn thing and change and to or would have saved the nonsense that this will create in conference committee, in the courts if the bill is signed into law with this amendment in place, &c. &c.

01 February 2009

Pearls Before Swine takes on English-only

sums it up pretty well, i should say (click for big):

27 August 2008

flavors of English on Google

i was just looking through the site statistics for this here blog. one of the most interesting and useful bits of information that statcounter provides me are the search terms that people use. i would say that 99% of these searches are done on Google — we really have drunk the pagerank kool-aid. a lot of searches are pretty lengthy and specific (e.g. "kobe bryant interview in italian" or "who is the girl in the benny lava video?"). one recent search stuck out to me, though. somebody searched for just the word "whomever", and wound up at my previous post "The Office on whomever". i thought that was pretty remarkable. i clicked through on the link that statcounter provided me and saw that the search was made on google.co.uk, and that descriptively adequate was on the front page of results, at position number 6.

then, for whatever reason, i decided to re-run the search using google.com. my post was nowhere to be found on the first page. the results were entirely different. descriptively adequate finally showed up at #14 on the list of results. what's going on? certainly google hasn't written different versions of pagerank to deal with different localizations of English? as far as cataloguing search results goes, the fact that a bunch of Americans in California wrote the algorithm shouldn't adversely affect Brits and the like.

i couldn't stop there. i ran the search on all of the English Google localizations that i could think of, and got even more different results. i've also noted the number of total results that Google estimates, which also (oddly) vary by localization.

localization#total hits
google.com147,480,000
google.co.uk68,200,000
google.ca78,180,000
google.com.au108,190,000
google.com.nz78,460,000

as i was compiling this table i remembered that Google mucks with your search results if you're signed in (which i of course had to be in order to access blogger, without which i couldn't be writing this post). i signed out, and on google.com the DA link rose to #4. i guess i should just be happy i'm on the front page on all of these searches. but there are still lingering, bizarre questions.

why does Google report different numbers of hits for different localizations?
no clue. (comments are open!)

what is causing the rank fluctuations even when i'm not logged in?
some clue. on all of the non-US localizations there is a feature "search pages from [country name]". perhaps i've got fewer australian sites linking to my blog, so my rank is slightly lower in australia than in the US or great britain.

why the hell is Google biasing my custom algorithm against my own damn blog?!
i mean throw me a bone here, guys.

and the baffler...
why do i get this on google.ca?
i mean, you're kidding, right? i'm sure that the frequency of whatever is much higher than that of whomever, but 8 million hits on a word that's in the dictionary should be enough data for google to not question my intent. and why only canadians, eh? this, of course, isn't the first time that i've seen weird spelling suggestions on Google. so perhaps they really do think they know something about English varieties that i don't?

23 August 2008

Malaysian government fails to ban feature reconstruction

please, don't judge me about the inspiration for this post. the short story is "sometimes you just get bored, and who knows where you could end up on Wikipedia!" tonight it was crappy pop song articles. thence comes this quote from the "Controversy" section of the article for this summer's top hit, I Kissed a Girl.

In Malaysian radio stations, the song has been retitled 'I Kissed...' with the words 'a girl' silenced throughout the chorus in the song.
never mind the odd choice of preposition (as a native English speaker i've never heard a song in a radio station; on works fine). the fact of the matter is that this censorship is about as effective as bleeping the -hole in asshole. if you take the phrase "i kissed a girl" and eliminate "a girl", then in isolation it becomes completely open-ended. it could be "i kissed a man" or "i kissed my mother" or "i kissed a frog". too bad there are more lyrics in the song's refrain!
i kissed a girl / and i liked it / the taste of her cherry chapstick
oops! there's a gendered pronoun hanging out there, eight words later. and it needs an antecedent. and the only preceding nominals are i and it. i can't be the antecedent, because then she would have said my, and it is decidedly neuter. so it can only be...gasp! she didn't! chances are nobody's getting the wool pulled over their eyes either; Wikipedia also says that increasing numbers of Malaysians are identifying English as a first language. they can put the pieces of this not-so-tricky linguistic puzzle back together as quickly as i did. censorship falls flat again.

i think i've gotten more linguistic enjoyment out of the song than by listening to it. there's one other bit of the chorus that intrigued me. it's the other pronoun in those lines, namely it. i'm sure that the intended antecedent is "[the fact that] i kissed a girl", but i can't help but get an ambiguous interpretation where it could be topicalized and actually refer to "the taste..." is this a weird judgement? comments are always open here.

18 August 2008

we want...a count noun!

it's great that Language Log has enabled comments on some of their posts, but it's all the more frustrating when i have something pithy to say and they're turned off. this is a would-be comment in response to Arnold Zwicky's post Countification.

in describing the difference between mass and count nouns in English, he says that shrub is a count noun while shrubbery is a mass noun. while i can certainly use shrubbery as a mass noun, i rarely talk about shrubbery at all, and when i do, i'm almost always quoting Monty Python.



throughout the Knights Who Say 'Ni!' sketch, shrubbery is used consistently as a count noun, taking determiners such as a and another, and having a plural form shrubberies. i never thought of these as ungrammatical in any way, although i suppose it would add to the humor, although the pure absurdity of the scene is plenty. we of course also have Monty Python to thank for the brilliant backformation shrubber (n.) - one who arranges, designs, and sells shrubberies.

aaaaaand back!

a barrage of posts is forthcoming! i'm officially declaring my summer blogging malaise to be over as a new school year is (sadly) just around the corner. as with all my blog revival phases in the past, things will probably slow down in a few weeks, but in the meantime, here we go!

04 June 2008

mind your [t]s and [ʔ]s

this is a few days old, coming from this past weekend's National Spelling Bee.  ah, America.  home of the only language in the world that actually creates a need for spelling bees (as far as i know), and word final glottalization of vowels.




Erin Andrews and numbnuts?  youtube gold.

12 May 2008

don't mention this one

a strange coincidence of events just occurred.  i was walking home from campus, where i had been working on a paper on some bizarre control phenomena, when "Pipe Dream" by the now-sadly-defunct Trendy came up on my ipod.  the first sentence of the song is:

would it be all right / if i asked a deaf guy / to borrow his headphones?
yeah, that's PRO, in a clause introduced by ask that also contains an object, being subject-controlled.  i've got enough to talk about in my paper (analyzing sentences of the type "John said PRO to leave"), so this oddball case is going to get conveniently ignored for the time being.

29 April 2008

WTF, Cupertino???

i have just stumbled upon a most peculiar Cupertino effect trap.  i am writing a paper for my phonology ii class on raddoppiamento in Italian.  it's a continuation of some previous work that i did last semester.  when writing my previous paper, i was running Mac OS X 10.4.  the system-wide spellchecker had no clue what this big, long, Italian word was, and offered up no possible corrections.  to avoid lots of little red underlines, i added the word to the dictionary and went merrily on my way.  since then i suffered a hard drive failure, which had two consequences: 1) the dictionary "forgot" that raddoppiamento is a word, as far as i'm concerned and 2) i've upgraded to Mac OS X 10.5.  no problem, i just have to re-teach the dictionary the word.


and that's when this happened:


whaaaaaaaaat?  where did that "d" come from??  i fired up dictionary.app (which presumably is the same dictionary resource used by the spellchecker) and, no surprise, raddoppiamdento is decidedly not in the dictionary either.  nor, as far as i can tell, is any substring that would fool the spelling suggestions algorithm (which i know for a fact does sometimes produce novel suggestions, especially when two words are accidentally conjoined) into thinking that raddoppiamdento was not only a possibility, but in fact preferable to raddoppiamento.

just for confirmation that i wasn't going crazy, i turned to Google, which seems to know which is the real word and which is the imposter (what the hell is going on, imposter just got flagged, despite the fact that dictionary.app gives both imposter and impostor):

um, yes?

if anyone has any hypothesis whatsoever as to what the great "improvement" Apple made to its spellchecker algorithm such that it produces these shenanigans, please leave a comment.

26 April 2008

i can haz fotos?

seriously, flickr? seriously?



a quick handful of refreshes gave me greetings in languages that i expected: Portuguese, Tagalog, Arabic, French, and English.  although i guess flickr isn't totally hanging on formality—both "hello" and "yo" are given as English greetings.

just as long as they don't publish an entire site localization in lolcat.