
Today, I was reading an article in the January/February 2010 issue of Speech Technology, titled ‘Text-To-Speech’ by Caroline Henton.
The article was about a test used for evaluating the latest text-to-speech (known as TTS) products.
I’m interested in the technology for several reasons; one being my friend Tori_Z uses it to “read” blogs and other web pages. Another reason is my younger brother’s PhD (in Computer Science) thesis was about using computers to parse human language – or machine language recognition – his work dealt with making a computer understand someone, anyone, that spoke to it with an emphasis on the ums, likes, you knows, ahs, and errs and backtracking that people do when they speak naturally.
As I read the article I was struck by how difficult the test was – it consisted of 10 utterances to test TTS performance in pronunciation accuracy, text normalization, pausing, and other aspects of prosody.
Here are the ten utterances – how well can you read and pronounce these phrases? And Tori? I hope your blog reading program doesn’t choke to badly here!
1) McCain called Obama a liberal, and then he insulted him.
2) Jenny gave Peter instructions to follow.
3) I want doors I can shut!
4) Was the red book read or do we have to read it?
5) Bring me a blue towel, and a red one.
6) Cumin, fenugreek, bouillabaisse, rouille, riesling.
7) 124th Avenue, 120 4th Avenue, 100 24th Avenue.
8) Suisun City, Poughkeepsie, Coeur d’Alene, Streatham, Guildford.
9) Barrasso, Boustany, Faleomavaega, Grijalva, Kratovil, Radanovich, Sebelius.
10) Ralph Vaughn Williams, Ralph Fiennes, Nicolas Sarkozy, Nicholas Nickleby, Maria Callas, Black Maria.
I found numbers 6, 7, 8, 9, and 10 to be near tongue twisters. Apparently, number 9 is a partial list of current members of Congress. Four vendors submitted their products for evaluation and they actually preformed rather well. I’m amazed by how fast the technology is improving.
The article closed with:
“For future progress in speech synthesis, more attention to general purpose TTS (and less to, say, in-vehicle navigation) and the semantic disambiguation of homographs (read/read; Maria/Maria) would be universally beneficial.”
Uh huh, I umm... yeah... like, I agree with that, for sure.
:)





