Using a Pareto approach with an LLM to learn Devanagari, Part 2.

Good morning, Gemini.

Let us refresh in our respective minds what we did with the approach to the Devanagari script last time, so that we might continue where we left off, in the same vein.

Good morning! We are officially back in the polyglot strategy room.

Let’s refresh our mental workspace and look at the foundation we built for our huliganov.tv masterclass.

Where We Left Off: The One-Hour Masterclass

We completely abandoned the traditional, overwhelming method of memorizing the alphabet from to . Instead, we deployed a high-frequency spy-code approach, targeting the letters most commonly used to write international names and loanwords in modern Hindi.

Here is what we unlocked in our first hour:

1. The High-Frequency Core (10 Consonants)

We mastered roughly 30% of the consonant lineup, which statistically unlocked over 80% of the sounds needed for global maps and transliterations:

  • The Liquids/Nasals: (), (), (), ()

  • The Hard Stops (Tongue Curled Back): (Hard ), (Hard )

  • The Global Heavyweights: (), (), (), ()

2. The Vowel Dashboard (The Matras)

We learned how Devanagari behaves as an abugida—where every standalone consonant has a built-in “ah” sound (schwa) that drops naturally at the very end of a word. We learned how to override that built-in vowel using graphic modifiers:

  • (Long ), ि (Short – written to the left), (Long ), / (/), / (Pure /).

  • Our Modern Upgrades: The English short vowel hacks—the Crescent Moon (ॉ) for the “O” in Boston/Coffee and the Double Feather (ै) for the “A” in Bank/Texas.

3. The Grand Finale: The 94% Literacy Proof

We put our toolkit to the test against all 50 US States. The result? A 94% functional literacy rate. Our readers could perfectly read or highly accurately guess almost the entire map (Texas = टेक्सस, Boston = बॉस्टन, California = कैलिफ़ोर्निया) using just that one hour of concepts.

Moving Forward: Writing Mechanics & Disguised Letters

Now, as we discussed, if we want to move our readers from a 67% overall capability toward that 100/100 total fluency point, we need to tackle how the letters are actually built.

 

In Devanagari, knowing the equivalent of stroke order isn’t just about good handwriting—it is a vital tool for optical recognition. Many letters look incredibly similar at a glance, and understanding which stroke is drawn first helps the brain instantly tell them apart.

The Golden Rule of Writing Devanagari

Unlike Western script where you write along a baseline, or Chinese characters drawn inside an imaginary box, Devanagari is built on a hanging line system.

  1. You write the distinctive curves and body of the letter first.

  2. You draw the vertical “anchor wall” (if the letter has one).

  3. Crucial Step: You draw the horizontal top bar (the shirorekha) last, effectively hanging the completed word from the line.

Are you ready to select the first cluster of letters from our high-frequency squad to analyze their stroke paths and compare them to the lookalikes that might confuse our readers in the next pass?

Continue reading “Using a Pareto approach with an LLM to learn Devanagari, Part 2.”

Using a Pareto approach with an LLM to learn Devanagari, Part 1

Using a Pareto approach with an LLM to learn Devanagari!

This article is basically an experiment in how linguists or polyglots or language enthusiasts (whatever we want to call ourselves, what does language matter 😉 ) can use LLMs, in this case Gemini on a PRO sub – you can try your own favorite one with the same prompts and see if it comes up with the same sort of thing) to learn something which is perhaps not altogether straightforward, but nevertheless very useful. In this case the writing system used, after Latin, Chinese characters and Arabic-based writing, by most people on the planet. The premise here is not to learn either Hindi or Sanskrit or any of the other languages that use Devanagari as languages. This is an exercise purely in getting to grips with a writing system. For this reason, we are looking at how this writing system transcribes international words, such as geographical or personal names, but not looking at any Hindi or Nepali per se, not every phrasebook phrases. Apart from anything else, the article should be as agnostic as possible with regard to which language form the set is of most interst to readers, whilst being aware that statistically Hindi will account for the greater part.

The prompts (that is, my bits) are in the H4 format, which for those unfamiliar with doing WordPress means the fourth header level format, which is what I am writing in now. What Gemini answers is in the other formats.

So hopefully you find this interesting on two levels: firstly the experiment as to whether, right now, in July 2026, I can simply dispense with an expensive book and ask Gemini to write me an instruction book on a topic like this as we go along, allowing me to dictate the order of play, the style and other matters, and the second point of interest is the content itself, that is, the actual learning of what in this article is actually only 20 of what it takes to know Devanagari inside out, but it’s been guided to follow a Pareto approach. The test at the close proves that Pareto has not entirely held true, because we have a 67:20 relation rather than a 80:20 as Wilf Pareto might have predicted based on his principle, but anyway I think it’s a good deal to give my readers 67% of a very useful skill while asking them to invest only one hour to gain it.

Obviously it will take a few iterations to get it well memorised, but each run at it is likely to be doable within an hour.

Here goes.

Continue reading “Using a Pareto approach with an LLM to learn Devanagari, Part 1”

Friday AI Day #1 – using AI to explain an intranslatable joke from German and then invent a dynamically equivalent one in English

This article is going to demand a bit of knowledge from the reader as it starts off in German. I was a bit stumped by a German joke I read online because it hinges on the Swiss German and Swabian dialects being almost like foreign languages to a Prussian who speaks a more standard German – called “Hochdeutsch” although in linguistic terms a Prussian would tend more to Low German and the Swabian and Alemannic forms of German are HIgh German dialects becuase originally that meant upland German, with the second German sound shifts and the vowel shifts.

The AI explains the joke in German and then takes my challenge to produce an equivalent joke in English. It first selects Yorshire and Scouser dialects to understand each other and helpfully “translate” the dialect for the benefit of the Londoner but to no effect, as he doesn’t get either. This choice of dialects however makes for a fairly confusing joke as the reader is led to the conclusion that the Londoner must be really thick not to have understood what the other two were saying, although in performance it conceivably might work. I suggested confronting the Londoner in the train witgh Scots and Geordie, and then the AI came up with a more convincing example of how those two dialects might understand one another better than a Londoner might understand either.

All of this I now share with you to illustrate how capable the AI – in this case Copilot – is to deal not only with standard forms of languages but also to communicate and even compose in less standard forms of the language.I’m still experimenting with the best ways to show who is speaking, me or the AI, so for now I am using header formats for when it’s me talking to the AI.

Kennst Du einen Witz über die Schweiz? Ein Schweizer, ein Schwabe und ein Preuße fahren mit der Bahn von Basel nach Stuttgart. Schweigen im Abteil, während die Landschaft vorbei zieht. Um die Stimmung aufzuheitern, fragt der Schweizer den Preussen „sind sie au in Basel gsie?“. Fragende Blicke des Preussen. Springt der Schwabe helfend ein „Err moint gwä“.

😄 Sehr schön erzählt! Das ist ein klassischer Dialekt‑Witz, der mit den feinen Nuancen zwischen Schweizerdeutsch, Schwäbisch und Hochdeutsch spielt. Der Preuße versteht „gsie“ nicht, der Schwabe übersetzt es mit „gwä“ – und beide meinen natürlich „gewesen“.

Das Lustige daran ist, dass sich die drei Dialekte gegenseitig fast wie Fremdsprachen anhören, obwohl sie alle deutsch sind. Gerade im Zug von Basel nach Stuttgart ist das eine sehr realistische Szene: man fährt durch die Sprachgrenzen und merkt, wie unterschiedlich die Regionen klingen.

Wenn du magst, kann ich dir noch ein paar andere Schweizer Witze erzählen, die mit Sprache und Missverständnissen spielen. Zum Beispiel:

  • Ein Tourist fragt in Zürich: „Sprechen Sie hier auch Hochdeutsch?“ – Antwort: „Ja, aber nur mit Gästen.“
  • Oder der Klassiker: „Was ist das Beste an der Schweiz?“ – „Ich weiß es nicht, aber die Flagge ist ein großes Plus.“ 🇨🇭

Möchtest du lieber noch ein paar Dialekt‑Witze (Schwaben, Bayern, Schweizer) oder eher die typischen Schweiz‑Klischee‑Witze hören?

If you had to make an equivalent of that joke using English dialects, how would that go?

Great challenge! Let’s mirror the Swiss–Swabian–Prussian joke with English dialects. The humor comes from one speaker using a regional verb form, another trying to “helpfully” translate it into another dialect, and the outsider still being lost. Here’s how it could work in English:

🚂 English Dialect Version

Three men — a Yorkshireman, a Scouser (Liverpool), and a Londoner — are riding the train from Manchester to London. Silence in the carriage as the countryside rolls by. To break the ice, the Yorkshireman asks the Londoner:

“Tha bin i’ Manchester, then?”
The Londoner looks puzzled.

The Scouser jumps in to help:

“He means, ‘Have yer been?’”

The Londoner still looks baffled, muttering:

“Been where?”

Why it works

    • Yorkshire dialect: “Tha bin” = “Have you been.”

 

    • Scouse dialect: “Yer been” = “You’ve been.”

 

    • Londoner: Used to standard English, still doesn’t get the point.

 

    • The joke plays on the same rhythm as the Swiss one: regional verb forms that sound alien to outsiders, and a “translation” that doesn’t actually clarify.

 

Would you like me to sketch a few alternative versions — say with Scots, Irish, and Cockney — so you can see how the humor shifts depending on which dialects clash?

 

It seems like the English accents are more mutually intelligible than the German ones.

That’s a sharp observation. You’re right: English dialects are usually more mutually intelligible than German dialects, and the reason lies in history and linguistic structure.

🗣 Why English Dialects Stay Closer

  • Standardization: English had early standardization through printing (London English became dominant in the 15th–16th centuries).
  • Mass media: Radio, TV, and now the internet reinforce a common “Standard English,” so even strong accents rarely block comprehension.
  • Shared vocabulary: Most dialects use the same core words, with differences in pronunciation or a few local terms.

🇩🇪 Why German Dialects Diverge More

  • Late standardization: Standard German (Hochdeutsch) only became widespread in the 19th–20th centuries. Before that, local dialects were the everyday language.
  • Regional isolation: Mountain valleys, rural communities, and political fragmentation meant dialects evolved independently.
  • Deep structural differences: Dialects like Swiss German or Bavarian aren’t just accents — they have distinct grammar, verb forms, and vocabulary. For example:
    • Swiss German “gsie” = “gewesen” (been).
    • Swabian “gwä” = also “gewesen.”
    • To a northern German (Preuße), both sound alien.

📊 Comparison

Language Dialect Variation Mutual Intelligibility
English Accents + some vocabulary High (Londoner understands Yorkshireman, even if puzzled)
German Dialects with grammar/vocab differences Lower (Preuße struggles with Swiss/Swabian speech)

🔑 Takeaway

English dialects are mostly pronunciation differences, while German dialects can be different languages in miniature. That’s why the Swiss–Swabian–Prussian joke works so well: the Preuße genuinely doesn’t understand, whereas in English the Londoner would at least catch the gist.

Would you like me to sketch a map of German dialect zones vs. English dialect zones, so you can visually see why German splits more sharply than English?

I think a Scot, a Geordie and a Londoner would be able to get closer to the rhythm.

Exactly — that trio captures the rhythm beautifully, because Scots and Geordies both have deeply distinctive dialects that can feel almost like separate languages to a Londoner. Here’s how an equivalent joke could play out:

🚂 English Dialect Joke (Scot, Geordie, Londoner)

A Scot, a Geordie, and a Londoner are riding the train from Edinburgh to London. Silence in the carriage as the countryside rolls by. To break the ice, the Scot turns to the Londoner:

“Were ye in Embra yestreen?” (Scots: “yestreen” = yesterday evening)
The Londoner looks puzzled.

The Geordie jumps in to help:

“He means, ‘Wor ye theer tha neet?’” (Geordie: “tha neet” or “last neet” = last night)
The Londoner still frowns:

“Sorry, mate… I’ve no idea what you’re on about.”
Why it works

  • Scots dialect: “yestreen” is a real Scots word, but opaque to outsiders.
  • Geordie dialect: “last neet” is clearer, but still foreign‑sounding to a Londoner.
  • Londoner: Even after the “translation,” still confused — just like the Preuße in the Swiss joke.

So yes, a Scot, a Geordie, and a Londoner recreate the same dynamic: two regional dialects bouncing off each other, while the “outsider” remains lost.

Would you like me to sketch a mini‑table of dialect equivalents (Swiss/Swabian/Preuße vs. Scot/Geordie/Londoner), so you can see the parallels side by side?

Yes, that gets closer to it.

I’m glad you feel that way — the Scot/Geordie/Londoner trio really does mirror the Swiss/Swabian/Preuße rhythm, because the dialect gap is wide enough to create genuine confusion, yet still close enough to be funny rather than incomprehensible.

Here’s a neat side‑by‑side comparison to show the parallel:

Trio Dialect Phrase “Helper” Translation Outsider Reaction
Swiss / Swabian / Preuße Swiss: „Sind Sie au in Basel gsie?“ Swabian: „Err moint gwä.“ Preuße: still puzzled
Scot / Geordie / Londoner Scot: “Were ye in Embra yestreen?” Geordie: “He means, wor ye theer last neet?” Londoner: still puzzled

Both jokes hinge on:

  • A regional verb form or time word that outsiders don’t recognize (gsie / yestreen).
  • A helper dialect that’s just as opaque (gwä / last neet).
  • An outsider who remains lost, even after the “translation.

That’s why the rhythm feels so similar — it’s the comedy of mutual intelligibility breaking down inside one language.

Jak mówić po polsku, żeby brzmieć jak z Doliny Krzemowej (ale nie być rozumianym ani tu, ani tam)

How to Speak Polish Like You’re from Silicon Valley (But Be Misunderstood Everywhere)

By David J. James |

Witajcie w świecie pseudoanglicyzmów—językowej krainie, gdzie „zrobić upload” brzmi jak operacja chirurgiczna, „być na callu” to stan egzystencjalny, a „deadlineować” to nowa forma cierpienia. W tym słowniczku pokazujemy, jak polski biznes i młodzieżowy slang tworzą hybrydy, które brzmią światowo, ale są zrozumiałe tylko dla wtajemniczonych.

📘 Słowniczek Pseudoanglicyzmów | Glossary of Pseudo-English Polishisms

🇵🇱 Wyrażenie Znaczenie w polskim kontekście 🇬🇧 Jak to brzmi po angielsku? 🧠 Komentarz
O co kaman? „O co chodzi?”, „Co tu się dzieje?” What’s going on? (but “come on” doesn’t fit) Brzmi jak angielski, ale nim nie jest
Zrobić research Poszukać informacji, przeanalizować Do research (not “make a research”) „Zrobić” wszystko to polska specjalność
Być na callu Uczestniczyć w rozmowie online Be on a call Korpo-slang w pełnej krasie
Zrobić feedback Dać informację zwrotną Give feedback Feedback jako rzeczownik, „zrobić” jako czasownik
Deadlineować Ustalać lub pilnować terminu Doesn’t exist in English Brzmi jak kara boska
Zrobić upload Wgrać plik Upload a file Technicznie brzmi, ale to kalka
Zrobić meeting Zorganizować spotkanie Set up a meeting / Have a meeting Brzmi jak „zrobić obiad”
Zrobić target Osiągnąć cel sprzedażowy Hit the target / Reach the goal Agresywna metafora łowiecka
Zrobić follow-up Wrócić do tematu, przypomnieć Follow up (verb, not noun) Często używane jako rzeczownik, co jest błędem
Zrobić deal Dogadać się, zawrzeć umowę Make a deal / Close a deal Brzmi jak Wall Street, ale to kalka

Continue reading “Jak mówić po polsku, żeby brzmieć jak z Doliny Krzemowej (ale nie być rozumianym ani tu, ani tam)”

A discussion about mathematics and Esperanto productive suffixes and wordbuilding

Sometimes when I talk to AI I feel like I am channeling Philomena Cunk, but occasionallu you get the odd moment when you wonder who is trolling whom…

As usual, my questions are in the headers and the repetitive part where the AI alsways asks if it can do anything else I have edited out.

I would like to be reminded how to calculate a correlation coefficient

Of course! The correlation coefficient, often denoted as r, measures the strength and direction of the linear relationship between two variables. Here’s a step-by-step guide to calculate it:

  1. Collect Data: Gather your paired data points, which are your two sets of variables (e.g., X and Y).
  2. Calculate the Mean: Find the mean (average) of both X and Y.
  3. Subtract the Mean: Subtract the mean of X from each X value to get the deviations from the mean (X – mean of X). Do the same for Y.
  4. Multiply Deviations: Multiply the deviations for each pair (i.e., (X – mean of X) * (Y – mean of Y)).
  5. Square Deviations: Square the deviations for each variable (i.e., (X – mean of X)² and (Y – mean of Y)²).
  6. Sum Up: Sum the multiplied deviations and the squared deviations:
    • Sum of (X – mean of X) * (Y – mean of Y)
    • Sum of (X – mean of X)²
    • Sum of (Y – mean of Y)²
  7. Calculate the Correlation Coefficient: Use the formula:

Where:

  • XiX_i and YiY_i are the individual data points
  • Xˉ\bar{X} and Yˉ\bar{Y} are the means of X and Y
  • The numerator is the sum of the product of the deviations of X and Y
  • The denominator is the square root of the product of the sum of squared deviations for X and Y

The value of r ranges from -1 to 1:

  • 1 indicates a perfect positive linear relationship.
  • -1 indicates a perfect negative linear relationship.
  • 0 indicates no linear relationship.

 

Which part of a fraction is the numerator, top or bottom? Top, right?

That’s right! The numerator is the top part of a fraction. It represents how many parts of the whole are being considered. The bottom part of the fraction is called the denominator, and it indicates the total number of equal parts the whole is divided into.

For example, in the fraction 34\frac{3}{4}:

  • 3 is the numerator (top part)
  • 4 is the denominator (bottom part)

 

And product is if you add something together or muliply it?

The product is the result of multiplying two or more numbers together. For example, the product of 3 and 4 is 3×4=123 \times 4 = 12.

If you add numbers together, the result is called the sum. For example, the sum of 3 and 4 is 3+4=73 + 4 = 7.

 

What if you divide, is that productee?

When you divide, the result is called the quotient. For example, if you divide 12 by 3, the quotient is 4, because 12÷3=412 \div 3 = 4.

So, to summarize:

  • The result of addition is called the sum.
  • The result of multiplication is called the product.
  • The result of division is called the quotient.

 

What’s the word for subtracting, also sum?

The word you’re looking for is difference. When you subtract one number from another, the result is called the difference. For example, if you subtract 5 from 8, the difference is 3, because 8−5=38 – 5 = 3.

So to summarize:

  • The result of addition is called the sum.
  • The result of subtraction is called the difference.
  • The result of multiplication is called the product.
  • The result of division is called the quotient.

 

I wonder whether in an artificial language like esperanto they would use words for sum, difference, product and quotient which would be related to the verbal roots for adding, subtracting etc. Do you know?

Yes, in Esperanto, mathematical terms are often derived from their related verbs, making them easy to understand and remember. Here’s how it works:

  • Sumo (sum) comes from the verb sumigi (to sum/add).
  • Diferenco (difference) comes from the verb subtrahi (to subtract).
  • Produkto (product) comes from the verb multipliki (to multiply).
  • Kvociento (quotient) comes from the verb dividi (to divide).

This system makes Esperanto quite logical and consistent, as the names for mathematical operations are directly related to the actions they represent. It’s one of the many ways Esperanto strives to be an accessible and easy-to-learn language.

 

I’m surprised they didn’t say sumenco, subtrahenco, multiplikenco and dividenco to keep it the same, for example.

That’s an interesting point! Esperanto often aims for simplicity and regularity, but it also takes into account international usage and familiarity. By using terms like “sumo,” “diferenco,” “produkto,” and “kvociento,” it aligns with the international vocabulary that many speakers might already know, making it more intuitive for learners from different linguistic backgrounds.

Using “-enco” endings could have been another logical approach, but Esperanto sometimes opts for more familiar forms to ease learning and understanding. This balance between regularity and international familiarity is part of what makes Esperanto a unique and accessible language.

 

Is an -enco ending productive in Esperanto?

No, the “-enco” ending is not a productive suffix in Esperanto. In Esperanto, word formation is typically very regular, and suffixes like “-o,” “-a,” “-e,” and “-i” are commonly used to create nouns, adjectives, adverbs, and verbs, respectively.

Productive suffixes in Esperanto include:

  • -o: for nouns (e.g., “hundo” – dog)
  • -a: for adjectives (e.g., “bela” – beautiful)
  • -e: for adverbs (e.g., “rapide” – quickly)
  • -i: for infinitive verbs (e.g., “vidi” – to see)

There are also various affixes like:

  • -eco: for abstract nouns (e.g., “beleco” – beauty)
  • -ado: for action nouns (e.g., “parolado” – speech)
  • -ant: for present participles (e.g., “parolanta” – speaking)
  • -int: for past participles (e.g., “parolinta” – having spoken)
  • -ont: for future participles (e.g., “parolonta” – going to speak)

The Esperanto language maintains simplicity and consistency by using these productive suffixes, making it easier to learn and use.

 

I quite like using -enco for the abstract result of applying a verb.

That’s a creative idea! While “enco” isn’t a standard suffix in Esperanto, the beauty of artificial languages is their flexibility and adaptability. You could certainly use “enco” as a way to represent the abstract result of applying a verb in your own usage or within a specific community of speakers.

For example:

  • Sumenco: The result of adding (sum)
  • Subtrahenco: The result of subtracting (difference)
  • Multiplikenco: The result of multiplying (product)
  • Dividenco: The result of dividing (quotient)

Creating your own conventions can be a fun and effective way to explore the language and make it your own. If you find that it catches on with others, it might even become a recognized feature within your Esperanto-speaking community.

Is this allowed?

In the context of Esperanto, language evolution and creativity are highly valued. While “enco” isn’t a standard suffix for abstract results, Esperanto’s flexible nature allows for innovative usage, especially within specific communities or contexts. If you and others find “enco” useful and start using it consistently, it could gain traction.

Esperanto is known for its adaptability, and many speakers enjoy experimenting with new words and forms. Just be aware that if you introduce new terms, it’s a good idea to explain their meaning to ensure clear communication, especially with those less familiar with your innovations.