(Overview of) Natural Language Processing Lecture 2: Morphology and finite state techniques

(Overview of) Natural Language Processing Lecture 2: Morphology and finite state techniques

Lecture 2: Morphology and finite state techniques

(Overview of) Natural Language ProcessingLecture 2: Morphology and finite state


Paula Buttery (Materials by Ann Copestake)

Computer LaboratoryUniversity of Cambridge

October 2019

Lecture 2: Morphology and finite state techniques

Outline of today’s lecture

Lecture 2: Morphology and finite state techniquesA brief introduction to morphologyUsing morphology in NLPAspects of morphological processingFinite state techniquesMore applications for finite state techniques

Lecture 2: Morphology and finite state techniques

A brief introduction to morphology

Morphology is the study of word structure

We need some vocabulary to talk about the structure:I morpheme: a minimal information carrying unitI affix: morpheme which only occurs in conjunction with

other morphemes (affixes are bound morphemes)I words made up of stem and zero or more affixes.

e.g. dog+sI compounds have more than one stem.

e.g. book+shop+sI stems are usually free morphemes (meaning they can

exist alone)I Note that slither, slide, slip etc have somewhat similar

meanings, but sl- not a morpheme.

Lecture 2: Morphology and finite state techniques

A brief introduction to morphology

Affixes comes in various forms

I suffix: dog+s, truth+fulI prefix: un+wiseI infix: (maybe) abso-bloody-lutelyI circumfix: not in English

German ge+kauf+t (stem kauf, affix ge_t)

Listed in order of frequency across languages

Lecture 2: Morphology and finite state techniques

A brief introduction to morphology

Inflectional morphemes carry grammatical information

I Inflectional morphemes can tell us about tense, aspect,number, person, gender, case...

I e.g., plural suffix +s, past participle +edI all the inflections of a stem are often referred to as a


Lecture 2: Morphology and finite state techniques

A brief introduction to morphology

Derivational morphemes change the meaning

I e.g., un-, re-, anti-, -ism, -ist ...I broad range of semantic possibilities, may change part of

speech: help→ helperI indefinite combinations:


Lecture 2: Morphology and finite state techniques

A brief introduction to morphology

Languages have different typical word structures

I isolating languages: low number of morphemes per word(e.g. Yoruba)

I synthetic languages: high number of morphemes per wordI agglutinative: the language has a large number of affixes

each carrying one piece of linguistic information (e.g.Turkish)

I inflected: a single affix carries multiple pieces of linguisticinformation (e.g. French)

What type of language is English?

Lecture 2: Morphology and finite state techniques

A brief introduction to morphology

English is an analytic language

English is considered to be analytic:I very little inflectional morphologyI relies on word order insteadI and has lots of helper words (articles and prepositions)I but not an isolating language because has derivational


Lecture 2: Morphology and finite state techniques

A brief introduction to morphology

English is an analytic language

English has a mix of morphological features:I suffixes for inflectional morphologyI but also has inflection through sound changes:

I sing, sang, sungI ring, rang, rungI BUT: ping, pinged, pingedI the pattern is no longer productive but the other inflectional

affixes areI and what about:

I go, went, goneI good, better, best

I uses both prefixes and suffixes for derivational morphologyI but also has zero-derivations: tango, waltz

Lecture 2: Morphology and finite state techniques

A brief introduction to morphology

Internal structure and ambiguity

Morpheme ambiguity: stems and affixes may be individuallyambiguous: e.g. paint (noun or verb), +s (plural or 3persg-verb)Structural ambiguity: e.g., shorts or short -sblackberry blueberry strawberry cranberryunionised could be union -ise -ed or un- ion -ise -edBracketing: un- ion -ise -edI un- ion is not a possible form, so not ((un- ion) -ise) -edI un- is ambiguous:

I with verbs: means ‘reversal’ (e.g., untie)I with adjectives: means ‘not’ (e.g., unwise, unsurprised)

I therefore (un- ((ion -ise) -ed))

Lecture 2: Morphology and finite state techniques

Using morphology in NLP

Using morphological processing in NLP

I compiling a full-form lexiconI stemming for IR (not linguistic stem)I lemmatization (often inflections only): finding stems and

affixes as a precursor to parsingmorphosyntax: interaction between morphology andsyntax

I generationMorphological processing may be bidirectional: i.e.,parsing and generation.party + PLURAL <-> partiessleep + PAST_VERB <-> slept

Lecture 2: Morphology and finite state techniques

Aspects of morphological processing

Spelling rules

I English morphology is essentially concatenativeI irregular morphology — inflectional forms have to be listedI regular phonological and spelling changes associated with

affixation, e.g.I -s is pronounced differently with stem ending in s, x or zI spelling reflects this with the addition of an e (boxes etc)

morphophonologyI in English, description is independent of particular


Lecture 2: Morphology and finite state techniques

Aspects of morphological processing

e-insertione.g. boxˆs to boxes

ε→ e/


ˆ s

I map ‘underlying’ form to surface formI mapping is left of the slash, context to the rightI notation:

position of mappingε empty stringˆ affix boundary — stem ˆ affix

I same rule for plural and 3sg verbI formalisable/implementable as a finite state transducer

Lecture 2: Morphology and finite state techniques

Finite state techniques

Finite state automata for recognitionday/month pairs:

0,1,2,3 digit / 0,1 0,1,2

digit digit

1 2 3 4 5 6

I non-deterministic — after input of ‘2’, in state 2 and state 3.I double circle indicates accept stateI accepts e.g., 11/3 and 3/12I also accepts 37/00 — overgeneration

Lecture 2: Morphology and finite state techniques

Finite state techniques

Reminder: Finite-State Automata

FSA are defined as M = (Q,Σ,∆, s,F) where:I Q = {q0,q1,q2...} is a finite set of states.I Σ is the alphabet: a finite set of transition symbols.I ∆ ⊆ Q× Σ×Q is a function Q×Σ→ Q which we write asδ. Given q ∈ Q and i ∈ Σ then δ(q, i) returns a new stateq′ ∈ Q

I s is a starting stateI F is the set of all end states

Lecture 2: Morphology and finite state techniques

Finite state techniques

Recursive FSAcomma-separated list of day/month pairs:

0,1,2,3 digit / 0,1 0,1,2

digit digit

1 2 3 4 5 6

I list of indefinite lengthI e.g., 11/3, 5/6, 12/04

Lecture 2: Morphology and finite state techniques

Finite state techniques


e.g. boxˆs to boxes

ε→ e/


ˆ s

I map ‘underlying’ form to surface formI mapping is left of the slash, context to the rightI notation:

position of mappingε empty stringˆ affix boundary — stem ˆ affix

Lecture 2: Morphology and finite state techniques

Finite state techniques

Finite State Transducers for MorphologyWe will be attempting to map between a word and its structureand to do this we will need an augmentation to the FSA;something called a Finite state transducer (FST).

q0start q1 q2 q3 q4b:b a:o a:o



I FST are used to map between representations.I You can think of a FST as being FSA which produces two

sequences for any given path through the states;I Or alternatively as an FSA which maps one string into


Lecture 2: Morphology and finite state techniques

Finite state techniques

The operation of a FST



q0start q1 q2 q3 q4b:b a:o a:o



Lecture 2: Morphology and finite state techniques

Finite state techniques

The operation of a FST



q0start q1 q2 q3 q4b:b a:o a:o



Lecture 2: Morphology and finite state techniques

Finite state techniques

The operation of a FST



q0start q1 q2 q3 q4b:b a:o a:o



Lecture 2: Morphology and finite state techniques

Finite state techniques

The operation of a FST



q0start q1 q2 q3 q4b:b a:o a:o



Lecture 2: Morphology and finite state techniques

Finite state techniques

Formal Definition of an FSTTo define a FST formally we need to tweak the definition of anFSA to include two more pieces of information.

q0start q1 q2 q3 q4b:b a:o a:o



OUTPUT ALPHABET ∆ Now rather than a single alphabet weneed two alphabets: the input alphabet; andoutput alphabet.

OUTPUT FUNCTION σ(q, i) The output function is amathematical function that takes two arguments(the current state q and a member of the inputalphabet i) and returns the associated outputcharacters o ∈ ∆.

Lecture 2: Morphology and finite state techniques

Finite state techniques

Formal Definition of an FSTOur sheep to ghost language converter example is thenformally defined as follow:

Q = {q0,q1,q2,q3,q4}Σ = {b,a, !}∆ = {b,o, !}

q0 = q0

F = {q4}

δ(q, i) =

b a !

q0 q1 − −q1 − q2 −q2 − q3 −q3 − q3 q4q4 − − −

δ(q, i) =

b a !

q0 b − −q1 − o −q2 − o −q3 − o !q4 − − −

Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish



q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

An FST for the language Opish




q1 q2







Lecture 2: Morphology and finite state techniques

Finite state techniques

Finite state transducer


e : eother : other

ε : ˆ


s : s



e : eother : other

s : sx : xz : z e : ˆ

s : sx : xz : z

ε→ e/


ˆ s

surface : underlyingc a k e s↔ c a k e ˆ sb o x e s↔ b o x ˆ s

Lecture 2: Morphology and finite state techniques

Finite state techniques

Analysing b o x e s


b : bε : ˆ

2 3

4Input: bOutput: b(Plus: ε . ˆ)

Lecture 2: Morphology and finite state techniques

Finite state techniques

Analysing b o x e s


o : o

2 3

4 Input: b oOutput: b o

Lecture 2: Morphology and finite state techniques

Finite state techniques

Analysing b o x e s

1 2 3


x : x

Input: b o xOutput: b o x

Lecture 2: Morphology and finite state techniques

Finite state techniques

Analysing b o x e s

1 2 3


e : ee : ˆ

Input: b o x eOutput: b o x ˆOutput: b o x e

Lecture 2: Morphology and finite state techniques

Finite state techniques

Analysing b o x e ε s


ε : ˆ

2 3


Input: b o x eOutput: b o x ˆOutput: b o x eInput: b o x e εOutput: b o x e ˆ

Lecture 2: Morphology and finite state techniques

Finite state techniques

Analysing b o x e s

1 2

s : s



s : s Input: b o x e sOutput: b o x ˆ sOutput: b o x e sInput: b o x e ε sOutput: b o x e ˆ s

Lecture 2: Morphology and finite state techniques

Finite state techniques

Analysing b o x e s


e : eother : other

ε : ˆ


s : s



e : eother : other

s : sx : xz : z e : ˆ

s : sx : xz : z

Input: b o x e sAccept output: b o x ˆ sAccept output: b o x e sInput: b o x e ε sAccept output: b o x e ˆ s

Lecture 2: Morphology and finite state techniques

Finite state techniques

Using FSTs

I FSTs assume tokenization (word boundaries) and wordssplit into characters. One character pair per transition!

I Analysis: return character list with affix boundaries, soenabling lexical lookup.

I Generation: input comes from stem and affix lexicons.I One FST per spelling rule: either compose to big FST or

run in parallel.I FSTs do not allow for internal structure:

I can’t model un- ion -ize -d bracketing.I can’t condition on prior transitions, so potential redundancy

Lecture 2: Morphology and finite state techniques

Finite state techniques

Lexical requirements for morphological processing

I affixes, plus the associated information conveyed by theaffixed PAST_VERBed PSP_VERBs PLURAL_NOUN

I irregular forms, with associated information similar to thatfor affixesbegan PAST_VERB beginbegun PSP_VERB begin

I stems with syntactic categories (plus more)

Lecture 2: Morphology and finite state techniques

More applications for finite state techniques

Some other uses of finite state techniques in NLP

I Grammars for simple spoken dialogue systems (directlywritten or compiled)

I Partial grammars for text preprocessing, tokenization,named entity recognition etc.

I Dialogue models for spoken dialogue systems (SDS)e.g. obtaining a date:

1. No information. System prompts for month and day.2. Month only is known. System prompts for day.3. Day only is known. System prompts for month.4. Month and day known.

Lecture 2: Morphology and finite state techniques

More applications for finite state techniques

Lecture 2: Morphology and finite state techniques

More applications for finite state techniques

Concluding comments

I English is an outlier among the world’s languages: verylimited inflectional morphology.

I English inflectional morphology hasn’t been a practicalproblem for NLP systems for decades.

I Limited need for probabilities, small number of possiblemorphological analyses for a word.

I Lots of other applications of finite-state techniques: fast,supported by toolkits (eg. openFST), good initial approachfor very limited systems.