β˜‘ MCQ PRACTICE

Natural Language Processing Unit 5

Practice objective questions for quick revision and examination preparation. Try answering each question before revealing the answer.

πŸ“š Natural Language Processing
πŸ“– Unit 5
🎯 MCQs

Natural Language Processing - Unit-5

1
A language model assigns
Aa part of speech to each word
Ba parse tree to a sentence
Ca sense to each word
Da probability to a sequence of words
Correct Answer a probability to a sequence of words
2
A bigram model predicts a word based on
Athe previous one word
Bthe previous two words
Cthe whole document
Dthe next word
Correct Answer the previous one word
3
A trigram model predicts a word based on
Athe previous one word
Bno previous words
Cthe previous two words
Dthe next two words
Correct Answer the previous two words
4
The Markov assumption in n-gram models states that
Athe next word depends only on a limited history
Bthe next word depends on the whole text
Cwords are independent of each other and the past
Dthe next word never depends on any word
Correct Answer the next word depends only on a limited history
5
N-gram probabilities are usually estimated from
Athe length of the words
Brelative frequency counts in a training corpus
Cthe number of sentences only
Drandom values
Correct Answer relative frequency counts in a training corpus
6
Perplexity is used to evaluate language models; a model with
Ahigher perplexity is better
Bzero words is better
Clower perplexity is better
Dnegative perplexity is best
Correct Answer lower perplexity is better
7
Perplexity is related to cross-entropy H as
AH squared
BH divided by 2
C2 raised to the power H
DH plus 2
Correct Answer 2 raised to the power H
8
The data sparsity problem in n-gram models means that
Athe corpus is too large
Ball n-grams occur many times
Cthere are no words
Dmany valid n-grams never occur in the training data
Correct Answer many valid n-grams never occur in the training data
9
Smoothing is used to
Adelete rare words
Bincrease the number of zero probabilities
Cstop training
Dgive some probability mass to unseen n-grams
Correct Answer give some probability mass to unseen n-grams
10
Add-one (Laplace) smoothing
Asubtracts one from every count
Badds one to every n-gram count
Cdoubles every count
Dremoves zero counts from the vocabulary
Correct Answer adds one to every n-gram count
11
Good–Turing smoothing re-estimates counts using
Athe number of n-grams that occur a given number of times
Bthe length of the sentences
Cthe alphabetical order of words
Dthe number of paragraphs
Correct Answer the number of n-grams that occur a given number of times
12
Kneser–Ney smoothing is based on
Athe word length
Bcontinuation probability, how many different contexts a word follows
Cthe sentence length only
Dthe alphabetical position
Correct Answer continuation probability, how many different contexts a word follows
13
In backoff, if a higher-order n-gram is not observed, the model
Afalls back to a lower-order n-gram
Bstops and returns an error
Cuses a higher-order n-gram
Ddeletes the sentence
Correct Answer falls back to a lower-order n-gram
14
Interpolation combines
Aonly the highest order
Bonly the unigram
Cno estimates
Destimates from different n-gram orders using weights
Correct Answer estimates from different n-gram orders using weights
15
Class-based language models
Ause a separate model for each letter
Bignore word classes
Cgroup words into classes to reduce data sparsity
Dstore every sentence
Correct Answer group words into classes to reduce data sparsity
16
Variable-length n-gram models
Avary the length of the context depending on the data
Balways use a fixed context of one word
Cuse no context
Duse only the next word
Correct Answer vary the length of the context depending on the data
17
LDA, used in Bayesian topic-based language models, stands for
ALinear Data Analysis
BLanguage Distribution Algorithm
CLatent Dirichlet Allocation
DLatent Decision Approach
Correct Answer Latent Dirichlet Allocation
18
Language model adaptation is the process of
Adeleting the model
Btuning a model to a new domain, topic or speaker
Ctranslating the model
Dcompressing the corpus
Correct Answer tuning a model to a new domain, topic or speaker
19
A cache language model
Adecreases the probability of recent words
Bstores images
Cremoves the vocabulary
Dincreases the probability of recently used words
Correct Answer increases the probability of recently used words
20
Cross-lingual language modeling
Auses only one language
Bignores all other languages
Cworks only for numbers
Duses data or models from other languages to help the target language
Correct Answer uses data or models from other languages to help the target language
21
Words that are not in the training vocabulary are called
Astop words
Bstems
Cout-of-vocabulary (OOV) words
Dlemmas
Correct Answer out-of-vocabulary (OOV) words
22
Bayesian parameter estimation for language models makes use of
Ano probabilities
Bonly the maximum count
Ca random sentence
Da prior distribution over the model parameters
Correct Answer a prior distribution over the model parameters
23
Which of the following is an application of language models?
AImage resizing
BDisk defragmentation
CNetwork routing
DSpeech recognition and machine translation
Correct Answer Speech recognition and machine translation
24
Multilingual language modeling aims to
Ahandle only one language
Bremove all words
Cbuild language models that can handle several languages
Dignore grammar and vocabulary entirely
Correct Answer build language models that can handle several languages
25
An n-gram model with a larger n generally
Aneeds no data
Bnever suffers from sparsity
Cneeds more training data and suffers more from sparsity
Dhas fewer parameters
Correct Answer needs more training data and suffers more from sparsity

Fill in the Blanks

26 A bigram model conditions on the previous __________ word(s).
Correct Answer one
27 A trigram model conditions on the previous __________ words.
Correct Answer two
28 __________ is the standard measure for evaluating language models; lower is better.
Correct Answer Perplexity
29 Perplexity equals 2 raised to the power of the __________ entropy.
Correct Answer cross
30 __________ smoothing adds one to every n-gram count.
Correct Answer Laplace
31 Kneser–__________ smoothing is based on continuation probability.
Correct Answer Ney
32 In __________–Turing smoothing, counts are re-estimated using counts of counts.
Correct Answer Good
33 Falling back to a lower-order n-gram when the higher-order one is unseen is called __________.
Correct Answer backoff
34 Combining estimates of different n-gram orders with weights is called __________.
Correct Answer interpolation
35 LDA stands for Latent __________ Allocation.
Correct Answer Dirichlet
36 __________-based language models group words into classes.
Correct Answer Class
37 Words that are not in the training vocabulary are called out-of-__________ (OOV) words.
Correct Answer vocabulary
38 The Markov assumption says the next word depends only on a limited __________.
Correct Answer history
39 Language model adaptation tunes a model to a new __________.
Correct Answer domain
40 A __________ language model increases the probability of recently used words.
Correct Answer cache
← Back to All MCQs