1.1 意味とは何か
1.2 伝統的な意味論
1.3 Wordnet
1.4 分布仮説と分散表現
演習問題
| [MALE] | [FEMALE] | [ADULT] | [HUMAN] | [MARRIED] | |
|---|---|---|---|---|---|
| bachelor | + | − | + | + | − |
| spinster | − | + | + | + | − |
| woman | − | + | + | + | n.a. |
| wife | − | + | + | + | + |
| girl | − | + | − | + | − |
| boy | + | − | − | + | − |
n.a.: not applicable (設定なし)
実は、これ、認知心理学っぽい話で「カテゴリーは存在するか」という議論でもあったりする。例えば、動物を鳥と魚に分けることはできそうだけど、もっと分かりやすい例を言えば、1から100までの整数を偶数と奇数に分けることはできる。偶数のカテゴリに入るすべての整数は「2で割ると余りが0」という特徴を持っていて、奇数のカテゴリに入る全ての整数は「2で割ると余りが1」という特徴を持っている。では、「ゲーム」というカテゴリに入るものの特徴を言うことは可能でしょうか。それからプロトタイプ理論というものが出てきて...。家族的類似性とか。
シソーラスとは類義語辞書であり、意味が似ている単語が列挙されています。
ここでは、WordNetというシソーラス兼辞書を使用します。
from nltk.corpus import wordnet as wn
wn.synsets("dog")
[Synset('dog.n.01'),
Synset('frump.n.01'),
Synset('dog.n.03'),
Synset('cad.n.01'),
Synset('frank.n.02'),
Synset('pawl.n.01'),
Synset('andiron.n.01'),
Synset('chase.v.01')]
dog = wn.synset("dog.n.01")
dog.definition()
'a member of the genus Canis (probably descended from the common wolf) that has been domesticated by man since prehistoric times; occurs in many breeds'
dog.examples()[0]
'the dog barked all night'
dog.lemma_names()
['dog', 'domestic_dog', 'Canis_familiaris']
wn.synset("dog.n.01").lemma_names("jpn")
['イヌ', 'ドッグ', '洋犬', '犬', '飼い犬', '飼犬']
dog.hypernym_paths()
[[Synset('entity.n.01'),
Synset('physical_entity.n.01'),
Synset('object.n.01'),
Synset('whole.n.02'),
Synset('living_thing.n.01'),
Synset('organism.n.01'),
Synset('animal.n.01'),
Synset('chordate.n.01'),
Synset('vertebrate.n.01'),
Synset('mammal.n.01'),
Synset('placental.n.01'),
Synset('carnivore.n.01'),
Synset('canine.n.02'),
Synset('dog.n.01')],
[Synset('entity.n.01'),
Synset('physical_entity.n.01'),
Synset('object.n.01'),
Synset('whole.n.02'),
Synset('living_thing.n.01'),
Synset('organism.n.01'),
Synset('animal.n.01'),
Synset('domestic_animal.n.01'),
Synset('dog.n.01')]]

computer = wn.synset("computer.n.01")
bicycle = wn.synset("bicycle.n.01")
computer.path_similarity(dog)
0.09090909090909091
computer.path_similarity(bicycle)
0.14285714285714285
それぞれの文におけるballの意味を考えてみましょう。
(奥村 (2010). 『自然言語処理の基礎』より)
分散表現とは単語を数値列(ベクトル)で表現する技術。
分散表現をどのようにして得るのか実際にやってみよう。

Distributional Semantic Models Tutorial at NAACL-HLT 2010より抜粋
import pandas as pd
import numpy as np
label = ["knife","cat","???","boat","cup","pig","banana"]
C = np.array([[51,20,84,0,3,0],
[52,58,4,4,6,26],
[115,83,10,42,33,17],
[59,39,23,4,0,0],
[98,14,6,2,1,0],
[12,17,3,2,9,27],
[11,2,2,0,18,0]])
C
array([[ 51, 20, 84, 0, 3, 0],
[ 52, 58, 4, 4, 6, 26],
[115, 83, 10, 42, 33, 17],
[ 59, 39, 23, 4, 0, 0],
[ 98, 14, 6, 2, 1, 0],
[ 12, 17, 3, 2, 9, 27],
[ 11, 2, 2, 0, 18, 0]])
# Standardization
C2 = []
for c in C:
sum_c = sum(c)
D = []
for c2 in c:
D.append(c2/sum_c)
C2.append(D)
pd.options.display.precision = 3
table = pd.DataFrame({"knife":C2[0],"cat":C2[1],"???":C2[2],"boat":C2[3],"cup":C2[4],"pig":C2[5],"banana":C2[6]},index=["get","see","use","hear","eat","kill"])
Decipher_hieroglyphs = table.T
Decipher_hieroglyphs
| get | see | use | hear | eat | kill | |
|---|---|---|---|---|---|---|
| knife | 0.323 | 0.127 | 0.532 | 0.000 | 0.019 | 0.000 |
| cat | 0.347 | 0.387 | 0.027 | 0.027 | 0.040 | 0.173 |
| ??? | 0.383 | 0.277 | 0.033 | 0.140 | 0.110 | 0.057 |
| boat | 0.472 | 0.312 | 0.184 | 0.032 | 0.000 | 0.000 |
| cup | 0.810 | 0.116 | 0.050 | 0.017 | 0.008 | 0.000 |
| pig | 0.171 | 0.243 | 0.043 | 0.029 | 0.129 | 0.386 |
| banana | 0.333 | 0.061 | 0.061 | 0.000 | 0.545 | 0.000 |
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
%matplotlib inline
pca = PCA(n_components=2)
X = pca.fit_transform(C2)
plt.figure(figsize=(9,6))
for i in range(len(label)):
plt.plot(X[:,0][i],X[:,1][i],"o",color="b")
plt.annotate(label[i],xy=(X[:,0][i],X[:,1][i]),fontsize=22)
それでは、同じような考え方で、以下にある果物と野菜の定義を用いて、単語を区別してみます(単語の類似度を可視化してみます)。
fruit_and_vegetable = {
# Fruits
"strawberry": "A soft red fruit with very small yellow seeds on the surface, that is sweet and grows on a low plant.",
"apple": "A hard, round fruit with smooth green, red, or yellow skin, white flesh, and seeds in the middle, that grows on trees.",
"blueberry": "A small, dark blue, sweet fruit that grows on a bush.",
"banana": "A long curved fruit with a thick yellow skin and soft, sweet white flesh inside.",
"orange": "A round citrus fruit with a thick reddish-yellow skin, sweet juicy flesh, and a lot of juice.",
"grape": "A small, round, green, purple, or red fruit that grows in bunches on vines and is used to make wine or eaten raw.",
"peach": "A round fruit with sweet juicy yellow or white flesh, a fuzzy skin, and a large hard seed inside.",
"watermelon": "A large, round or oval fruit with a thick green skin, sweet watery pink or red flesh, and black seeds.",
# Vegetables
"tomato": "A soft, round red vegetable with a lot of seeds, eaten raw in salads or cooked.",
"cabbage": "A large, round vegetable with thick green, white, or purple leaves that can be eaten raw or cooked.",
"carrot": "A long, thin, orange vegetable that grows underground and has green leaves on top.",
"onion": "A round vegetable with many layers inside and a brown, yellow, or white skin, with a strong sharp smell and taste.",
"potato": "A round white vegetable with a brown, red, or yellow skin, which grows underground from the roots of a plant.",
"cucumber": "A long, thin vegetable with a dark green skin and a light green, watery flesh with small seeds, usually eaten raw in salads.",
"lettuce": "A plant with large, green leaves that are eaten raw in salads.",
"broccoli": "A vegetable made up of small green or purple flower heads that are eaten cooked or raw.",
"eggplant": "A vegetable with a shiny, dark purple skin and soft, white flesh inside, usually eaten cooked."
}
# sklearnのクラスのインポート
from sklearn.feature_extraction.text import CountVectorizer
# 単語ベクトルの生成のためのインスタンス生成
# 引数の"mind_df=1"は「1回以上出現した単語を数える」という意味
# 2回以上出現した単語だけをカウントするには"min_df=2"とする
vectorizer = CountVectorizer(min_df=1)
# 定義のみを取り出す
L = []
for i,j in fruit_and_vegetable.items():
L.append(j)
# 単語ベクトルの生成のプロセス1
vectorizer.fit(L)
# 単語ベクトルの生成のプロセス2
X = vectorizer.transform(L)
# strawberryの単語ベクトル
vectorizer.transform([fruit_and_vegetable["strawberry"]]).toarray()
array([[1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0,
0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 2, 0, 0,
0, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 1, 0, 1, 1,
0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 1]])
# このままだと次元数が大きいので、可視化のために単語ベクトルを2次元に圧縮
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
%matplotlib inline
pca = PCA(n_components=2)
X2 = pca.fit_transform(X)
# 2次元に圧縮した単語ベクトルをプロット
labels = list(fruit_and_vegetable.keys())
for i in range(len(labels)):
plt.plot(X2[:,0][i],X2[:,1][i],"o",color="b")
plt.annotate(labels[i],xy=(X2[:,0][i],X2[:,1][i]),fontsize=12)