{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "intro",
   "metadata": {},
   "source": [
    "# 18　Boostingで予測を育てる\n",
    "\n",
    "FUJIMOTO LAB 深層学習コース。Web教材の図と説明を読んでから実行してください。Pythonの計算はColabの実行環境で行います。\n",
    "\n",
    "**この回の目標**：Boostingの『順番に補う』流れと、テスト用データで評価する理由を説明できる\n",
    "\n",
    "このノートブックは小さな合成データを使用し、元の教材の固定Driveパスや動画を必要としません。"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "predict",
   "metadata": {},
   "source": [
    "## 1　まず予想する\n",
    "\n",
    "コードを実行する前に、表示される値や形を予想してください。"
   ]
  },
  {
   "cell_type": "code",
   "id": "demo-one",
   "metadata": {},
   "source": [
    "from sklearn.tree import DecisionTreeClassifier\n",
    "from sklearn.ensemble import GradientBoostingClassifier\n",
    "from sklearn.datasets import make_moons\n",
    "from sklearn.model_selection import train_test_split\n",
    "X, y = make_moons(n_samples=240, noise=0.24, random_state=7)\n",
    "X_train, X_test, y_train, y_test = train_test_split(\n",
    "    X, y, test_size=0.3, stratify=y, random_state=7)\n",
    "small_tree = DecisionTreeClassifier(max_depth=1, random_state=7).fit(X_train, y_train)\n",
    "boost = GradientBoostingClassifier(n_estimators=20, max_depth=1, random_state=7).fit(X_train, y_train)\n",
    "print('学習用', len(X_train), 'テスト用', len(X_test))"
   ],
   "execution_count": null,
   "outputs": []
  },
  {
   "cell_type": "markdown",
   "id": "explain-one",
   "metadata": {},
   "source": [
    "**確かめ方**：2種類の曲線状データを作り、学習用だけでモデルを作りました。テスト用は学習に渡していません。"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "part-two",
   "metadata": {},
   "source": [
    "## 2　1本の木と、組み合わせた木を比べよう\n",
    "\n",
    "同じテスト用データで正解率を比べます。数値は設定やデータで変わるので、いつでもBoostingが勝つとは考えません。"
   ]
  },
  {
   "cell_type": "code",
   "id": "demo-two",
   "metadata": {},
   "source": [
    "print('1本の木:', round(small_tree.score(X_test, y_test), 3))\n",
    "print('Boosting:', round(boost.score(X_test, y_test), 3))\n",
    "print('木の数:', boost.n_estimators)"
   ],
   "execution_count": null,
   "outputs": []
  },
  {
   "cell_type": "markdown",
   "id": "explain-two",
   "metadata": {},
   "source": [
    "**結果を読む**：条件をそろえて比較します。木を増やした結果が学習用で良くても、テスト用で良くなるとは限りません。"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "experiment",
   "metadata": {},
   "source": [
    "## 3　値を変えて比べる\n",
    "\n",
    "次のセルでは、表示される値や条件を変えて、何が結果を決めるかを確かめます。長くかかる実験は、少数の合成データで行います。"
   ]
  },
  {
   "cell_type": "code",
   "id": "lab",
   "metadata": {},
   "source": [
    "# 複数の木が、予測の間違いを減らせたか調べよう\n",
    "tree_pred = small_tree.predict(X_test)\n",
    "boost_pred = boost.predict(X_test)\n",
    "print('1本の木の誤り:', int((tree_pred != y_test).sum()))\n",
    "print('Boostingの誤り:', int((boost_pred != y_test).sum()))\n",
    "print('注意: この小さな例だけで常に優れるとは言えません。')\n"
   ],
   "execution_count": null,
   "outputs": []
  },
  {
   "cell_type": "markdown",
   "id": "extra-guide-1",
   "metadata": {},
   "source": [
    "## 追加実習 1　木の数を変える\n",
    "\n",
    "学習用とテスト用の両方を比べます。テスト用の値が下がるなら、木を増やすだけでは解決しません。\n",
    "\n",
    "**実行前に予想**：何が変わり、何が変わらないでしょうか。"
   ]
  },
  {
   "cell_type": "code",
   "id": "extra-code-1",
   "metadata": {},
   "source": [
    "for count in [1, 5, 20, 80]:\n",
    "    trial = GradientBoostingClassifier(n_estimators=count, max_depth=1, random_state=7)\n",
    "    trial.fit(X_train, y_train)\n",
    "    print(count, '本 学習', round(trial.score(X_train, y_train), 3),\n",
    "          'テスト', round(trial.score(X_test, y_test), 3))\n"
   ],
   "execution_count": null,
   "outputs": []
  },
  {
   "cell_type": "markdown",
   "id": "extra-reflect-1",
   "metadata": {},
   "source": [
    "**確認**：予想と違った点を一つ書き、値を一つ変えて再実行してください。"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "extra-guide-2",
   "metadata": {},
   "source": [
    "## 追加実習 2　予測と正解を数件だけ見る\n",
    "\n",
    "正解率の裏側には、当たった例と外れた例があります。\n",
    "\n",
    "**実行前に予想**：何が変わり、何が変わらないでしょうか。"
   ]
  },
  {
   "cell_type": "code",
   "id": "extra-code-2",
   "metadata": {},
   "source": [
    "guessed = boost.predict(X_test)\n",
    "for index in range(8):\n",
    "    print(index+1, '正解', y_test[index], '予測', guessed[index])\n",
    "print('テスト用で外した件数:', int((guessed != y_test).sum()))\n"
   ],
   "execution_count": null,
   "outputs": []
  },
  {
   "cell_type": "markdown",
   "id": "extra-reflect-2",
   "metadata": {},
   "source": [
    "**確認**：予想と違った点を一つ書き、値を一つ変えて再実行してください。"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "reflection",
   "metadata": {},
   "source": [
    "## 自分の言葉で答えよう\n",
    "\n",
    "木を増やした後、改善したか確かめるには？\n",
    "\n",
    "- まず予想を書く\n",
    "- コードのどの行が答えを決めるか指す\n",
    "- 条件や値を1つ変えて、予想と実行結果を比べる\n",
    "\n",
    "**ヒント**：初めて見るデータでも役立つかを見るため、テスト用で比較します。"
   ]
  }
 ],
 "metadata": {
  "colab": {
   "name": "18-boosting.ipynb",
   "provenance": []
  },
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
