{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "6b4fce18-128e-441d-9040-7d738b1ac1c8",
   "metadata": {},
   "source": [
    "## PDSP 2026, Lecture 10, 10 September 2026"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b02df753-a171-4b21-9afa-8c3dc2f74d78",
   "metadata": {},
   "source": [
    "### Arrays, lists and dictionaries"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "3f585472",
   "metadata": {
    "id": "k3jspHRc8oLq"
   },
   "source": [
    "### Arrays\n",
    "- Contiguous block of memory\n",
    "- Typically size is declared in advance, all values are uniform\n",
    "- `a[0]` points to first memory location in the allocated block\n",
    "- Locate `a[i]` in memory using index arithmetic\n",
    "  - Skip `i` blocks of memory, each block's size determined by value stored in array\n",
    "- **Random access** -- accessing the value at `a[i]` does not depend on `i`\n",
    "- Useful for procedures like sorting, where we need to swap out of order values `a[i]` and `a[j]`\n",
    "  - `a[i], a[j] = a[j], a[i]`\n",
    "  - Cost of such a swap is constant, independent of where the elements to be swapped are in the array\n",
    "- Inserting or deleting a value is expensive\n",
    "- Need to shift elements right or left, respectively, depending on the location of the modification"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "fad3a11b",
   "metadata": {
    "id": "afE0tg-G9h6T"
   },
   "source": [
    "### Lists\n",
    "- Each location is a *cell*, consisiting of a value and a link to the next cell\n",
    "  - Think of a list as a train, made up of a linked sequence of cells\n",
    "- The name of the list `l` gives us access to `l[0]`, the first cell\n",
    "- To reach cell `l[i]`, we must traverse the links from `l[0]` to `l[1]` to `l[2]` $\\ldots$ to `l[i-1]`] to `l[i]`\n",
    "  - Takes time proportional to `i`\n",
    "- Cost of swapping `l[i]` and `l[j]` varies, depending on values `i` and `j`\n",
    "- On the other hand, if we are already at `l[i]` modifying the list is easy\n",
    "  - *Insert* - create a new cell and reroute the links\n",
    "  - *Delete* - bypass the deleted cell by rerouting the links\n",
    "- Each insert/delete requires a fixed amount of local \"plumbing\", independent of where in the list it is performed"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1ecdb204",
   "metadata": {
    "id": "8yTypELQ-dcz"
   },
   "source": [
    "### Dictionaries\n",
    "- Values are stored in a fixed block of size $m$\n",
    "- Keys are mapped to $\\{0,1,\\ldots,m-1\\}$\n",
    "- Hash function $h: K \\to S$ maps a *large* set of keys $K$ to a *small* range $S$\n",
    "- Simple hash function: interpret $k \\in K$ as a bit sequence representing a number $n_k$ in binary, and compute $n_k \\bmod m$, where $|S| = m$\n",
    "- Mismatch in sizes means that there will be *collisions* -- $k_1 \\neq k_2$, but $h(k_1) = h(k_2)$\n",
    "- A good hash function maps keys \"randomly\" to minimize collisions\n",
    "- Hash can be used as a *signature* of authenticity\n",
    "  - Modifying $k$ slightly will drastically alter $h(k)$\n",
    "  - No easy way to reverse engineer a $k'$ to map to a given $h(k)$\n",
    "  - Use to check that large files have not been tampered with in transit, either due to network errors or malicious intervention\n",
    "- Dictionary uses a hash function to map key values to storage locations\n",
    "- Lookup requires computing $h(k)$ which takes roughly the same time for any $k$\n",
    "  - Compare with computing the offset `a[i]` for any index `i` in an array\n",
    "- Collisions are inevitable, different mechanisms to manage this, which we will not discuss now\n",
    "- Effectively, a dictionary combines flexibility with random access"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "d5718ce2",
   "metadata": {
    "id": "pmdHUp9catwl"
   },
   "source": [
    "### Lists in Python\n",
    "- Flexible size, allow inserting/deleting elements in between\n",
    "- However, implementation is an array, rather than a list\n",
    "- Initially allocate a block of storage to the list\n",
    "- When storage runs out, double the allocation\n",
    "- `l.append(x)` is efficient, moves the right end of the list one position forward within the array\n",
    "- `l.insert(0,x)` inserts a value at the start, expensive because it requires shifting all the elements by 1\n",
    "- We will run experiments to validate these claims"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "d9c10e9b",
   "metadata": {},
   "source": [
    "### Measuring execution time\n",
    "- Call `time.perf_counter()`\n",
    "- Actual return value is meaningless, but difference between two calls measures time in seconds"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 1,
   "id": "3ba6e7ea",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "The history saving thread hit an unexpected error (OperationalError('attempt to write a readonly database')).History will not be written to the database.\n"
     ]
    }
   ],
   "source": [
    "import time"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 2,
   "id": "2f4e334c-a8a8-48f4-873a-6df829712c5d",
   "metadata": {},
   "outputs": [
    {
     "data": {
      "text/plain": [
       "121540.768069176"
      ]
     },
     "execution_count": 2,
     "metadata": {},
     "output_type": "execute_result"
    }
   ],
   "source": [
    "time.perf_counter()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "9404f16f",
   "metadata": {},
   "source": [
    "- $10^7$ appends to an empty Python list"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 3,
   "id": "274f1c27",
   "metadata": {
    "colab": {
     "base_uri": "https://localhost:8080/"
    },
    "id": "sJEWTNND7lNj",
    "outputId": "22c49e9c-4a48-4cda-b06f-cb65a29a73f3"
   },
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "0.9761906810017535\n"
     ]
    }
   ],
   "source": [
    "start = time.perf_counter()\n",
    "l = []\n",
    "for i in range(10000000):\n",
    "    l.append(i)\n",
    "elapsed = time.perf_counter() - start\n",
    "print(elapsed)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "dbc629f4",
   "metadata": {},
   "source": [
    "- Doubling the work approximately doubles the time, linear"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 4,
   "id": "ab3672e6",
   "metadata": {
    "colab": {
     "base_uri": "https://localhost:8080/"
    },
    "id": "sJEWTNND7lNj",
    "outputId": "22c49e9c-4a48-4cda-b06f-cb65a29a73f3"
   },
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "2.1174399349984014\n"
     ]
    }
   ],
   "source": [
    "start = time.perf_counter()\n",
    "l = []\n",
    "for i in range(20000000):\n",
    "    l.append(i)\n",
    "elapsed = time.perf_counter() - start\n",
    "print(elapsed)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 5,
   "id": "8c4133bf-d851-4435-8ae3-4b810054cb52",
   "metadata": {
    "colab": {
     "base_uri": "https://localhost:8080/"
    },
    "id": "sJEWTNND7lNj",
    "outputId": "22c49e9c-4a48-4cda-b06f-cb65a29a73f3"
   },
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "4.2518419139960315\n"
     ]
    }
   ],
   "source": [
    "start = time.perf_counter()\n",
    "l = []\n",
    "for i in range(40000000):\n",
    "    l.append(i)\n",
    "elapsed = time.perf_counter() - start\n",
    "print(elapsed)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "c01b9010",
   "metadata": {},
   "source": [
    "- $10^5$ inserts at the beginning of a Python list"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 6,
   "id": "c1324aa0",
   "metadata": {
    "colab": {
     "base_uri": "https://localhost:8080/"
    },
    "id": "NXfPd11Q7pon",
    "outputId": "7850cc4e-8087-41cd-de21-72f8ebf9a1ee"
   },
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "2.318510551995132\n"
     ]
    }
   ],
   "source": [
    "start = time.perf_counter()\n",
    "l = []\n",
    "for i in range(100000):\n",
    "    l.insert(0,i)\n",
    "elapsed = time.perf_counter() - start\n",
    "print(elapsed)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "487c5a59",
   "metadata": {},
   "source": [
    "- Doubling and tripling the work multiplies the time by $4$ and $9$, respectively, so quadratic"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 7,
   "id": "c04aa692",
   "metadata": {
    "colab": {
     "base_uri": "https://localhost:8080/"
    },
    "id": "NXfPd11Q7pon",
    "outputId": "7850cc4e-8087-41cd-de21-72f8ebf9a1ee"
   },
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "7.243922003006446\n"
     ]
    }
   ],
   "source": [
    "start = time.perf_counter()\n",
    "l = []\n",
    "for i in range(200000):\n",
    "    l.insert(0,i)\n",
    "elapsed = time.perf_counter() - start\n",
    "print(elapsed)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 8,
   "id": "520440ca",
   "metadata": {
    "colab": {
     "base_uri": "https://localhost:8080/"
    },
    "id": "NXfPd11Q7pon",
    "outputId": "7850cc4e-8087-41cd-de21-72f8ebf9a1ee"
   },
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "17.871282854001038\n"
     ]
    }
   ],
   "source": [
    "start = time.perf_counter()\n",
    "l = []\n",
    "for i in range(300000):\n",
    "    l.insert(0,i)\n",
    "elapsed = time.perf_counter() - start\n",
    "print(elapsed)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 9,
   "id": "47680763-71c8-431c-9160-2e0b5b43afa9",
   "metadata": {
    "colab": {
     "base_uri": "https://localhost:8080/"
    },
    "id": "NXfPd11Q7pon",
    "outputId": "7850cc4e-8087-41cd-de21-72f8ebf9a1ee"
   },
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "35.851695748991915\n"
     ]
    }
   ],
   "source": [
    "start = time.perf_counter()\n",
    "l = []\n",
    "for i in range(400000):\n",
    "    l.insert(0,i)\n",
    "elapsed = time.perf_counter() - start\n",
    "print(elapsed)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "468142ee-71b3-412e-b912-ef93124c0868",
   "metadata": {},
   "source": [
    "- Another experiment\n",
    "- First create a list with 5000, 10000, ... items\n",
    "- Then do 10000, 20000, ... repetitions of del(l[0]) and l.insert(0,v)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 10,
   "id": "0b56f855-6146-4522-a400-82fc90cfa62f",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "10000 0.02885618900472764\n",
      "20000 0.11141112798941322\n",
      "30000 0.2379553990031127\n",
      "40000 0.4336804780032253\n",
      "50000 0.6927165079978295\n",
      "60000 0.9982619740039809\n",
      "70000 1.3697435880021658\n",
      "80000 1.802433042001212\n",
      "90000 2.3142533600039314\n",
      "100000 2.8545684500131756\n"
     ]
    }
   ],
   "source": [
    "for j in range(1,11):\n",
    "    l = []\n",
    "    for i in range(j*5000):\n",
    "        l.append(i)\n",
    "\n",
    "    start = time.perf_counter()\n",
    "    for i in range(j*10000):\n",
    "        del(l[0])\n",
    "        l.insert(0,i)\n",
    "    elapsed = time.perf_counter() - start\n",
    "    print(j*10000,elapsed)"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3 (ipykernel)",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.13.5"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
