{"id":29,"date":"2026-09-10T13:05:43","date_gmt":"2026-09-10T20:05:43","guid":{"rendered":"https:\/\/blogs.reed.edu\/datalab\/?p=29"},"modified":"2026-09-11T15:13:49","modified_gmt":"2026-09-11T22:13:49","slug":"commons-common-meals","status":"publish","type":"post","link":"https:\/\/blogs.reed.edu\/datalab\/2026\/09\/10\/commons-common-meals\/","title":{"rendered":"Commons&#8217; Common Meals"},"content":{"rendered":"\n<div style=\"border-left: 4px solid #ffc846;padding: 20px 24px;margin: 20px 0;font-size: 1.3em\">The most commonly served meal was Fish &amp; Chips at 14 times! It was closely followed by Fried Tofu &amp; Chips at 12 times. <\/div>\n\n\n\n<p>This week we wanted to know what types of meals were being served at commons. We found out that a ton of data on that is available and the chart in the Quest just skims the surface. If you want to learn more about the data, or use it yourself, keep reading for a walk-through or email <a href=\"mailto:data@reed.edu\">data@reed.edu<\/a>. <\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"768\" src=\"http:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/ingredient_plot-1024x768.png\" alt=\"plot of common main ingredients served by Commons\" class=\"wp-image-19\" srcset=\"https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/ingredient_plot-1024x768.png 1024w, https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/ingredient_plot-300x225.png 300w, https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/ingredient_plot-768x576.png 768w, https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/ingredient_plot-1536x1152.png 1536w, https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/ingredient_plot-2048x1536.png 2048w, https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/ingredient_plot-1200x900.png 1200w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Where did we find the data?<\/h2>\n\n\n\n<p>Every day Bon Appetit&#8217;s menu is posted on the <a href=\"https:\/\/reed.cafebonappetit.com\/cafe\/commons-cafe\/\">Commons homepage<\/a>. Since we wanted to get access to past menus, the first thing to do was see if that information was cached somewhere. When you go to any Commons&#8217; menu that is not todays, you will see the URL look something like this: <a href=\"https:\/\/reed.cafebonappetit.com\/cafe\/commons-cafe\/2026-09-11\/\">https:\/\/reed.cafebonappetit.com\/cafe\/commons-cafe\/2026-09-11\/<\/a> The important thing there is that it ends in the specific date. That means each menu was theoretically stored as its own page that we could access using a web-scraping script. <\/p>\n\n\n\n<h2 class=\"wp-block-heading\">How did we web scrape the data?<\/h2>\n\n\n\n<p>The individual URL dates indicate that each day&#8217;s data exists as an inline JS object in the page&#8217;s HTML code. So we can write a script to access that and pull out the code. I prefer working in R, so I wrote this script in R. But it could have been written in Python or a number of other languages. <\/p>\n\n\n\n<p>There is actually a lot of data contained in each day&#8217;s page. To make the graph simpler (because it has to print in black and white and be legible when small), I decided to only pull dinner menus and only from the SimplyOasis and Classics stations. But more data than this is available if you want to use it!<\/p>\n\n\n\n<p>Here&#8217;s the script. The main things it relies on are the packages <a href=\"https:\/\/httr.r-lib.org\/\">{httr}<\/a> and <a href=\"https:\/\/www.rdocumentation.org\/packages\/jsonlite\/versions\/2.0.0\">{jsonlite}<\/a>. {httr} handles the HTTP request \u2014 downloading each day&#8217;s page \u2014 while {jsonlite} takes the JSON-formatted menu data embedded in that page and turns it into something R can actually work with.<\/p>\n\n\n\n<details>\n<summary style=\"font-size: 1.3em;font-weight: bold;color: #e84a37\">Show the script<\/summary>\n<pre class=\"wp-block-code\"><code>\n\n# load libraries\nlibrary(httr)\nlibrary(jsonlite)\nlibrary(stringr)\nlibrary(readr)\n\n####  set selections for scraping\n###############################\n\n# get url for menu\ncafe_url &lt;- \"https:\/\/reed.cafebonappetit.com\/cafe\/commons-cafe\/%s\/\"\n\n# select what meal (can be single value or vector)\ndaypart &lt;- \"Dinner\"\n# select station (can be single value or vector)\nstation_wanted &lt;- c(\"SimplyOASIS\", \"Classics\")\n# select date range\ndates &lt;- seq(as.Date(\"2026-01-26\"), as.Date(\"2026-05-14\"), by = \"day\")\n\n\n####  create functions to:\n####  fetch webpage,  \n####  extract the menu,\n####  extract the correct part of day\n\n# downloads one date's menu page as raw HTML\nfetch_page &lt;- function(d) {\n  url &lt;- sprintf(cafe_url, format(d, \"%Y-%m-%d\"))\n  resp &lt;- GET(url, timeout(30))\n  stop_for_status(resp)\n  content(resp, as = \"text\", encoding = \"UTF-8\")\n}\n\n# pulls the menu_items from the JavaScript and parses it as JSON\nextract_menu_items &lt;- function(html) {\n  pattern &lt;- regex(\"Bamco\\\\.menu_items\\\\s*=\\\\s*(\\\\{.*?\\\\});\", dotall = TRUE)\n  match &lt;- str_match(html, pattern)\n  fromJSON(match&#091;1, 2], simplifyVector = FALSE)\n}\n\n# pulls the relevant daypart and returns as a named list\nextract_dayparts &lt;- function(html) {\n  pattern &lt;- regex(\n    \"Bamco\\\\.dayparts\\\\&#091;'(\\\\d+)'\\\\]\\\\s*=\\\\s*(\\\\{.*?\\\\});\\\\s*\\\\n\\\\s*\\\\}\\\\)\\\\(\\\\);\",\n    dotall = TRUE\n  )\n  matches &lt;- str_match_all(html, pattern)&#091;&#091;1]]\n  dayparts &lt;- list()\n  for (i in seq_len(nrow(matches))) {\n    dayparts&#091;&#091;matches&#091;i, 2]]] &lt;- fromJSON(matches&#091;i, 3], simplifyVector = FALSE)\n  }\n  dayparts\n}\n\n\n####  create function extract the actual meals\n###############################\n\n# helper function for pulling menu\n# if a is NULL, return b; otherwise return a \n# (can't get combinable rows if NULL values are present)\n`%||%` &lt;- function(a, b) if (is.null(a)) b else a\n\n# extracts meals based on previous parameters\nget_meal_items &lt;- function(d) {\n  html &lt;- fetch_page(d)\n  menu_items &lt;- extract_menu_items(html)\n  dayparts &lt;- extract_dayparts(html)\n\n  # create empty list for output\n  rows &lt;- list()\n  for (dp in dayparts) {\n    # only do this for the correct meal time\n    if (!(dp$label %in% daypart)) next\n\n    # find every station that matches the station(s)_wanted\n    # using Filter because it's lists not dataframes\n    stations &lt;- Filter(\\(s) str_trim(str_remove_all(s$label, \"&lt;&#091;^&gt;]+&gt;\")) %in% station_wanted, dp$stations)\n\n    # make sure there's something there to find\n    if (length(stations) == 0) next\n\n    # for every matching station, grab data for each of its items\n    for (station in stations) {\n      for (item_id in station$items) {\n        # get the item's full list of things\n        item &lt;- menu_items&#091;&#091;item_id]]\n        # make a dataframe out of the following\n        rows&#091;&#091;length(rows) + 1]] &lt;- data.frame(\n          date = as.character(d),\n          meal = dp$label,\n          item_name = item$label,\n          description = item$description %||% \"\", # put \"\" instead of NULL\n          price = as.character(item$price %||% \"\"), # put \"\" instead of NULL\n          dietary_tags = paste(unlist(item$cor_icon), collapse = \", \")\n        )\n      }\n    }\n  }\n  # bind everything together\n  do.call(rbind, rows)\n}\n\n\n####  actually pull the data\n###############################\n\n# create empty list for things to go in\nrows &lt;- list()\n# get meal for each date\nfor (i in seq_along(dates)) {\n  d &lt;- dates&#091;i]\n  rows&#091;&#091;i]] &lt;- get_meal_items(d) # this is the line that actually runs everything\n\n  # prints each date just to show progress \n  cat(as.character(d), \"\\n\")\n  # small pause between requests (technically don't need if it's too slow)\n  Sys.sleep(0.1)\n}\n\nall_meals &lt;- do.call(rbind, rows)\n\n\n####  write file\n###############################\n\nwrite_csv(all_meals, \"data\/dinner_jan26_may14.csv\")\n\n&lt;\/details&gt;\n<\/code><\/pre>\n<\/details>\n<\/ br>\n\n\n\n<p><\/p>\n\n\n\n<p>Out of this script, I get a csv file that looks like this: <\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"865\" height=\"317\" src=\"https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/image.png\" alt=\"\" class=\"wp-image-37\" srcset=\"https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/image.png 865w, https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/image-300x110.png 300w, https:\/\/blogs.reed.edu\/datalab\/files\/2026\/09\/image-768x281.png 768w\" sizes=\"auto, (max-width: 865px) 100vw, 865px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Wrangling the data and graphing<\/h2>\n\n\n\n<p>The way I designed the scraping leaves the data pretty clean, so not a lot needs to be done to work with it. The only thing I need to do is decide what things I want to pull from it. After looking through it, I decided that the best thing would be to look at main ingredients from the item name. I decided to focus on proteins (and mushrooms). Full disclosure, I used Claude to generate this list. I have strong and mixed feelings about AI, but this is one of the uses where it excels and it was much faster than trying to just identify things by most common word because things like &#8220;pasta&#8221; or &#8220;seasoned&#8221; came up a lot.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Meat Substitute:<\/strong> Beyond, Gardein, Chik, Impossible<\/li>\n\n\n\n<li><strong>Chicken:<\/strong> Chicken, Pollo<\/li>\n\n\n\n<li><strong>Beef:<\/strong> Beef, Steak, Brisket, Tri-Tip<\/li>\n\n\n\n<li><strong>Pork:<\/strong> Pork, Sausage, Bratwurst<\/li>\n\n\n\n<li><strong>Turkey:<\/strong> Turkey<\/li>\n\n\n\n<li><strong>Fish:<\/strong> Fish, Salmon<\/li>\n\n\n\n<li><strong>Tofu:<\/strong> Tofu<\/li>\n\n\n\n<li><strong>Mushroom:<\/strong> Mushroom, Portobello<\/li>\n\n\n\n<li><strong>Chickpea:<\/strong> Chickpea, Garbanzo<\/li>\n\n\n\n<li><strong>Soy Tempeh:<\/strong> Tempeh (tempeh not described as chickpea or lentil)<\/li>\n\n\n\n<li><strong>Lentil:<\/strong> Lentil<\/li>\n\n\n\n<li><strong>Bean:<\/strong> Bean<\/li>\n<\/ul>\n\n\n\n<p>I organized the data so that each dish had a column classifying it as the above or &#8220;other&#8221;. Then the script just counted up the frequency of each. I also classified as &#8220;meat&#8221;, &#8220;veggie&#8221;, or &#8220;mixed&#8221; (because other contained meat and veggie dishes). <\/p>\n\n\n\n<p>Then I made a bar graph with flipped coordinates so the foods would be on the y-axis. When you have a lot of bars to show or things with long names, this is a good thing to do. People can see the difference in bars just as easily, and it makes reading the names simpler. So, consider doing this with your bar graphs. <\/p>\n\n\n\n<p>I colored the bars by meat\/veggie\/other. I made other striped and to do this I needed to add a bit of more specialized code with the package <a href=\"https:\/\/coolbutuseless.github.io\/package\/ggpattern\/\">{ggpattern}<\/a>. Then I made adjustments to the axis labels, the theme background, and the display of the gridlines.<\/p>\n\n\n\n<details>\n<summary style=\"font-size: 1.3em;font-weight: bold;color: #e84a37\">Show the script<\/summary>\n<pre class=\"wp-block-code\"><code>\n\n# load library\nlibrary(tidyverse)\nlibrary(ggpattern)\n\n# load data\nraw_menu &lt;- read_csv(\"data\/dinner_jan26_may14.csv\")\n\n# most popular single meal\nmost_popular &lt;- raw_menu |&gt; \n  group_by(item_name) |&gt; \n  summarize(total = n()) |&gt; \n  arrange(desc(total)) |&gt; \n  slice(1) |&gt;\n  pull(item_name)\n\n# search item_name to tag main ingredient\n# group into categories\n# order is hierarchical (helps with sausage vs fake sausage, diff kinds of tempeh)\nmenu &lt;- raw_menu %&gt;%\n  mutate(main_ingredient = case_when(\n    str_detect(item_name, regex(\"Beyond|Gardein|Chik|Impossible\", ignore_case = TRUE)) ~ \"Meat Substitute\",\n    str_detect(item_name, regex(\"Chicken|Pollo\", ignore_case = TRUE)) ~ \"Chicken\",\n    str_detect(item_name, regex(\"Beef|Steak|Brisket|Tri-Tip\", ignore_case = TRUE)) ~ \"Beef\",\n    str_detect(item_name, regex(\"Pork|Sausage|Bratwurst\", ignore_case = TRUE)) ~ \"Pork\",\n    str_detect(item_name, regex(\"Turkey\", ignore_case = TRUE)) ~ \"Turkey\",\n    str_detect(item_name, regex(\"Fish|Salmon\", ignore_case = TRUE)) ~ \"Fish\",\n    str_detect(item_name, regex(\"Tofu\", ignore_case = TRUE)) ~ \"Tofu\",\n    str_detect(item_name, regex(\"Mushroom|Portobello\", ignore_case = TRUE)) ~ \"Mushroom\",\n    str_detect(item_name, regex(\"Tempeh\", ignore_case = TRUE)) &amp;\n      str_detect(item_name, regex(\"Chickpea|Garbanzo\", ignore_case = TRUE)) ~ \"Chickpea\",\n    str_detect(item_name, regex(\"Tempeh\", ignore_case = TRUE)) &amp;\n      !str_detect(item_name, regex(\"Lentil\", ignore_case = TRUE))~ \"Soy Tempeh\",\n    str_detect(item_name, regex(\"Chickpea|Garbanzo\", ignore_case = TRUE)) ~ \"Chickpea\",\n    str_detect(item_name, regex(\"Lentil\", ignore_case = TRUE)) ~ \"Lentil\",\n    str_detect(item_name, regex(\"Bean\", ignore_case = TRUE)) ~ \"Bean\",\n    # real dishes with other main ingredient \n    item_name %in% c(\n      \"Garlic Roasted Lamb\",\n      \"Stuffed Peppers with Smashed Potatoes\",\n      \"Eggplant Stew over Polenta\",\n      \"Seared Yams over Polenta\",\n      \"Vegetable Pozole\",\n      \"Pozole Chile Verde\"\n    ) ~ \"Other\",\n    TRUE ~ NA_character_\n  )) %&gt;%\n  # drop things like ice cream desserts that aren't dinner\n  filter(!is.na(main_ingredient)) \n\n# summarize and arrange by count\nsummary_menu &lt;- menu |&gt; \n  group_by(main_ingredient) |&gt; \n  summarize(count = n()) |&gt; \n  arrange(desc(count))\nsummary_menu\n\n\n# add meat vs veggie\n# other has lamb, so it's a combo of both\nfood_type &lt;- c(\n  \"Chicken\" = \"Meat\",\n  \"Beef\" = \"Meat\",\n  \"Pork\" = \"Meat\",\n  \"Turkey\" = \"Meat\",\n  \"Fish\" = \"Meat\",\n  \"Mushroom\" = \"Veggie\",\n  \"Tofu\" = \"Veggie\",\n  \"Meat Substitute\" = \"Veggie\",\n  \"Soy Tempeh\" = \"Veggie\",\n  \"Chickpea\" = \"Veggie\",\n  \"Lentil\" = \"Veggie\",\n  \"Bean\" = \"Veggie\",\n  \"Other\" = \"Mixed\"\n)\n\nsummary_menu &lt;- summary_menu |&gt;\n  mutate(food = food_type&#091;main_ingredient],\n         food_pattern = case_when(food == \"Mixed\" ~ \"Y\",\n                                  TRUE ~ \"N\"))\n\n# horizontal bar plot\ningredient_plot &lt;- summary_menu |&gt;\n  ggplot(aes(x = reorder(main_ingredient, count), y = count, fill = food, pattern = food_pattern)) +\n  geom_col_pattern(\n    color = NA,\n    pattern_fill = \"black\",\n    pattern_color = NA,\n    pattern_density = 0.3,\n    pattern_spacing = 0.03,\n    pattern_angle = 45\n  ) +\n  coord_flip() +\n  scale_y_continuous(expand = expansion(mult = c(0, 0.05))) +\n  scale_fill_manual(\n    values = c(\"Meat\" = \"gray25\", \"Veggie\" = \"gray65\", \"Mixed\" = \"gray65\"),\n    breaks = c(\"Meat\", \"Veggie\", \"Mixed\"),\n    name = NULL\n  ) +\n  scale_pattern_manual(values = c(\"N\" = \"none\", \"Y\" = \"stripe\")) +\n  guides(\n    pattern = \"none\",\n    # this forces only \"Both\" to look striped\n    fill = guide_legend(override.aes = list(pattern = c(\"none\", \"none\", \"stripe\")))\n  ) +\n  labs(\n    x = NULL, y = \"Number of Times Served\",\n    title = \"What's for Dinner?\") +\n  theme_minimal(base_size = 16) +\n  theme(\n    plot.title = element_text(face = \"bold\", size = 28, hjust = 0.5),\n    plot.subtitle = element_text(size = 13, color = \"black\", margin = margin(b = 10), hjust = 0.5),\n    axis.text.x = element_text(color = \"black\", size = 12),\n    axis.text.y = element_text(color = \"black\", size = 12, margin = margin(r = 0)),\n    axis.title.x = element_text(face = \"bold\", size = 14),\n    legend.text = element_text(size = 14),\n    legend.position = \"top\",\n    panel.grid.minor.y = element_blank(),\n    panel.grid.major.y = element_blank()\n  )\n\ningredient_plot\n\n# save image and a vector copy for print \nggsave(\"ingredient_plot.png\", ingredient_plot, width = 8, height = 6)\nggsave(\"ingredient_plot.svg\", ingredient_plot, width = 8, height = 6)\n<\/code><\/pre>\n<\/details>\n<\/ br>\n\n\n\n<p><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Things you could do with these scripts<\/h2>\n\n\n\n<p>So there&#8217;s a lot more to the data than just what I got. I only looked at last semester, but the data goes back to 2014 (I think, you&#8217;d have to double check). The JSON data has meals from all times of day and all stations present in commons. It also has prices and dietary tags, like gluten free. Here are a few ideas:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Does the menu change when there are holidays approaching? More pie at Thanksgiving or more halal food during Ramadan?<\/li>\n\n\n\n<li>How many new dishes are added each semester or is the menu consistent over time?<\/li>\n\n\n\n<li>How do the Farm to Fork offerings change with seasons?<\/li>\n\n\n\n<li>Was there any shift after Covid?<\/li>\n<\/ul>\n\n\n\n<div style=\"border-left: 4px solid #ffc846;padding: 20px 24px;margin: 20px 0;font-size: 1.3em\">If you&#8217;re interested in exploring any of these, you can email <a href=\"mailto:data@reed.edu\">data@reed.edu<\/a> or come by the <a href=\"https:\/\/www.reed.edu\/data-at-reed\/data_lab\/\">DataLab<\/a>!<\/div>\n","protected":false},"excerpt":{"rendered":"<p>The most commonly served meal was Fish &amp; Chips at 14 times! It was closely followed by Fried Tofu &amp; Chips at 12 times. This week we wanted to know what types of meals were being served at commons. We&nbsp;&hellip; <a href=\"https:\/\/blogs.reed.edu\/datalab\/2026\/09\/10\/commons-common-meals\/\">finish&nbsp;reading&nbsp;Commons&#8217; Common Meals<\/a><\/p>\n","protected":false},"author":3079,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[4],"class_list":["post-29","post","type-post","status-publish","format-standard","hentry","category-uncategorized","tag-r"],"_links":{"self":[{"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/posts\/29","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/users\/3079"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/comments?post=29"}],"version-history":[{"count":18,"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/posts\/29\/revisions"}],"predecessor-version":[{"id":54,"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/posts\/29\/revisions\/54"}],"wp:attachment":[{"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/media?parent=29"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/categories?post=29"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.reed.edu\/datalab\/wp-json\/wp\/v2\/tags?post=29"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}