There are many Masters in Data Science programs that have been developed over the last few years including at universities, massive open online courses (MOOCs) such as the Data Science Specialization on Coursera or if you're really self motivated an Open Source Data Science Masters curriculum that you can work through at your own pace. A few weeks ago I was happy to see my graduate school alma mater (Rice University in Houston, TX) announced that it was awarded a $1.4 million dollar grant from the NSF to start a Data Science program between faculty in the Statistics and Computer Science departments. The money will be used to develop courses by faculty in both departments and to allow undergraduate students, graduate students and postdoctoral researchers to take courses and complete research projects. This comes after the university announced last fall a $43 million dollar investment of establishing a world-class program in data science. I look forward to seeing exciting things from these two initiatives!
Congrats to Rice!
Wednesday, August 3, 2016
Friday, July 29, 2016
Attending JSM Conference 2016 in Chicago
This weekend I'm headed to Chicago to the big, annual conference for statisticians in North America: the Joint Statistical Meetings (JSM). I'll be there Sunday through Wednesday afternoon and I'm giving a talk on Tuesday in Session #405 titled 'Statistical Challenges in the Analysis of Single-Cell RNA-Seq Data'. In preparation for the conference, I started a set of notes on GitHub of the sessions I'm mostly interested in attending. As I attend various sessions, I'll add more notes of software, papers and links from the talks. Looking forward to catching up with some old friends and meeting new friends!
Wednesday, July 20, 2016
Problems installing latest R version on cluster
For my research, I often need to access R on a cluster instead of on my personal laptop. This week I tried to update to the latest version of R-3.3.1 and I found out it is no longer as simple as it used to be. Previously, I would just follow my notes on installing R and R packages locally which is posted on GitHub. To install the latest R (currently version 3.3), I needed to access newer versions of some libraries (zlib, bzip2, curl, pcre and xz). The './configure' step would fail without access to the updated libraries. Usually these types of packages would be included in the installation of R on linux, but it is no longer included. Instead, it assumes these libraries are already installed and updated.
This led to some frustration. Then a postdoc in our lab (Mike Love) sent me this wonderful step-by-step tutorial on getting the updated libraries and successfully installing the latest R. As the tutorial mentions, if the system admins update these libraries, you won't run into this problem. If they don't keep this libraries up-to-date, then you will find this tutorial to be very helpful.
P.S. I tweeted the link and Gabe Becker responded that he had a more complicated set up. His solution was to statically link to the libraries.
This led to some frustration. Then a postdoc in our lab (Mike Love) sent me this wonderful step-by-step tutorial on getting the updated libraries and successfully installing the latest R. As the tutorial mentions, if the system admins update these libraries, you won't run into this problem. If they don't keep this libraries up-to-date, then you will find this tutorial to be very helpful.
P.S. I tweeted the link and Gabe Becker responded that he had a more complicated set up. His solution was to statically link to the libraries.
@stephaniehicks Ive been going through this, but our R has to live in larger ecosystems, so had to statically link libz, bz2, pcre and curl— Gabe Becker (@groundwalkergmb) July 20, 2016
Sunday, July 17, 2016
Advice on Following an Academic Career Path
I've been putting together some general advice on following an academic career path. This ranges from being a graduate student to faculty at a university or institution. I put the information in a GitHub repo, so more information could be added be me and others (pull requests are welcomed!) over time. The advice generally comes from being in an applied statistics / genomics field, but I think it mostly applies to a broader set of fields. As I move through my career, I will continue to add resources, links and notes to it. Hopefully others will find it useful!
Thursday, July 14, 2016
Maryland-Style Crab Cakes
After recently visiting Baltimore, we wanted to learn more about the local food culture. For anyone who has been there, you'll know that crab cakes and Old Bay seasoning are staples of the area.
Old Bay was originally developed by a Jewish man named Gustav Brunn, who fled Germany in 1939. After establishing himself on the Chesapeake Bay (which stretches from Virginia Beach to the Pennsylvania border), his seasoning took the name "Old Bay" from the iconic overnight mail service that ran from Baltimore, MD to Norfolk, VA. The Baltimore Steam Packet Company was nicknamed "the Old Bay Line," which survived attacks during the The Civil War and government requisition for WWII. After reopening as an automobile ferry service, the company went under... but the Old Bay seasoning continued to gain popularity as a seafood seasoning, since it kept tavern-goers thirsty. In 1990, McCormick & Co bought the Old Bay brand (with Zatarain's and Lowry's following).
Contents of Old Bay seasoning are:
- celery salt (ground celery seed and salt)
- paprika
- black pepper
- cayenne pepper
- ground mustard powder
- bay leaf powder
- other trace spices (cinnamon, allspice, nutmeg, ginger, clove, cardamom, etc)
Today's recipe is based on one developed by Jenn @ Once Upon A Chef: LINK
- 2 large eggs
- 2-1/2 tablespoons mayonnaise
- 1 teaspoon Worcestershire sauce
- 1 teaspoon Old Bay seasoning
- 1 teaspoon garlic powder
- 1/2 teaspoon chili powder
- 1/2 teaspoon black pepper
- 1/4 teaspoon salt
- pinch of ground mustard
- 1 large stalk of celery, finely diced
- 1/4 large onion, finely diced
- 2 tablespoons celery leaves, finely chopped
- 1 pound lump crab meat
- ~1 cup panko
- Vegetable or canola oil, for cooking
To the chopping block!
Finely chop celery and onion.
Mix together all ingredients.
Scoop the melange into patties.
Shallow fry the patties in canola oil.
Flip when they are golden brown (and delicious).
Choose your favorite sauce and enjoy!
Monday, March 14, 2016
Swedish Almond Cake
Happy Pi Day! Today is my one my husband's favorite days of the year, partly because pie usually appears in our kitchen around this time, but mostly because we are both have a love for all things math-y.
Spring has definitely arrived early to our town and for some reason I have been on an almond kick lately. While I know this recipe is not officially considered a "pie", it was what I was craving this weekend. I found this recipe for a Swedish Almond Cake that looked too good to pass up. I did not have sliced almonds, so I substituted chopping up almonds myself.
Hope you enjoy as much as we did!
Ingredients:
- 1 cup granulated sugar
- zest from 1 lemon
- 1/2 tsp almond extract
- 1/2 tsp vanilla extract
- 1/4 tsp salt
- 2 eggs
- 1 cup all-purpose flour
- 1/2 cup butter, melted (1 stick)
- 1/8 cup almonds, chopped
- 1 tbsp coarse sugar (or I just used granulated sugar)
Recipe:
1) In a bowl, combine granulated sugar and the zest from 1 lemon.
2) Add vanilla and almond extract and salt. Mix until sugar mixture is well blended.
3) Add in eggs one at a time. Mix until wet ingredients are frothy.
4) To the wet ingredients, add 1 stick of melted butter and 1 cup flour. Grease an 8 or 9 inch cake pan or springform pan. Add cake batter to the pan. Top with chopped (or sliced) almonds and coarse sugar.
5) Bake at 350 degrees for 25 minutes. Remove from the oven and let cool for 20 minutes before taking it out of the pan.
Unfortunately, I forgot to take a picture of the final cake after it was done (yes, it was that good)! We finished almost half it over the weekend before we realized that. So, I leave you with a picture our half eaten almond cake.
Highly recommended recipe! :)
Thursday, February 25, 2016
Using Version Control and GitHub in the Classroom
This semester I'm co-instructing a course called Introduction to Data Science (BIO 260 and CSCI E-107) with Rafael Irizarry at the Harvard School of Public Health and Harvard Extension School. It is similar to a course that I was the head TA for in Fall 2014 taught at Harvard University called CS 109. We have a fantastic group of people involved with the course this year, which has made developing a course from scratch run much more smoothly than it could have been.
We spent one lecture teaching students the importance of version control. Version control is a way of tracking the change history of a project. Even if you have never heard of version control, you have probably already done it manually. For example, if you have ever written a document or paper, you may have tried copying and renaming the file multiple times as it went through different stages ("paper-v1.doc", "paper-v2.doc", "paper-final.doc", "paper_finalFINALdraft.doc", etc.). If at any point you wanted to see an older version of the paper, you could simply open the file. That is a form of version control. It's not very efficient, but it is in fact a form of version control. One improvement over this would be to have a way only keeping one file (e.g. "paper.doc") AND being able to see older versions of it as it changed through time. You can think of these older versions as little snapshots of the paper as it changed through time. The same idea can be applied to code that you write.
In data science, it's important to know how to keep track of your code as it changes over time. On top of that, when you are writing code in a collaborative setting, it is almost required that that you know something about version control. This is how a group of people can collaboratively contribute code to the same project using the same file. Git is a tool that automates and enhances a lot of the tasks that arise when dealing with larger, longer-living, and collaborative projects. It has also become the common underpinning to many popular online code repositories, GitHub being the most popular.
We spent one lecture teaching students the importance of version control. Version control is a way of tracking the change history of a project. Even if you have never heard of version control, you have probably already done it manually. For example, if you have ever written a document or paper, you may have tried copying and renaming the file multiple times as it went through different stages ("paper-v1.doc", "paper-v2.doc", "paper-final.doc", "paper_finalFINALdraft.doc", etc.). If at any point you wanted to see an older version of the paper, you could simply open the file. That is a form of version control. It's not very efficient, but it is in fact a form of version control. One improvement over this would be to have a way only keeping one file (e.g. "paper.doc") AND being able to see older versions of it as it changed through time. You can think of these older versions as little snapshots of the paper as it changed through time. The same idea can be applied to code that you write.
In data science, it's important to know how to keep track of your code as it changes over time. On top of that, when you are writing code in a collaborative setting, it is almost required that that you know something about version control. This is how a group of people can collaboratively contribute code to the same project using the same file. Git is a tool that automates and enhances a lot of the tasks that arise when dealing with larger, longer-living, and collaborative projects. It has also become the common underpinning to many popular online code repositories, GitHub being the most popular.
One of the unique aspects of the course is that we are requiring the students to work on their homework assignments and submit their homework assignments using GitHub. We started by creating the GitHub organization datasciencelabs-students. Then, we followed the GitHub Education Classroom guide. To create private repositories for each student for each homework assignment, we followed the sandboxing setup. Sandboxing is the idea of creating duplicated repositories (one for each student) in an automated fashion. The tool that physically creates the private repositories for each student for each homework assignment is called "teachers_pet".
For anyone interested in using teachers_pet in your own classroom, I created a set of notes on GitHub describing how to install and use the tool. Once you have set up all the authentication steps with GitHub, you can clone my teachers_pet GitHub repository, which is an enhanced version made up of multiple repositories from the web. It is enhanced because it can push files to specific branches in the private repositories AND it can delete repositories at the end of your course so you can re-use the private repositories in the future.
I hope others find the teachers_pet tutorial useful in using GitHub in the classroom!
Subscribe to:
Posts (Atom)










