Showing posts with label conference. Show all posts
Showing posts with label conference. Show all posts

Friday, August 17, 2018

How to successfully submit a conference session proposal

Conferences are a wonderful place to learn about exciting research happening in your field, to meet new people and/or potential collaborators and to catch up with old friends. Recently, I attended the annual Joint Statistical Meetings (JSM) in Vancouver, Canada, which is the largest conference for statisticians in North America. One of the features of this conference is the program committee invites anyone to submit a proposal for an Invited Session. In terms of JSM, the due date is almost a year in advance of the conference. For example, Invited Session proposals for JSM 2019 (July 27-Aug 1, 2019) are due September 6, 2018. There are also other types of sessions called Topic-contributed Sessions and Contributed Sessions with due dates a bit later in the year. The problem is because this conference is so large, it requires a long time to read through the proposals and to organize all the sessions.

In previous years, I helped organize and participated in an invited session at JSM. This year, I submitted a proposal for a topic-contributed session, but it was elevated to a late-breaking session. Most recently, I organized an Invited Session, which was just accepted, for the 2019 Eastern North American Region (ENAR) International Biometric Society conference. Finally, I am on the program committee for the 2019 Symposium on Data Science and Statistics (SDSS).

Given my recent experiences organizing sessions for conferences and the JSM 2019 Invited Session proposals are due in a few weeks, I thought I would it would be relevant to write a blog post on strategies I use when putting together a proposal.

1. Come up with a interesting, timely and relevant topic. You might bounce ideas off of your colleagues to see if the topic would have wide enough interest and might be of interest or be relevant to a particular conference. This is often the most difficult part of organizing a session. If your topic is not really of interest to a wide enough audience, it is highly unlikely that it will be selected. However, you want it to be focused enough that you can reasonably talk about the topic in 1-2 hours.

2. Create a title and abstract for your session proposal. If you have thought carefully about the topic, this should be an easier step. The title should be succinct and representative of what you want your session to be about. The abstract should contain (1) why this topic is important and relevant, (2) the focus and goal of the session, (3) what the speakers will discuss. Ask a colleague to review the title and abstract to give you feedback.

3. Decide on speakers and send out invitation emails. First, think carefully about who is in your audience and who you want to invite. Some things to think about when coming up a list of potential speakers: their backgrounds, their expertise, their perspective, their ability to give good presentations and the diversity of the speakers. The last one is most commonly overlooked, but can bring such rich and valuable discussions if you have a diverse set of speakers.

Here is a suggested format to send the invitation emails:

Hi ____, 

I am organizing a <add name of session> session for the <add name of conference> conference in <add location> next year. The session will be focused on <add topic of session>. 

My goal with this session is for <add goal of session>. Instead of focusing on <a previously discussed topic>, I want to focus on <a new topic>. My hope for the session is that audience members will be able to <add what you want the audience to get out of the session>. 

As the <add the person's title, etc>, I would like to invite you to speak in the session <(or) join as a panel member to share your insight and perspectives (if a panel)>. Your expertise in <all the reasons why this person would be a good speaker> would be highly valuable and a great contribution to the session. 

I hope you will join the session if you plan to attend and aren’t otherwise committed. Could you let me know by <fill in the date> if you would be willing to speak? I plan to include 3-4 speakers <(or) panel members> and welcome suggestions for additional speakers. 

I am happy to answer any other questions that you may have. 

All the best, 
<add your name here>
<add your affiliation here>

Some will say yes, some will say no. If needed, send out more email requests. The main things are to explain (1) the details of the conference and session, (2) the focus and goal of the session, (3) why you are inviting them or why you think they would be a good contribution to the session, (4) and the date you need for them to respond to you by.

4. Submit a proposal to the conference by the due date. This usually includes at minium a title and an abstract. You also want to include the name of all the speakers who have agreed to participate in your session, their affiliations (departments / institutions / company name, etc), and usually an email address. This helps the conference organizers get a better idea of what will be discussed. This last one is often required, so check out what is needed for your specific conference.

5. Wait for a response from the conference organizers. This can be a quick or very long process, depending on how large the conference is. Typically the larger the conference, the longer it takes to go through all the proposals. JSM's Invited Session proposals are notoriously long:

And my favorite response from the amazing Shannon Ellis:

If your proposal is accepted, CONGRATULATIONS! Send an email to all the speakers to share the good news. They will need to register and may need to submit individual abstracts for each of their respective talks. Remind the speakers of any upcoming deadlines. If not, consider trying again next year!

Most importantly, if you are a student or postdoc, organizing a conference session can be a fantastic way to meet people in your field! You have the opportunity to craft a session on a topic that excites you the most and chair that session. Conferences need fresh perspectives and new ideas, and I find students and postdocs actively working on a research topic have some of the most insightful suggestions.

Wednesday, June 6, 2018

Addressing Sexual Misconduct in Statistics

The cultural revolution from the recent #metoo movement has demonstrated that sexual harassment and sexual assault are far too common and often go unreported in professional settings, including in the field of Statistics. In response to this, I have organized a late breaking session at the Joint Statistical Meetings on Addressing Sexual Misconduct in Statistics on Monday July 30, 2018 2-3:50pm in Vancouver, Canada.

The focus of this session is to bring together seven panelists to discuss how sexual misconduct can negatively impact the careers of the victims and how we as a community can make positive changes. Our goal is to open the dialogue on recognizing and condemning predatory sexual behavior, and provide support and inclusion to all members of the statistics community.

This session will be a positive discussion to address this topic and will offer perspectives from conference participants, elected leaders of professional societies, academic journal editors, academic department chairs, and program committee members. In addition, this session will also include the perspective of individuals who have been directly impacted by sexual harassment or assault, which is particularly relevant to other individuals in the audience who have also been impacted by sexual harassment or assault, but may not have discussed it publicly. Keegan Korthauer will chair the session and our panelists will include:

If you are planning on attending #JSM2018, I invite you to join us.  

Thursday, January 5, 2017

Women in Statistics and Data Science Conference 2016

Happy New Year everyone! After a wonderful holiday break, I was excited to find my copy of AMSTAT News from the American Statistical Association in my mailbox! Someone pointed out to me that if you look close enough, I can be found standing in the background on the front cover of the Dec 2016 issue. So, apparently I can check off 'being on the cover a magazine' from my bucket list. ;)


The picture was a snapshot of an audience at talk from the Women in Statistics and Data Science Conference (WSDS) 2016 which was held in Charlotte, North Carolina Oct 20-22, 2016. The cover story can be found here. In 2014, I attended the inaugural Women in Statistics Conference and was fortunate enough to attend the conference again this year! Looking back through blogposts this fall, I realized I did not write a blog entry after I returned from the conference, but I did manage keep a few notes in a Github repo of a few of the talks. Here, I want to summarize some of my thoughts and experiences of WSDS 2016 and at the end I describe a few suggestions for future WSDS conferences. I was not able to attend all the talks mostly because there were many concurrent sessions happening at the same time, but I hope this highlights at least a portion of the conference!

The picture above was taken in a talk by one of my favorite speakers from the 2016 conference, Erin Anika Wiley from Westat, titled "Do you Hear What I Hear?: An Examination of Effective Communication". I managed to slip into the back of jammed packed room and furiously write down notes on the results of a survey she conducted asking about opinions of presentations from statisticians. I love that this was the front cover of the magazine because Erin really had the audience laughing and engaged for her entire talk. Also, it was humbling and enlightening to hear the survey responses on how non-statisticians perceive statisticians based on talks and presentations. Definitely motivation for how we can communicate our results more effectively!


The first keynote address was from Cynthia Clark titled "Consider your legacy". She gave inspirational talk discussing what her contributions have been in her personal and professional life. As a new mother, I sincerely appreciated hearing how she prioritized her family throughout her life.



Another great keynote address was from Stacy Lindborg at Biogen with a talk titled "Know your power". Stacy is a natural at connecting with the audience by sharing personal stories from her life. In this talk, she shared five reflections/tips on having a successful career in the face of many challenges (both personal and professional). One of my favorite quotes from her talk was "We love the things that we are good at", which really resonated with me.

Keeping with theme of careers, a topic I found particularly interesting is how to navigate changes in your career. Michelle Dunn, Donna LaLonde, and Nancy Flournoy shared very honest and personal experiences of following non-traditional career paths, recognizing and moving forward when you have failed, and knowing when to leave a job. A central theme from each of their talks was to always be growing your network and find trusted individuals/mentors to help you navigate these experiences.

Similar to the session in 2014, the 2016 conference included a fantastic panel of the past, current and future ASA Presidents. These women were able to persevere in the face of a highly male-dominated field to become ASA Presidents and become role models for women in the early stages of their career like me. Even with all this progress, I agree with Mary Ellen Bock that there is so much more work still to do with supporting more specific minority groups in this field. For example, she eloquently described her hope of one day seeing a non-white woman on this panel of presidents.


In my opinion, one of the coolest sessions was listening to the amazing Mary W. Gray from American University giving a fascinating discussion on what US laws exist to protect women. If you don't know who she is, you can read a summary on wikipedia, which is pretty incredible! One of my favorite things that I learned about her was that she wears a 3/4 euro lapel pin (noting pay inequality is not limited to the US) to promote "equal pay for equal work" referring to the Equal Pay Act in 1963 and Title VII in 1964.  Sadly I didn't get a picture of the lapel pin, but I love the idea!


Wendy Martinez from the Bureau of Labor Statistics gave a great keynote address on what and how federal statistical agencies are thinking about when it comes to data science. Data science is something I am passionate about, so I really enjoyed this talk. I would have loved to see a larger emphasis on data science education though.


Unfortunately Bin Yu from UC Berkeley was unable to join us in person, but the conference organizers were able to set up a video chat and connect to the big screen!  Her talk was titled "A holistic approach to interdisciplinary research" where she began discussing how she has a people centric view of life and research (people are mysteries to unveil just like research). It was amazing to hear a little bit about her background and how she grew up in China during the cultural revolution (1966-1976) where the universities stopped for 10 years. Something I think that cannot be overstated was she noted that she most appreciates intellectual diversity in collaboration and research. Overall, it was an inspiring talk!


The conference ended with a festive dinner filled with new friends, awards, presentations and delicious desserts!


As the conference took place at a hotel, I managed to find a nice location to write down some notes from the talks on my GitHub page. :) Feel free to check them out for more details on the talks that I attended.



Finally, I want to finish this blogpost with a few suggestions for future WSDS conferences. My intention is for these suggestions to be viewed as constructive to help make this conference stronger and more accessible to women in the future.

  1. Fewer talks running concurrently and more posters/speed sessions. I appreciated that the keynote talks did not have other scheduled talks, but at other times I had to choose between 5-6 sessions running concurrently, which led to frustration and disappointment that I missed so many other talks. My arguments for this are similar to the ideas previously described to improve JSM (decreasing the number of contributed sessions and increasing the presence and importance of poster sessions). 
  2. More talks from women in data scientist positions in academia, government, and industry. I recognize that the term "data science" is very much being actively debated in many settings, but as this conference now has "data science" in the title, I would love to see a larger discussion of data science and how that relates to statistics, here too. There was such a great representation of perspectives from women who pioneered this field. I think it would be equally beneficial to include the perspective from a new generation of women in data science positions. 
  3. Childcare. I was surprised to find out that no childcare was available at this conference. Considering a good portion of this conference was dedicated towards discussing issues related to balancing careers with families, I found a lack of childcare a bit ironic (?). This suggestion is motivated by my own personal experience of recently having a baby. I saw several other WSDS attendees who brought their children and probably could have benefited from a childcare service. There are great childcare examples broken down by cost, size of conference, childcare agency, fees charged, etc.  I hope this will be incorporated in future WSDS conferences. 
Thank you to all the conference organizers for putting together this conference. I look forward to the next one in La Jolla, CA Oct 19-21, 2017! 

Wednesday, September 21, 2016

2016 Single Cell Genomics Conference

Last week I attended the 2016 Single Cell Genomics (SCG) Conference at the Wellcome Genome Campus in Hinxton, Cambridge, UK. This conference started in 2013 at the Weizmann Institute of Science, moved to the Karolina Institute in 2014, and was most recently held in the Hubrecht Institute in 2015. A PDF of the program is located here and my notes on the talks including links to various methods, publications and twitter accounts for the speakers are on GitHub. You could follow the tweets online using the twitter hashtag of #SCGen16.
This was a fantastic conference that I would highly suggest to anyone working in the single cell genomics world. Some random take aways of things I noticed compared to other conferences I recently attended were:

1. I was bummed that many of the speakers did not have their talks streamed. As I attended in person, I was fortunate enough to see all the talks, but apparently there was a pretty big waitlist for this conference. The organizers decided to offer a stream for a reduced registration fee with the caveat that not all talks may be streamed.  I assume this was primarily driven by the fact that many presenters discussed unpublished work, but still I'm sure this was frustrating for the waitlisted individuals who were not able to attend in person.


2. Similar to #1, almost none of the speakers made their slides available online (even the ones that agreed to have their talk streamed). I recently attend the Joint Statistical Meetings in Chicago, IL July 31-Aug 3 this year where Karl Broman kept a great list of links to talks from speakers at JSM 2016 (I also kept my own notes). I was really hoping to be able to go back and review slides after the conference ended, but to my knowledge, I was the only one to post slides. To be fair, JSM is significantly larger than the single-cell genomics conference, but I was expecting at least some speakers to post their slides online. I gave a talk on progress in deal with batch effects and biases in single-cell RNA-seq data.

3. The power of the pre-print! Much of the work discussed was either published or in pre-prints.  Sarah Teichmann presented work on the sensitivity, specificity and accuracy of different scRNA-seq protocols using ERCC spike-ins and has a pre-print available on bioRxiv. She humorously pointed out a sense of relief after Lior Pachter sounding positive about the paper by tweeting it.

4. As someone who's very much interested in promoting women in STEM, I was interested to see the ratio of female to male speakers and attendees at the conference.  Though there were fewer women, I was excited to see such prominent women showcasing their work, but as always, there is room for improvement. Sean Davis from NCI has a single cell GitHub repository, which contains list of people (males and females) working in this field. It would be great place to start for looking for additional female speakers.

5. Finally, there were many fantastic biological talks and a lesser number of equally wonderful methods talks, but a few of my favorite methods talks included:


6. Lack of twitter presence. I don't believe having a twitter presence is relevant to having good research in genomics, but I have noticed a large community of people who work in genomics on twitter.  Therefore, it was interesting to me that so few of the speakers at this conference had a presence on twitter. Just for my own curiosity, I did a quick analysis of the tweets from there conference [code in RMarkdown] to find out the top 50 most tweeted words at the conference (after removing the #scgen16 hashtag).


If you want to know more about some of the talks, I have some notes here on GitHub.

Friday, July 29, 2016

Attending JSM Conference 2016 in Chicago

This weekend I'm headed to Chicago to the big, annual conference for statisticians in North America: the Joint Statistical Meetings (JSM). I'll be there Sunday through Wednesday afternoon and I'm giving a talk on Tuesday in Session #405 titled 'Statistical Challenges in the Analysis of Single-Cell RNA-Seq Data'. In preparation for the conference, I started a set of notes on GitHub of the sessions I'm mostly interested in attending. As I attend various sessions, I'll add more notes of software, papers and links from the talks.  Looking forward to catching up with some old friends and meeting new friends!

Tuesday, May 20, 2014

Inaugural Women in Statistics 2014: Highlights and Discussion Points

This week I attended the Women in Statistics conference which was held May 15-17 in the Raleigh-Durham area in North Carolina. I wrote a blog post prior to the conference and this is my follow up post. The theme of the conference was "Know Your Power" in which women discussed transformative moments in their lives and discussed ways to make positive changes in our field. To see more details on individual talks, you can search for tweets with the hashtag #WiS2014. The conference was filled with phenomenal talks/discussions, but I want to give a few highlights from the conference. 


[Pictured (bottom row, left to right): Stephanie Hicks, Jenna Krall, Alyson Wilson, Alicia Carriquiry]
[Pictured (top row, left to right): Cal Tate Moore, Rachel Schutt, Sally C. Morton, Samantha Tyner]

Here are a few key discussion points I took away from the conference:
  1. Social media (blogging, Twitter, LinkedIn, etc) is a great way to build a brand for yourself. Arati Mejdal gave several examples of statisticians and data scientists who have done this such as Hilary Mason (popular blog and twitter feed), Emma Pierson (recent graduate from Stanford who wrote a hilarious article on FiveThirtyEight showing people really just want to date themselves) and Andrew Gelman who says he uses his blog as a way to "steer statistics in a useful way". Two key points to make the most of social media are post regularly and actively comment / engage in discussions. As statisticians or data scientists, the best posts are visual and brief and they are different from academic articles (expert, but friendly).  
  2. Start networking now. Alicia Carriquiry gave a beautiful talk on how to build and nurture your professional network.  Some of the advice included: attend professional meetings, never turn down the opportunity to present your work, chat with people who have similar interests and those who have different interests, be willing to introduce yourself to people you would like to meet, create & practice your elevator pitch and get objective reviews of your performance early in your career.  If you are a young professor, invite other young professors from different departments to give talks and you may have the opportunity to do the same in their department. Jessica Utts (newly elected ASA president for 2016) said she came to "know her power" when she recognized the value of networking.  
  3. Do what makes you happy. It does not matter if your career takes you into academia, industry, government or a bit of all three: as Sally Morton said "Go where you will have the most impact and be most happy. If you are happy, that's where you'll be the most productive".  Rachel Schutt discussed how she did not know at the time how all the pieces of her career (e.g. graduate school, teaching, working at Google, professor Columbia University, etc) would come to fit together at current position. She just did what made her happy. Francesca Dominici led a discussion on Why women can't have it all? in which she stated "It is OK to want to spend time with your children. It OK to be passionate and committed about your work". She argued "a new definition of academic success should be defined to include rewards for teaching and mentoring".  No simple fix, but rather there needs to be a cultural change amongst both men and women to redefine the idea of "academic success". 
  4. The Imposter Syndrome is a real thing. Don't be discouraged by it, but rather recognize the problem if it's affecting you and focus your strengths. Focus on what you have accomplished versus the things you have not. The imposter syndrome is not the same thing as low self-esteem: low self-esteem is boosted when you have a success, but the imposter syndrome makes you feel more terrified if you have a success. For some additional thoughts on this, check out Lean In: Women, Work and the Will to Lead and The Confidence Gap. 
  5. Grace Wahba is simply a hero.  I'm not sure I could ever do her talk justice by trying to summarize it. I will say listening to her talk about her early career was a very surreal and a humbling experience. I feel fortunate to not have to face many of the challenges she faced, but listening to her talk was one of the highlights of the entire conference fore me!  I just encourage everyone to attend her COPSS Fisher Lecture at JSM August 6, 2014 at 4pm.  
Final thoughts: The conference was filled with enlightening talks from speakers of all backgrounds and of all ages who challenged the conference participants to "know your power" through sharing their own stories and experiences. These women are an inspiration and I know many younger women attending the conference felt very encouraged to take on the challenges that lie ahead of us.  I learned a great deal of professional and career development tools and felt men could have just as easily benefited from them too.  Thank you to the organizers and everyone who spent countless hours putting together an extraordinary conference.  I would highly recommend Women in Statistics to future participants!

I leave you with a few more pictures from the conference:

Panel of past and future presents of the American Statistical Association

Mixing and mingling at the poster session Friday night

A little bit of fun: superhero statisticians to the rescue (post-poster session)! 

Sally Morton sharing some of her experiences from the conference including her first "selfie" 

 The amazing Grace Wahba and her "Ah-ha" moments

Thanks to all sponsors.
Platinum: Duke U, NIGMS/NIH, ASA, Minerva Research Foundation, Walmart
Gold: IBM, Lowe's
Silver: Biogen Idec, Experian, Lilly, Minitab, Morestream, SAS
Bronze: Berry Consultants, Cytel, JMP, Nielsen, NC State, Rho, RTI, Stata, UNC, Westat

Wednesday, May 14, 2014

The inaugural Women in Statistics Conference

This week is the inaugural Women in Statistics Conference being held May 15-17, 2014 in Cary, North Carolina. This conference is targeted at women at varying stages starting from graduate school all the way through tenured professors or well-funded CEOs in industry.  As a female statistician (and a postdoctoral fellow), I am very excited to attend this conference celebrating women in statistics!  Here are a few of the reasons why: 
  1. The opportunity to listen to and to interact with an entire community of female statisticians from industry, academia & government is one of the most attractive aspects of this conference. Not only will these talks/breakout sessions focus on a diverse set of career opportunities, they will also focus on useful topics on how to obtain these positions e.g. Answering tricky interview questions, Things I wish I knew when I started working, Optimizing your job search, How to negotiate what you are worth, The value of internships, Preparing for promotion in academia, etc. These are all topics both men and women in our field can benefit from, so I plan to create a second blog post summarizing ideas/notes that are relevant for the entire statistical community.  
  2. Statistics as a discipline is currently facing its own set of challenges within the larger community of Science, Technology, Engineering and Mathematics (STEM), one including being able to attract women to the STEM fields. Many people have suggested ideas and discussed reasons why this is happening. I cannot speak for other women, but I can say one of the reasons why I am I where I am today is the copious amount of support that I have receieved from not only my family and friends, but most importantly from my mentors, faculty advisors and peers.  I was fortunate enough to not have "a terrible graduate school experience", but rather one filled with mentoring, guidance and patience. I know this conference will also be filled with mentoring and guidance from other female statisticians, many of which I consider to be role models. Conferences like this provide women with the information and tools needed to thrive not only in statistics, but in the larger STEM fields as well.  
  3. The idea of gender inequality in the field of statistics is not a new story, but it has been recently discussed in several articles. Ingram Olkin and Terry Speed both discussed the fact that at JSM 2012 "of the four named lectures (i.e. Wald, Rietz, Neyman, Fisher), the seven medallion lectures, and the two invited lectures, none of them were women". Amanda Golbeck wrote an Op-ed titled Where Are the Women in the JSM Registration Guide? in which stated "a productive way to help recruit, retain and nourish women professionals is to provide strong role models for them".  I completely agree and this conference will discuss several of these issues in talks and breakout sessions on topics such as Increasing Visibility of Women in Statistics, Increasing the Number of Women Awards, Recruiting and Retaining Women and Minorities in Statistical Science, Women in Science: Contributions, Inspirations, and Rewards, and Finding Our Place in History: Decades of Women Pioneers and Trail Blazers to name a few.  
  4. I'm particularly excited about the Internet Activism: Using social media to enhance your career breakout session. I learned this idea of using social media as tool to keep up with the literature & make a internet presence for yourself fairly late in my graduate school career. It's a way academic departments and industries can learn about your research interests and contributions.  I know the use of social media has absolutely transformed the way I function as a researcher. I was introduced to this idea actually from the genetics/genomics community by attending the American Society of Human Genetics for the past several years.  I think statisticians haven't quite caught on to the social media bug like the world of genomics, but as statistics departments are grappling with the debate of adsorbing statistics into incredibly popular emerging field of "data science", this is a topic I think many statisticians would find particularly useful.  
In addition, I have been asked to lead a discussion on Taking on Leadership Positions on Saturday morning.  I thought about what questions might be the most useful to ask and here are a few ideas that I have come up with:
  1. What defines a good leader? Is it innovation, focus, communication, ability to hire creative people with diverse backgrounds, ability to risk failing?  Some articles I found relevant were the Harvard Business Review put out an article on Real Leadership Lessons of Steve Jobs and the Forbes Women Leaders Must Dive In, Not Just Lean In. What other articles are good reference points? 
  2. Who are some examples of great leaders inside or outside the field of statistics? 
  3. What are some examples of positions require leadership skills inside or outside the field of statistics?  What do these positions have in common? 
  4. In what ways might someone who does not have an innate ability to lead learn to lead?   Are the qualities (from Q1) usually inherited or can they be learned?  
  5. What are the different styles of leaders? 
  6. How do you balance a position of leadership and maintain a balanced life either with your research and/or family life? 

I welcome other thoughts/suggestions! I plan to live tweet as many talks/breakout sessions as I can (you can follow me @stephaniehicks), but I will definitely write a second blogpost summarizing my thoughts and key points taken away from the conference.  

Thursday, February 6, 2014

Creating A New Ground for 'Data Science' outside of just Statistics or Computer Science

In a recent AMSTAT News article, Terry Speed wrote a fantastic and inspirational article on the field of statistics.   He began by summarizing some themes from the IYS 2013 "The Future of Statistical Science" workshop he attended.  Some themes he noted were "what an excellent job statisticians are doing" particularly in the areas of "genomics, cancer biology, the study of diet, the environment and climate, in risk and regulation, neuroimaging, confidentiality and privacy and autism research" but saw a lack of representation from "social, agricultural, government, business and industrial statistics".

He openly acknowledged our field of statistics faces many challenges (I direct you to the link for the complete list). In particular: (1) statistics departments are grappling with the debate of adsorbing statistics into incredibly popular emerging field of "data science" and (2) statistics departments are not able to deliver the type of graduates that companies such as Google, Apple, Amazon want which "perhaps involves adopting a more engineering approach to our work".  Yet, students across the world are seeking out majors and/or courses in statistics and "data science" through many different venues e.g. taking courses at a universities, via MOOCs, or online tutorials, etc. Should Statistics be renamed "Data Science"? What is the overlap? What is the best way to train "Data Scientists"?

Terry's response was essentially that data science is not the same as statistics and he saw "no evidence that data science ... has any prospect of replacing our discipline" because statistics "is far wider and deeper than data science".  He encourages statisticians to embrace this emerging field of data science and not fear it: "As with mathematics more generally, we are in this business for the long term. Let's not lose our nerve."

These ideas were all echoed at a symposium I attended last Friday called 'Paths to Precision Medicine: The Role of Statistics' hosted by the Department of Biostatistics at Harvard School of Public Health. At the end of the symposium, there was a panel discussion on 'Education of Future Statisticians in the Big Data Era' with Giovanni Parmigiani, Corsee Sanders, Rafael Irizarry and Marc Pfeffer as the panelists.

Giovanni said when transformations in technology occur, this is a characteristic of the "Big Data Era" and no single field is able to take on "Big Data" as a whole by saying this is a subset of what they do (similar to the argument made by Terry).  Not computer science. Not electrical engineering.  Not statistics. When it comes to creating a curriculum to train individuals who are seeking to become "data scientists" in the Big Data Era, he says we should "teach less and do more".  Rather than starting with a predefined idea of what we should be teaching and/or adding more things to a curriculum for "data science", we should let these individuals dive right into projects as early as possible (with mentorship!). This will also help develop essential skills such as communication and teamwork which cannot be taught in the classroom.

Furthermore, Rafa proposed creating a new, uncharted territory between statistics and computer science to provide the training to individuals who are seeking to become data scientists with as much emphasis as the individuals want in the direction of statistics, computer science or other fields.  This will more than likely require faculty to step out of their comfort zone and even possibly connect with faculty outside their department to jointly teach a course. When it comes to data science, there are many people who believe data science is more about the computing than the statistics and others who would emphasize the statistics more than the computing.  Regardless of the degree of emphasis of various fields on data science, the one thing I do know is that I agree with Jeff Leek's view that "the key word in 'Data Science' is not Data, it is Science".  We should be emphasizing the science (whether it is statistics, computer science, hacking skills, etc) in our curriculums for data science and allow the students to have an opinion in how diverse of a training they want.  As Rafa said, "the faculty in statistics and biostatistics departments are becoming increasingly more diverse and we should make these graduate students more diverse too."

Tuesday, October 2, 2012

Beyond the Genome 2012: Highlights and Discussion Points

This past week I attended the Beyond the Genome conference Sept 27-29 held in Boston at Harvard Medical School this year.  Genome Medicine and Genome Biology have hosted this conference for three years now and it was my first time attending.  To see more details on individual talks, check out the blogposts from Oliver Hofmann tagged with btg2012 or you can search for tweets with the hashtag #BtG2012. I want to give a few highlights from the conference, but first I'll begin with a cartoon which was a very fitting way to describe the conference!


Here are a few key discussion points I took away from the conference:
  1. Genomic or bioinformatic tools used to analyze next-generation sequencing data need to be reproducible, accessible, fast, interactive and web-based. James Taylor from Galaxy advocated for making bioinformatics more reproducible and accessible and even gave an example of in a review only 7 out of 50 papers using BWA provided all the parameters necessary to be able to reproduce their research. Gabor Marth argued bioinformatic tools shouldn't be useable by only informaticians, but also should be useable by biologists.
  2. Bioinformaticians are re-inventing the wheel. I found two recent papers giving a review of batch annotation tools for variants obtained using next-gernation sequencing (Sifrim et al. 2012; Lyon and Wang 2012) with a list of over 19 tools published just since 2010!  At the conference, several of the speakers noted in their talks, "As bioinformaticians we apparently like to keep reinventing the wheel".  I would agree with this statement.  In addition to genomic tools needing to be reproducible, accessible, fast, interactive and web-based, I would argue we need to create a standardized tool or format for annotating variants.  
  3. With mutations, context matters. Why do some of the "right mutations" not respond to treatment? Josh Stuart advocated for using pathway-based analyses to assess the impact of mutations. Should focus on recurrent, rare variants (most likely to be impactful), but also need to make sure the background mutation rate is correct and determine if mutations are in key domains such as DNA binding or conserved domains. Daniel MacArthur said we need to analyze variants in the context of tens of thousands of genomes because the human genomic landscape is dominated by ultra-rare variation and we need to require consistent variant calls across studies. He also gave a set of recommendations on establishing causality from the NHGRI (see picture below).  Lynda Chin argued even if we know the mutation changes the function of the protein, we don't necessarily know the biological consequence. Therefore she argues we need to use a model systems approach to characterize and interpret a complete catalogue of driver mutations with functional validation.  
  4. The idea of using whole exome sequencing as a diagnostic test in clinic has arrived (especially for rare genetic disorders), but we are just now starting to deal with all the challenges that come along with it.  Sharon Plon discussed her first year experience with clinical whole exome sequencing and said almost 25% of sequenced patients have "medically actionable" findings (mostly cardiac, but some cancer). Elizabeth Worthey said we need to be very careful because she explained "medically actionable" doesn't always mean "can be treated". With secondary variants and findings in clinic, Leslie Biesecker argued context matters. In his experience, the response of the patient varies depending on prior family history: with previous diagnosis - mundane response, with prior family history - mild surprise, without family history - dazed and confused.  Amy McGuire gave a beautiful talk from the legal perspective on the reasons to disclose or not disclose incidental findings and included survey results from actual GWAS researchers. In terms of drug discovery, Stuart Schreiber gave many examples of relating the genomic alterations in cancer to the small-molecule sensitivity. 

Finally, I thought the talk by Richard Gibbs deserved it's own paragraph. His talk focused on who should we be sequencing when it comes to human genetic diseases.  Though next-generation sequencing technology has come a long way, it's still not perfect.  Performing whole exome sequencing as a service can be effectively done by experts, but this is inefficient and not scalable.  Automating the process of interpretation often ends with dumb results. Sequencing healthy individuals should be a low priority especially when not tying phenotypes to the control individuals (crazy!).  Sequencing individuals with complex diseases is still tricky because of the debate regarding CDCV v. CDRV hypothesis.  He argued that if complex diseases are caused by rare variants, then we should focus within families and not broader populations.  Sequencing individuals with mendelian disease should be high on the priority list because there is a huge value in obtaining a molecular diagnosis even without a treatment.  Finally recreational sequencing is low on the priority list.  He has even coined the term 'narcciss-ome'!  Overall, he predicts sequencing will become standard of care and soon all the excitement will be passé. 

Final thoughts:  The conference was filled with fantastic talks given by world-renowned speakers.  The level of genetic complexity of diseases never ceases to amaze me, but at the same time seeing such great research being produced to answer some tough genetic questions is always exciting.  I learned a great deal and would highly recommend Beyond the Genome to future participants!  

Thursday, May 17, 2012

Interface 2012: Day 2

JCGS Highlights at the Interface session 
With several great choices to pick from, my Day 2 of the Interface conference began in the JCGS Highlights at the Interface session to listen to Jennifer Le-Radamacher of University of Georgia give a talk on using symbolic-coveriance PCA and visualization techniques for interval-valued data.  The visualizations she showed were very interesting, but it was suggested she check out the colorspace palette in R to further improve her figures.

Information Mining session
Next, I moved to the Information Mining session to hear William Szewcyzk from the NSA give a talk on streaming exploratory data analysis.  He began by summarizing the process of data analysis in five steps: 1) choose a default model which is very controversial subject in itself because he said even if you just describe the data with its mean and variance, you are implicitly assuming the data can be described by the first and second moments of the distribution 2) project your data onto the model 3) examine the fit or lack thereof of the data to the model 4) adjust the model accordingly 5) repeat.  For streaming data, you have to make a slight modification to this process because people incorrectly assume they think they are the only ones working on that flow of data and their process is the last one to touch the data.  Finally he proposed a "default model for streaming data" similar to the way people often assume a gaussian distribution as the default model for static data. The last speaker of the session Andy Frenkiel of IBM gave a thought provoking talk on filling in the gaps of news stories when there is missing information using keyword searches.

Contributed Paper Session II
Xueying Chen of Rutgers University began this session with her split and conquer approach for extremely large datasets.  In her talk, she randomly splits a data set into subsets, estimates a penalized logistic regression model within each subset and finally combined the estimates from the subsets in a final set of coefficient estimates. The second speaker was Garrett Grolemund of Rice University (pictured below) who gave a wonderful demo of his R package Lubridate which greatly simplifies the process of working with dates, times and time zones.  Some great features include the ability to display the same instant of time in different time zones, to save and use time intervals as a class object in R and the test whether certain dates fall "%within%" a different set of dates.  I'll also advertise for his online course for Visualization in R with ggplot2 on June 19-20!




David Kahle of Baylor University (pictured below) gave the final talk of the session on his useful R package mpoly which allows user to work with multivariate polynomials within R.  There are three other packages in R which work with polynomials, but they are not very intuitive or efficient to work with.  Some features include a new class of mpoly objects, basic arithmetic/calculus such as gradients, algebra, and finally evaluating polynomials.




Woman VS Machine: The Inference Battle session
After lunch, quite a few conference attendees move into the Woman VS Machine: The Inference Battle session. The session began with Andreas Buja who gave a thought-provoking talk on the problems with post-selection inference and proposed the Post Selection Inference (PoSi) constant which allows valid post-selection inference. Interestingly, PoSi guarantees coverage of CIs and Type I errors of tests and is not specific for any type of model selection. The second speaker in the session Heike Hofmann outlined the concepts of visual inference within the framework of exploratory data analysis. In a classical statistical setting, we reject the null hypothesis if the test statistic is past some threshold, but in a visual setting, she argued we would reject the null hypothesis (i.e a plot is not distinguishable from null plots) if the data plot is identifiable.  A great example was shown in which she simulated random data in four dimensions and included one plot with the real signal. Due to the artifact of high dimensionality, the audience was not able to pick it out (including me!).  Finally, she described a set of experiments they designed using Amazon Mechanic Turk in which they recruited people to look at a line-up of plots to pick out which ones were different from the rest using criteria such as bi-modality, outliers and mean shift (and showed how the power estimates). Very curious results. The last speaker Mahbub Majumder described these "Turk experiments" in greater detail. The session ended with a great question from the audience about the reproducibility of this type of research (which I believe should be as long as the same plots could be used again).

Banquet Keynote
Mark Hansen of UCLA gave an amazing keynote talk for the banquet tonight!  I was very impressed with the quality of graphics and art projects he has produced over the last decade.

I will wrap up with a few pictures from the banquet of some current and past PhD students at Rice.





Wednesday, May 16, 2012

Interface 2012: Day 1

Today was the first day of the 43rd Interface Conference 2012 which is being held at Rice University this year (follow me with updates on twitter with hashtag #Interface12) There were several concurrent technical sessions going on through out the day, so I only post about the ones I attended.  It was an early start to the morning, but the coffee definitely helped.  :)



Keynote speaker
The keynote speaker Trevor Hastie from Stanford gave a wonderful talk this morning on methods for low-rank factorization with missing data (perfect application for the Netflix data).  He specifically discussed the methods Soft-Impute (soft threshold SVD) and Hard-Impute and showed their relationship to the Maximum Margin Matrix Factorization (MMMF).  His method has an expectation-maximization flavor to it and is similar to alternating ridge regression.  Finally he ended with a few generalizations including Convex Robust Completion (Robust SVD).  When the data matrix X can be approximated by L (a low-rank matrix) + S (sparse matrix), then the method just adds a penalty parameter on a sparse matrix in addition to penalty parameter on the low rank matrix.  I enjoyed the level of detail in this talk.



Software Development in R session
After the keynote, I decided to attend the Software Development in R technical session.  The first speaker JJ Allaire (founder of Rstudio) gave a great high level demo of many useful features in Rstudio.  He stressed the importance of "reproducible research" and "trustworthy computing".  Some of the most exciting things in Rstudio include: a searchable history for any piece of code ever run through the console, page back through plots, quickly traverse through nested functions, interact with Git and SVN and incorporation of Sweave and knitr.  You can easily navigate between chunks of code and even be pointed back to the source code after clicking on a complied pdf.  Rstudio now has the feature of writing in the markdown language to quickly publish high quality web pages instead of having to deal with html.  The second speaker in the session Norm Matloff of UC Davis discussed parallel computing in R. He reviewed classical shared-memory loop scheduling methods (static, dynamic, time-varying chunk size, etc) and how to these might be adapted to R. The example he discussed was how to parallelize all possible regressions in a given data set with dim(X) = n x p.  The available R packages discussed for parallel computing were:
  1) snow - serializes/deserializes communications which takes time; most used R package; the functions clusterApply() is static and clusterApplyLB() is dynamic; both limited to a fixed chunk size of 1 (small chunk sizes not good because of high overhead); chunk size > 1 must be programmed by user
  2) Rmpi - more flexible than snow, but still has serialization and network problems
  3) mclappy/multicore - each call involves new unix process creation
  4) gputools - each call involves a GPU kernel invocation, time intensive; major overhead
He suggested a new scheduling method called 'Random Scheduling'.  After making this small adjustment, then you can use the R packages as before (e.g. snow).  For his presentation slides go here and his open source book go here.  The final speaker of the session Duncan Murdoch gave a great overview of the older tools available for debugging and some examples of visual debuggers.


Statistical Models for Complex Functional Data session
The session started out with Todd Ogden of Columbia University who discussed sparse functional principal component regression to predict depression using MRI images as the functional data.  He mentioned other statistical learning tools such as random forests may be more accurate, but he advocated for using a regression-based method with functional data because functional regression has a clear interpretation of the weight function.  After expressing the functional components in the wavelet domain he applies penalization techniques such as wavelet-based LASSO to the functional model.    The basic idea is to perform sparse functional principal component analysis and then use the loadings as the predictors.  The second speaker Lan Zhou of Texas A&M showed how to use penalized bivariate B-splines in functional data analysis to estimate variability in Texas temperatures over the past 100 years.  Because Texas or the "domain" is complicated (e.g. not rectangular, contains holes, etc), she uses the idea of triangulations.  The goal is to estimate a bivariate smooth function over the domain to create a temperature map using data from weather stations and to investigate the variability in temperatures over the years.  Veera Baladandayuthapani of UT MD Anderson wrapped the session up by discussing a bayesian functional mixed model for copy-number variation data measured by aCGH array data and extending it for SNP array data (higher resolution).  The goal was to do a joint analysis on set of samples to look for a small signal by borrowing strength between the samples.

So far the talks have been excellent and I'm looking forward to the rest of the sessions tomorrow and Friday!