Thursday, May 13, 2010

The Shape of Shape Analysis Research: Part I

Shape analysis is a topic that is almost a killer app for computational geometry. Where the word 'almost' comes in is an interesting story about the tension between computational power and mathematical elegance.

Shape analysis is defined by the generic problem:
Given two (or more) shapes, determine whether they are similar.
Simple, no ? I don't need to list the applications for this problem. Or maybe I should: computer vision, computational biology, graphics, imaging, statistics, computer-aided design, and so many more.

It would not be an exaggeration to say that in many ways, shape analysis is as fundamental a problem as clustering. It's a problem that people have been studying for a very long time (I heard a rumor that Riemann once speculated on the manifold structure of shapes). Especially in the realm of biology, shape analysis is not merely a way to organize biological structures like proteins: it's a critical part of thinking about their very function.

Like any problem of such richness and depth, shape analysis has spawned its own ecology of concepts. You have the distances, the transformation groups, the problem frameworks, the feature representations, the algorithms, and the databases. You now even have the large data sets of very large shapes, and this brings nontrivial computational issues to the forefront.

Shape analysis (or shape matching) has been a core part of the computational geometry problem base for a long time. I've seen papers on point pattern matching from the late 70s. There's been a steady stream of work all through the past 30 years, introducing new concepts, distances and algorithm design principles.

In parallel, and mostly within the computer vision community, there have been other efforts, focused mainly on designing more and more elegant distance measures. Computer graphics got in on the action more recently as well, with a focus on different kinds of  measures and problems.

What I think is unfortunate (and this is entirely my own opinion) is that there's a strong disconnect between the developments happening in the computational geometry community, and the parallel efforts in the more 'applied' communities.

I'm not merely talking about lack of awareness of results and ideas. I believe there are fundamentally different ways in which people go about attacking the problems of shape analysis, and I think all sides could benefit greatly from understanding the strengths of the others.

Specifically, I think that within our community, we've focused for far too long on measures that are easy to define and relatively easy to compute (while not being trivial to work with). On the more 'math/vision' side, researchers have focused much more on measures that have the 'right' kind of mathematical structures, but have bailed miserably on computational aspects.

Over the next post or two, I want to develop this argument further, and hopefully end with a set of problems that I think are worthy of study both from their intrinsic appeal and the unsolved computational issues they present.

Stay tuned....

p.s The clustering series will also continue shortly. Oh do I love summer :)

Wednesday, May 12, 2010

SoCG 2010: Come one, come all :)

There's about a month left to go for SoCG, and I just returned from a walk through at Snowbird. Surprisingly, it was still snowing - the ski season has wound down though. Most of the snow will disappear by early June - we're seeing the last confused weather oscillations right now before the steady increase in temperature.

We went up there to check out the layout of the room(s) for the conference - the main conference room is nice and large, and the parallel session room is pretty big as well, there'll be wireless access throughout (with a code), and there's a coffee place right next to the rooms. There are nice balconies all around, and of course you can wander around outside as well.

Registrations have been trickling in, a little slower than my (increasingly) gray hair would like. If you haven't yet registered, consider this a gentle reminder :). It helps to have accurate numbers when estimating food quantities and number of proceedings etc.

See you all in a month !

Monday, May 10, 2010

Hirsch Conjecture disproved

Via polybot, Gil Kalai posts that Francisco Santos has disproved the Hirsch Conjecture. This is big news.

The conjecture:
the edge-vertex graph of an n-facet polytope in d-dimensional Euclidean space has diameter no more than n − d.
The result:
I will describe the construction of a 43-dimensional polytope with 86 facets and diameter bigger than 43. The proof is based on a generalization of the d-step theorem of Klee and Walkup.

The proof will be presented at the 100 Years in Seattle (The mathematics of Klee and Grunbaum) conference in end-July, this year. 

Saturday, May 08, 2010

Future-proofing research

Every now and then, we get called upon to project our work into the future. Usually, it's in a grant proposal (especially in a CAREER proposal). Sometimes it might be part of strategic planning at a faculty retreat. It even shows up in solicitations for position papers at various venues (for example this recent one that was circulating on a faculty list). It definitely comes up at faculty interviews, although I usually view it as a hazing ritual or the equivalent of "Nice weather we're having, aren't we?"

I understand the short-term imperative for such things: it's good to know that there are timelines in which your work has some kind of measurable impact, and even better to know that there's more than one (BPP vs NP, anyone?).

But I get the sense (and maybe I'm just off base here) that this kind of future prediction business is more common in non-theoryCS areas. My archetypical story for what happens if you ask theoreticians about future directions is Jeff Erickson's hilarious tale about his interview at MIT.

Of course the most famous example of future projection is in mathematics ! So maybe my premise is doomed ? But somehow I don't think so. I don't think mathematicians since Hilbert go around proposing future directions for entire areas (although there might be general consensus on key open problems), and I think theoryCS has absorbed much of this ethos (although I don't think that's true in theoretical physics).

I ask because I always feel awkward when asked questions like "where is going in the next X years ?" or even worse, "where SHOULD be going in the next X years". Maybe the more reasonable question is "where's all the activity and ferment happening right now". 

Tuesday, May 04, 2010

Ranking departments topologically rather than totally.

The US news rankings came out a while back (Jon Katz had two posts on this). As usual, this will prompt a round of either back-slapping or back-stabbing, depending on whether your department ranking went up or down (ours didn't change at all, which could also be a bad thing).

What I'd like to propose is a completely different way of doing rankings.

It's generally accepted that the place where rankings make the most difference is in graduate admissions, and there's a secondary effect in faculty hiring (since faculty want to get good students to work with). The general belief is that students will tiebreak between universities based on ranking, in the absence of more contextual information.

But it's also insanely silly to obsess about the relative rankings of (say) the top 5 schools, or to exult in your movement from 53 to 47 in the rankings. What I believe is generally true is that there are rough strata (antichains in a partial order, if you will) in which departments are generally of equivalent rank. Spending time and energy trying to optimize within such a statum is a useless waste (which doesn't mean that people don't LOVE to do it, because any activity is positive activity, right ? ... right ? ....)

What we do keep track of, and is interesting, is which universities our admits reject us for, or accept us in place of. If I'm not Stanford or MIT, but students are rejecting me only to go there, then I'm not happy, but I feel minor relief that at least they're not rejecting me for the University of Obscurity in Scarceville, Podunkistan.

But of course we know what this is ! it's a topological order ! So I propose the following tiering scheme:
A department is at tier k if "all" departments it is rejected for are at tier k-1 or less. 
Note 1: We have to define "all" carefully - there's always someone who's (say) following a boyfriend or girlfriend, or really wants to live in some town, etc etc. My preferred definition of "all" would be "at least 80%" or some large figure like that.

Note 2: If in fact people did select universities based on the "current" ranking scheme, this order would reflect that. Of course I don't believe this will happen

Note 3: This might even allow for more fine grained analysis based on subject area. Depending on the areas of the admitted students, one could create stratified orders by area.

Note 4: No I have no clue how to get this data, but many departments informally maintain this information (I know we try to get this info when we can), and it's not like the current approach is dripping with rigor anyway.

Note 5: If you're an administrator, you'll hate this when you're trying to move to a higher level, and you'll love it when you actuall make the move. The problem with the lack of granularity might annoy some people though.

Metrics on distributions defined over metric spaces

Now that's a title to make your head spin !

Semester is over, which means I hope to get back into blogging form again, picking up my clustering series where it left off, and also starting some new rants on problems in shape matching (which more and more to me look like problems in clustering).

But for today, just a little something to ponder. The following situation often occurs in data analysis. You have some data that inhabits a metric space - usually Euclidean space, but that doesn't really matter. You also have distributions over the data, by which I mean some kind of weight vector with one "component" for each data point, and components summing to one. The goal now is to compare these weight vectors in a way that takes into account the structure of the space.

The standard construction that one uses here is the earthmover distance, also known as the Wasserstein distance, or the Monge-Kantorovich distance, or the Mallows distance, or the transportation metric (you pick your favorite one). It's very intuitive (which is probably why it's been invented so many times) and works like this. Imagine piles of earth at each data point, with each pile having mass equalling the weight at that point. We have a "starting configuration" consisting of the first weight vector, and an "ending configuration" consisting of the second weight vector. The goal is to figure out how to "transport" the earth with minimum effort (effort = weight X distance moved) so that the first configuration becomes the second. Formally, this amounts to a generalized matching problem that can be solved via the Hungarian algorithm. The earthmover distance is very popular in computer vision (see the Wikipedia article for details)

Another metric over distributions of the kind above is the Levy-Prokhorov metric, which for two distributions u and v is defined as the smallest e such that on any neighborhood A, the measure of u is within e of the measure of v on A inflated by e (i.e by constructing a ball of size e around A), and vice versa. I haven't seen this metric used much in practice, and it seems hard to compute.

Another approach that I realized recently uses a method that thus far has been used only in shape analysis. I've been working with a shape matching measure called the current distance of late (more on this in another post - it's a fascinating measure). Roughly speaking, it works like this. It starts with a "shape" defined anyway you like (clouds of points, a curve, a surface, whatever), and a similarity function (a positive definite kernel actually) defined on this space. It then allows you to compare these shapes by using the kernel to create a "global signature" that lifts each shape to a point in a Hilbert space, where the induced distance captures the shape distance.

It also works with weighted point sets, which is the relevant point here. Suppose I give you a space with distance defined indirectly via a kernel similarity function (rather than via a distance function). The current distance then gives me a way of comparing distributions over this shape just like the above measures, and the kicker is that this approach is WAY more efficient then any of the above methods, taking near-linear time instead of needing the rather expensive Hungarian algorithm. Moreover, the current distance has a built-in isometric embedding into a Hilbert space, something the earthmover distance cannot have.

If you're curious for more details, wait for my post on the current distance - in the mean time, there are two papers we've written (one online, the other you should email me for) that explore the theory and practice behind the current distance. I'm curious now as to whether the current distance can be used an efficient replacement for the earthmover distance in applications that rely on the EMD, but don't have a natural relationship to shape analysis.

p.s Shape matching in particular, and data analysis in general, is a rich source of ever more exotic metric spaces. I've been working with a number of these, especially in non-Euclidean spaces, and there are lots of interesting algorithms questions here.

Saturday, May 01, 2010

SoCG 2010 Early Registration Deadine

SoCG 2010 Early Registration deadline coming up tomorrow night. Make sure you register so you don't pay the full price. Also make sure to lock down your hotel reservations.

Friday, April 02, 2010

SoCG 2010 News

There's been a flurry of activity on the SoCG 2010 front, and while I've been running around getting things done, I haven't been diligent enough about announcing the activity here. 
  1. The registration site is up and running, as has been noted. The deadline for early registration is May 2, so hurry up and get those registrations in ! The earlier you register, the easier it is for us to do our planning for conference activities.
  2. You can (and should) reserve your hotel rooms at the same time. We have a group block rate at both the Cliff Lodge and the Lodge, (students can share rooms at the latter) and the block will expire in the middle of May. Snowbird is up in the mountains, and so if you plan on reserving a hotel room elsewhere, you'l need to arrange to get yourself up the mountain - not hard, but something to plan for. 
  3. MADALGO is running a second installment of the very successful MASSIVE workshop on massive data algorithmics. This will be held on June 17, the day after SoCG ends, and so make sure to reserve an extended hotel stay if you plan on staying over. The group block applies through MASSIVE. Moreover, the deadline for submissions is Apr 14, so get those papers ready !
  4. We've arranged a group discount rate for shuttles to and from the airport to Snowbird. If you click on the Travel section of the conference website, the last link will take you to a PDF form that you can use to reserve the discounted shuttle ($29 each way). Make sure to mention the group name (SoCG 2010) and group number (6922) if you call in a reservation.
  5. The final program is available on the website. Some unusual elements worth noting: 
    • There's a long break for a hike on Tuesday. The plan (and hope) is to make our way up the mountain for some spectacular views of the valley. 
    • The conference ends relatively early on Wednesday (4pm) if you want to make a quick dash home. But of course you won't, because you'll be staying for MASSIVE !
  6. If you need a letter of invitation for visa processing, send me an email with your full name and address, and title(s) of the papers/videos you'll be presenting. I can generate the letters pretty quickly, but I do need this information. 

Tuesday, March 30, 2010

Why Conference Review Must End !

Exhibit A: Matt Welch's post on how PC deliberations happen.

Notes:
  • Yes, I realize it's partly tongue-in-cheek. But it's not far from the truth !
  • No, going to all-electronic meetings doesn't solve the problem. It merely replaces one set of group dynamics by another
  • Yes, we can't hope to remove irrational biases in the review process. That's why all we can hope for is to force them to be exposed and questioned. A back-and-forth between author and reviewer can help do that. 
  • And no, it's not true that "theory PCs are much better". 
I've been to NSF panels where the program manager does an excellent job of forcing people to state precisely what they mean by "interesting", "infeasible", "novel contribution" and other such weasel-words. When that happens, it's a bit easier to assess the contribution. One could imagine enlightened PC chairs doing this at paper review time, but there's really no time, given the number of papers that need to be processed in 2 days or so. 

Saturday, March 13, 2010

Choosing the number of clusters II: Diminishing Returns and the ROC curve

(this is part of an occasional series of essays on clustering: for all posts in this topic, click here)

In the last post, we looked at the elbow method for determining the "right" number of clusters in data. Today, we'll look at generalizations of the elbow method, all still based on the idea of examining the quality-compression tradeoff curve.

For the purpose of discussion, let's imagine that we have a plot in which the quality of the clustering (measured anyway you please, as long as 0 means all items in one cluster, and 1 means all items in separate clusters) is measured along the y-axis, and the representation complexity (i.e the space needed) is measured on the x axis, where again 0 corresponds to all items in a single cluster (least representation cost), and 1 corresponds to all items in separate clusters.

The ideal cluster is located at (0,1): cheapest representation and perfect quality. In general, as we use more and more clusters to organize the data, we can trace out a curve that starts at (0,0) and ends at (1,1). For most sane functions, this cuve is concave, and lives above the diagonal x=y.

This curve contains lots of useful information about the behavior of the data as the clustering evolves. It's often called the ROC curve, in reference to the curve used to capture the tradeoff between false positive and false negatives in classification. But what can we glean from it ?

We can ask what the curve would look like for "unclusterable" data, or data that has no definite sense of the "right number" of clusters. Such data would look self-similar: you could keep zooming in and not find any clear groupings that stood out. It's not too hard to see that data that looked like this would have a ROC curve that hugs the diagonal x=y, because there's a relatively smooth tradeoff between the quality and compressibility (so no elbow).

Conversely, data that does appear to have a definite set of clusters would try to get closer to the (0,1) point before veering off towards (1,1). This suggests a number of criteria, each trying to quantify this deviation from unclusterability.

  • you could measure the area between the curve and the diagonal. This is (essentially) the AUC (area-under-curve) measure. 
  • You could measure the closest distance between (0,1) and the curve. This also gives you a specific point on the curve, which you could then associate with the "best" clustering. 
  • You could also find the point of diminishing returns - the point of slope 1, where the clustering starts costing more to write down, but yields less quality. I've used this as the start point of a more involved procedure for finding a good clustering - more on that later. 
The ROC curve gives a slightly more general perspective on the elbow method. It's still fairly heuristicy, but at least you're trying to quantify the transition from bad clustering to good in somewhat more general terms. 

Ultimately, what both the ROC curve and the elbow method itself are trying to find is some kind of crucial transition (a phase transition almost) where the data "collapses" into its right state. If this sounds like simulated annealing and bifurcation theory to you, you're on the right track. In the next part, I'll talk about annealing and phase-transition-based ideas for finding the right clustering - all fairly ingenious in their own right. 

(ed. note: comments on the last post have convinced me to attempt a post on the nonparametric approaches to clustering, which all appear to involve eating ethnic food. Stay tuned...)

Monday, March 08, 2010

Who pays for submissions ?

Writing a paper takes a tremendous amount of time. So, one of the frequent complaints that authors  make is when PC members submit half-baked, clearly below-threshold reviews on a paper just to get the resume bullet and claim to have done their reviewing duties. Personally, I feel intense anger when receiving  crappy reviews that come with not the slightest bit of insight, and then am expected to rebut them or accept them. Not to mention the long-term psychological damage incurred by having papers rejected one after another. 

The problem is that reviewing a paper for a conference is free: all it takes is a few clicks of the mouse to upload your PDF file. (Of course, I'm not accounting for the cost of doing the research  (ha!) and actually reviewing the paper.)

Let's estimate the costs associated with doing research and submitting papers to conferences. I spend many months working, writing and submitting papers to conferences. A highly competitive conference will assign three reviewers to my paper, and with a lot of luck one of them might even tangentially be aware of my research area. After I make up a bunch of numbers, the cost of rejection of my paper amounts to over 3 gazillion dollars, none of which I can recoup. It's clear that conferences, which only survive if people submit, should be paying me to submit !

Of course, imposing this kind of a fee would no doubt drastically reduce the number of papers that are submitted. But this seems like a good thing: it would probably reduce the number of conferences, and remove the fiction that conferences actually do "quality control", leaving them with their original purpose of networking and creating a sense of community. Conferences could generate revenue by charging reviewers for the opportunity to preview the new works being submitted: this  would potentially also improve the quality of the reviews as well.  Although the financial incentive is not that great, getting paid should encourage TPC members to take the process more seriously.

The only downside I can see is people who submit a ton of papers everywhere and become "professional paper writers", but TPC chairs would clearly have to balance this against the research credentials of the people submitting papers. Note that many journals impose author fees for publication of the paper, so this provides a nice offset against that cost. 

It just seems crazy to me that the research community provides this free paper previewing service for committees with no negative ramifications for writing totally bogus reviews.

Disclaimers for the sarcasm-challenged:
  • Yes, I am obviously aware of Matt Welsh's post on this topic
  • Yes, this is a total ripoff/parody of his post
  • Yes in fact, I disagree with his point.

Sunday, March 07, 2010

Choosing the number of clusters I: The Elbow Method

(this is part of an occasional series of essays on clustering: for all posts in this topic, click here)

It's time to take a brief break from the different clustering paradigms, and ponder probably THE most vexing question in all of clustering.
How do we choose k, the number of clusters ?
This topic is so annoying, that I'm going to devote more than one post to it. While choosing k has been the excuse for some of the most violent hacks in clustering, there are at least a few principled directions, and there's a lot of room for further development.

(ed. note. The Wikipedia page on this topic was written by John Meier as part of his assignment in my clustering seminar. I think he did a great job writing the page, and it was a good example of trying to contribute to the larger Wikipedia effort via classroom work)

With the exception of correlation clustering, all clustering formulations have an underconstrained optimization structure where the goal is to trade off quality for compactness of representation. Since it's always possible to go to one extreme, you always need a kind of "`regularizer"' to make a particular point on the tradeoff curve the most desirable one. The choice of $k$, the number of clusters, is one such regularizer - it fixes the complexity of the representation, and then asks you to optimize for quality.

Now one has to be careful to see whether 'choosing k' even makes sense. Case in point: mixture-model clustering. Rather than asking for a grouping of data, it asks for a classification. The distinction is this: in a classification, you usually assume that you know what your classes are ! Either they are positive and negative examples, or one of a set of groups describing intrinsic structures in the data, and so on. So it generally makes less sense to want to "`choose"' $k$ - $k$ usually arises from the nature of the domain and data.


But in general clustering, the choice of $k$ is often in the eyes of the beholder. After all, if you have three groups of objects, each of which can be further divided into three groups, is $k$ 3 or 9 ? Your answer usually depends on implicit assumptions about what it means for a clustering to be "`reasonable"' and I'll try to bring out these assumptions while reviewing different ways of determining $k$.

The Elbow Method

The oldest method for determining the true number of clusters in a data set is inelegantly called the elbow method. It's pure simplicity, and for that reason alone has probably been reinvented many times over (ed. note: This is a problem peculiar to clustering; since there are many intuitively plausible ways to cluster data, it's easy to reinvent techniques, and in fact one might argue that there are very few techniques in clustering that are complex enough to be 'owned' by any inventor). The idea is this:

Start with $k=1$, and keep increasing it, measuring the cost of the optimal quality solution. If at some point the cost of the solution drops dramatically, that's the true $k$.

The intuitive argument behind the elbow method is this: you're trying to shoehorn $k$ boxes of data into many fewer groups, so by the pigeonhole principle, at least one group will contain data from two different boxes, and the cost of this group will skyrocket. When you finally find the right number of groups, every box fits perfectly, and the cost drops.

Deceptively simple, no ? It has that aspect that I've mentioned earlier - it defines the desired outcome as a transition, rather than a state. In practice of course, "`optimal quality"' becomes "`whichever clustering algorithm you like to run"', and "`drops dramatically"' becomes one of those gigantic hacks that make Principle and Rigor run away crying and hide under their bed.


The Alphabet-Soup Criteria

So can we make the elbow method a little more rigorous ? There have been a few attempts that work by changing the quantity that we look for the elbow in. A series of "`information criteria"'  (AIC, BIC, DIC, and so on) attempt to measure some kind of shift in information that happens as we increase $k$, rather than merely looking at the cost of the solution.

While they are all slightly different, they basically work the same way. They create a generative model with some kind of term that measures the complexity of the model, and another term that captures the likelihood of the given clustering. Combining these in a single measure yields a function that can be optimized as $k$ changes. This is not unlike the 'facility-location' formulation of the clustering problem, where each "`facility"' (or cluster) must be paid for with an 'opening cost'. The main advantage of the information criteria though is that the quantification is based on a principled cost function (log likelihood) and the two terms quantifying the complexity and the quality of the clustering have the same units.

Coming up next: Diminishing returns, the ROC curve, and phase transitions.

Tuesday, February 23, 2010

Guest Post: Update from the CRA Career Mentoring Workshop, Day II

(ed note: Jeff Phillips is at the CRA Career Mentoring Workshop. His Day 1 dispatch is here)

It is day two at the CRA Career Mentoring Workshop.

Today was all about funding, with speakers from NIH (Terry Yoo), DARPA (Peter Lee), Laboratory for Telecommunications Science (Mark Segal), and NSF (Jan Cuny). Jeanette Wing, the assistant director at NSF CISE also made an appearance at reception yesterday.

(ed. note: I just heard that Jeannette Wing is leaving CISE in July. This is sad news - she was a strong and dynamic presence at CISE)

NIH advertised having a lot of money (about $30 billion, compared to $7 billion in NSF). The NIH has many sub-institutes with many different topics, but all applications are funneled through grants.gov. Terry Yoo was very enthusiastic about us applying for a piece of his large pie. It seems a bit tricky, however, to fit a pure computer science project into one of these institutes, specific health-related applications are enough.

We all (CI Fellows) thanked Peter Lee who helped spearhead the CI Fellows program. He recently joined DARPA to head the Transformational Convergence Technology Office (TCTO or "tic-toe"), a new program that will oversee many funded computer science programs. See the DARPA_News twitter feed for information on DARPA solicitations. Among other goals of this office, is to eliminate harsh "go or no go" conditions associated with DARPA grants.
For young researchers, look for CSG or YSA programs, similar in some ways to NSF CAREER awards.

The Laboratory of Telecommunications Science is part of NSA. They hire many many Ph.D.s for advanced computer science research. He could not tell us specifics about what they do, but compared it to an industrial research labs (e.g. AT&T, Yahoo Research, etc.). Movement between research parts and non-research parts is more fluid and is pretty hands on. Even the theoretical computer scientists and mathematicians they hire often build systems to implement their work.

To get funding through them, they generally have close and specific collaborations with faculty. The best way to start a relationship is sending a student on a summer internship (maybe even an undergrad) or via a sabbatical.

Jan Cuny from NSF decided she did not need to convince us that we should apply to NSF; rather, she just assumed we would and gave a howto on applying for NSF grants. Most important tip: **talk to program officers!** (before you submit). Otherwise, it is hard to give specific summaries from her talk (the slides will eventually be online--definitely look for them). The presentation nicely demystified some of the reviewing process; such as how grants are reviewed and why she may choose a certain proposal (for diversity) that scored slightly lower than another unfunded proposal. The other key advice: follow the guidelines precisely and carefully, make it easy for reviewers, and focus the content section on proposed work, not existing work.


A parting thought. It has been great to see many friends who are recent faculty or postdocs, in areas who I might not meet in my normal conferences. But it was a bit odd to have such a large fraction of my competition for jobs in the next year or two in the same room. The funding agencies were definitely here advertising how to get their funding, but if we did not realize that this was important, we would probably not have much luck getting jobs. If you were a department looking to hire to someone, perhaps it would have made sense to come here to recruit postdocs to apply ? Although I guess that is a bit optimistic, as it is a hirer's market.

Is there some way that having many people looking for jobs all in one place can facilitate the hiring process, or has this been out-dated with the electronic age? I would argue that personal interaction is underrated, and would help universities figure out not just who has a great resume on paper, but is also great to personally interact with.

Monday, February 22, 2010

Guest Post: Update from the CRA Career Mentoring Workshop

(ed note: Jeff Phillips is at the CRA Career Mentoring Workshop today and tomorrow, and filed this dispatch from Day 1.)


I am reporting from the CRA Career Mentoring Workshop. So far it has been excellent.

Frankly, I would not have even thought of coming (I did not even know about it), but it was "highly recommended" for all CI Fellows. But, having been here the first day, I would now recommend it to postdocs, early faculty, and senior graduate students set on a career in academics. Someone could obtain all of the information presented here by asking the right people around your department or field, but this workshop has really stressed what are the right questions to ask, and whom to ask.

The topics today were "Planning Your Research Career," "Career Networking," "Teaching," "Managing and Mentoring Students," "Preparing a Tenure Dossier," "Time Management and Family Life," and "Advice from Early Career Faculty." I thought an important aspect of how it was organized was that each topic had at least two speakers. This kept presentations short, and always provided at least two (often differing) perspectives. This ensured that there was just not someone lecturing us on their opinions on a topic, but instead demonstrating to us that there was no one right way to approach being a young faculty. The slides from all of the talks should eventually be online, found from one of the above links.

It's hard to pick out a handful of pieces of advice to share with you. Maybe I had heard 80% of the suggestions before, and 20% were new. But I would guess for an average person in my position a different 20% would be new. For instance, I was surprised by how different the tenure process can be from university to university. The solution: ask your department chair what are the key factors (journal vs. conference papers, is there a funding dollar threshold, student progress, who can write letters), and ask more senior faculty members in your department who recently got tenure for copies of their dossier. In general, the advice was to "ask for advice" and sometimes, according to Kim Hazelwood, "if you don't like the answer, then keep asking until you get an answer you like."

(ed. note: I've also heard the counter-advice "focus on doing the work to get tenure at a high ranking place, rather than just your department" - the rationale being that having a generically strong tenure case makes you more mobile if you need to be)

Also, your department is making a 5-6 year investment in you. So they should be there to help you succeed; if you don't get tenure, then everyone loses. This has all helped realize the great demands of the tenure process, but also make it seem quite possible. Intimidating and comforting at the same time.

Wednesday, February 17, 2010

SoCG author feedback

David Eppstein has an analysis of the SoCG author feedback with respect to his papers. Worth perusing: his overall conclusion is that having the rebuttal was a good idea, but he'd like to hear from the committee (perhaps at the business meeting?).

I had two papers rejected. For one there was no feedback requested, and the paper was rejected. The final reviews made it clear that the reviewers understood the main contributions of the paper - what was under contention (among other things) was how the material was presented, and that's obviously not something author feedback can help with.

The other paper had one request for feedback which was basically a long negative review, again focusing on the presentation. We tried to respond as best we could, but it didn't make too much of a difference.

He did concur that the quality of reviewing was very high.

SoCG Feedback...

It's my 1000th post !! Took longer than I expected, but then my average has been slipping. Here's hoping I get to 2000 !

Just got back my reviews from the Valentine's day massacre. I have to say that I'm stunned, and not for the reason you think.

I've been writing papers and getting reviews for a LONG time now, and I have to say that these are the best reviews I've ever seen, and are way beyond the standard for the typical theory conference. Both papers were rejected, and so the reviews necessarily were negative, but in an extremely constructive way. They critiqued without being critical, were detailed and thorough, and clearly got to the heart of what the paper was about. The comments pointed out what was good in the paper, and more importantly, pointed out what the reviewers felt was missing, and how best to address the problems. I actually felt better about the rejection after reading the reviews, because they came across as genuinely liking the work.

Now it wasn't all good. There was a basic premise at the heart of the rejection that I disagree with, but it's a reasonable point to disagree on, and I can at least see a way to resolving that problem.

At least one other author agrees with me on this assessment - you're welcome to share your experiences in the comments. Congratulations to the PC - theory conference reviews are often slammed, and rightly so, but these reviews stand out for their high quality.


Tuesday, February 16, 2010

SoCG accepted papers

After the Valentine's day massacre, comes the list. I'll link to PDFs if I'm pointed to them (add in the comments)
  • David Millman and Jack Snoeyink. Computing Planar Voronoi Diagrams in Double Precision: An Example of Degree-driven Algorithm Analysis
  • György Elekes and Micha Sharir. Incidences in Three Dimensions and Distinct Distances in the Plane
  • Micha Sharir, Adam Sheffer and Emo Welzl. On Degrees in Random Triangulations
  • Pankaj K. Agarwal, Rinat Ben Avraham and Micha Sharir. The 2-Center Problem in Three Dimensions
  • Florian Berger and Rolf Klein. A Traveller's Problem
  • Sergio Cabello and Bojan Mohar. Adding one edge to planar graphs makes crossing number hard
  • Tobias Christ, Dömötör Pálvölgyi and MiloÅ¡ Stojaković. Consistent digital line segments
  • Dominique Attali and Andre Lieutier. Optimal reconstruction might be hard
  • Dominique Attali and Andre Lieutier. Reconstructing shapes with guarantees by unions of convex sets
  • Mark de Berg. Better Bounds on the Union Complexity of Locally Fat Objects
  • Marc Glisse and Sylvain Lazard. On the complexity of the sets of free lines and free line segments among balls in three dimensions
  • Tamal Dey, Jian Sun and Yusu Wang. Approximating loops in a shortest homology basis from point data
  • Mohammad Ali Abam and Sariel Har-Peled. New Constructions of SSPDs and their Applications
  • Sergio Cabello, Éric Colin de Verdière and Francis Lazarus. Output-sensitive algorithm for the edge-width of an embedded graph
  • Sergio Cabello, Éric Colin de Verdière and Francis Lazarus. Finding Shortest Non-Trivial Cycles in Directed Graphs on Surfaces
  • David Eppstein and Elena Mumford. Steinitz Theorems for Orthogonal Polyhedra
  • Akitoshi Kawamura, Jiri Matousek and Takeshi Tokuyama. Zone diagrams in Euclidean spaces and other normed spaces
  • Eryk Kopczynski, Igor Pak and Piotr Przytycki. Acute triangulations of polyhedra and the space
  • Keiko Imai, Akitoshi Kawamura, Jiri Matousek, Daniel Reem and Takeshi Tokuyama. Distance k-sectors exist
  • Joseph Mitchell. A Constant-Factor Approximation Algorithm for TSP with Neighborhoods in the Plane
  • Benjamin A. Burton. The complexity of the normal surface solution space
  • Karl Bringmann. Klee's Measure Problem on Fat Boxes in Time $O(n^{(d+2)/3})$
  • Natan Rubin. Lines Avoiding Balls in Three Dimensions Revisited
  • Pankaj K. Agarwal, Jie Gao, Leonidas Guibas, Haim Kaplan, Vladlen Koltun, Natan Rubin and Micha Sharir. Kinetic Stable Delaunay Graphs
  • Haim Kaplan, Micha Sharir and Natan Rubin. A Kinetic Triangulation Scheme For Moving Points in The Plane
  • Anne Driemel, Sariel Har-Peled and Carola Wenk. Approximating the \Frechet Distance for Realistic Curves in Near Linear Time
  • Joachim Gudmundsson and Pat Morin. Planar Visibility: Testing and Counting
  • Ken-ichi Kawarabayashi, Stephan Kreutzer and Bojan Mohar. Linkless and flat embeddings in 3-space and the Unknot problem
  • Roel Apfelbaum, Itay Ben-Dan, Stefan Felsner, Tillmann Miltzow, Rom Pinchasi, Torsten Ueckerdt and Ran Ziv. Points with Large Quadrant-Depth
  • Sunil Arya, David Mount and Jian Xia. Tight Lower Bounds for Halfspace Range Searching
  • Don Sheehy, Benoit Hudson, Gary Miller and Steve Oudot. Topological Inference via Meshing
  • Gur Harary and Ayellet Tal. 3D Euler Spirals for 3D Curve Completion
  • Peyman Afshani, Lars Arge and Kasper Dalgaard Larsen. Orthogonal Range Reporting: Query lower bounds, optimal structures in 3-d, and higher-dimensional improvements
  • Abdul Basit, Nabil Mustafa, Saurabh Ray and Sarfraz Raza. Improving the First Selection Lemma in $\Re3$
  • Bernard Chazelle. A Geometric Approach to Collective Motion
  • Bernard Chazelle. The Geometry of Flocking
  • Pankaj Agarwal, Boris Aronov, Marc van Kreveld, Maarten Löffler and Rodrigo Silveira. Computing Similarity between Piecewise-Linear Functions
  • Jean-Daniel Boissonnat and Arijit Ghosh. Manifold Reconstruction using Tangential Delaunay Complexes
  • Pankaj K. Agarwal. An Improved Algorithm for Computing the Volume of the Union of Cubes
  • William Harvey, Yusu Wang and Rephael Wenger. A Randomized $O(m\log m)$ Time Algorithm for Computing Reeb Graphs of Arbitrary Simplicial Complexes
  • Omid Amini, Jean-Daniel Boissonnat and Pooran Memari. Geometric Tomography With Topological Guarantees
  • Lars Arge, Morten Revsbaek and Norbert Zeh. I/O-efficient computation of water flow across a terrain
  • Timothy M. Chan. Optimal Partition Trees
  • Afra Zomorodian. The Tidy Set: Minimal Simplicial Set for Computing Homology of Clique Complexes
  • Umut Acar, Andrew Cotter, Benoit Hudson and Duru TürkoÄŸlu. Dynamic Well-Spaced Point Sets
  • David Mount and Eunhui Park. A Dynamic Data Structure for Approximate Range Searching
  • Janos Pach, Andrew Suk and Miroslav Treml. Tangencies between families of disjoint regions in the plane

Friday, February 12, 2010

Papers and SVN

Way back when, I had promised to do a brief post on the use of SVN (or other versioning systems) for paper writing. Writing this post reminds of all the times I've sniggered at mathematicians unaware of (or just discovering) BibTeX: I suspect all my more 'systemsy' friends are going to snigger at me for this.

For those of you not familiar with the (cvs, svn, git, ...) family of software, these are versioning systems that (generally speaking) maintain a repository of your files, allow you to check files out, make local changes, and check them back in, simultaneously with others who might be editing other files in the same directory, or even the same file itself.

This is perfect come paper writing time. Rather than passing around tokens, or copies of tex files, (or worse, zip files containing images etc), you just check the relevant files into a repository and your collaborator(s) can check them out at leisure. SVN is particularly good at merging files and identifying conflicts, making it easy to fix things.

My setup for SVN works like this: Each research project has a directory containing four subdirectories. Two are easy to explain: one is a "trunk" directory where all the draft documents go, and another is an "unversioned" directory for storing all relevant papers (I keep these separate so that when you're checking out the trunk, you don't need to keep downloading the papers that get added in)

The other two come in handy for maintaining multiple versions of the current paper. The 'branches' directory is what I use when it comes close to submission deadline time, and the only changes that need to be made are format-specific, or relate to shrinking text etc. The 'tags' directory is a place to store frozen versions of a paper (i.e post-submission, post-final version, arxiv-version, journal version, etc etc)

It seems complicated, but it works quite well. The basic workflow near deadline time is simply "check out trunk; make changes, check in trunk; repeat...". A couple of things make the process even smoother:
  • Providing detailed log messages when checking in a version: helps to record what exactly has changed from version to version - helpful when a collaborator needs to know what edits were made.
  • Configuring SVN to send email to the participants in a project whenever changes are committed. Apart from the subtle social engineering ("Oh no ! they're editing, I need to work on the paper as well now!"), it helps keep everyone in sync, so you know when updates have been made, and who made them.
  • Having a separate directory containing all the relevant tex style files. Makes it easy to add styles, conference specific class files etc.
I can't imagine going back to the old ways now that I have SVN. It's made writing papers with others tremendously streamlined.

Caveat:
  • SVN isn't ideal for collaborations across institutions. Much of my current work is with local folks, so this isn't a big problem, but it can be. Versioning software like git works better for distributed sharing, from what I understand.

Thursday, February 11, 2010

Graphs excluding a fixed minor

Families of graphs that exclude a fixed minor H have all kinds of nice properties: for example
  • If G excludes a fixed minor H, then G has O(n) edges
  • If G excludes a fixed minor H, then G has O(\sqrt{n}) treewidth
  • If G excludes a fixed minor H, then G has a nice decomposition into a few graphs of small treewidth
  • If G excludes a fixed minor H, then G can be decomposed into a clique sum of graphs that are almost embeddable on surfaces of bounded genus.
(all of these and more in Jeff Erickson's excellent comp. topology notes).

In all these cases, the O() notation hides terms that depend on the excluded graph H. for example, if H is a clique on k vertices, then G has at most O(nk\sqrt{log k}) edges.

So the question is: given a graph G, what's the smallest graph H that G excludes ? This problem is almost certainly NP-hard, and probably at least somewhat hard to approximate, but some approximation of the smallest graph (measured by edges or vertices) might be useful.

I was asked this question during our algorithms seminar a few days ago, and didn't have an answer.

Monday, February 08, 2010

Good prototyping software

All the code for my recent paper was written in MATLAB. it was convenient, especially since a lot of the prior work was in MATLAB too. I actually know almost no MATLAB, preferring to do my rapid protoptyping in C++ (yes, I'm crazy, I know).

Which brings me to the question that I know my lurking software hackers might have an answer to. If I have to invest in learning a new language for doing the empirical side of my work, what should it be ? Here are some desiderata:
  • Extensive package base: I don't want to reinvent wheels if I can avoid it. In this respect, the C++ STL is great, as is Boost, and Python has PADS, as well as many nifty packages. MATLAB is of course excellent.
  • Portability: I'm sure there's an exotic language out there that does EXACTLY what I need in 0.33 lines of code. But if I write code, I want to be able to put it out there for people to use, and I'd like to use it across multiple platforms. So Lua, not so great (at least in the research community) (sorry Otfried)
  • Good I/O modules: if I write code, I'll often want to send output to a graph, or plot some pictures etc. Some systems (MATLAB) are excellent for graphing data.
  • Performance: I don't want to sacrifice performance too much for ease of coding. I've always been afraid of things like Java for this reason. Of course, I'm told I'm dead wrong about this.
I deliberately haven't listed 'learning curve' as an option. If the language merits it, I'm willing to invest more time in switching, but obviously the benefits have to pay for the time spent. In terms of background, I'm most familiar with C++, and am a nodding acquaintance of python, and perl. I occasionally nod at MATLAB in the street when I see it, but usually cross the road to the other side if I see Java approaching. I used to be BFF with OpenGL, but then we broke up over finances (specifically the cost of GPUs).

Thoughts ?

Sunday, February 07, 2010

The challenge of doing good experimental work

I recently submitted (with Arvind Agarwal and Jeff Phillips) a paper on a unified view of multidimensional scaling. It's primarily empirical, in that the main technique is a heuristic that has many nice properties (including providing a single technique for optimizing a whole host of cost measures for MDS). For anyone interested though, there's also a nice JL-style theorem for dimensionality reduction from (hi-D) sphere to (lo-D) sphere, which gets the "right" bound for the number of dimensions.

But this post isn't really about the paper. It's more about the challenges of doing good empirical work when you're trained to think formally about problems. This post is influenced by Michael Mitzenmacher's exhortations (one, two) on the importance of implementations, and Mikkel Thorup's guest lament on the lack of appreciation for simple algorithmic results that have major consequences in practice.

So you're looking at practical ramifications of some nice theory result, or you're trying to apply formal algorithmic tools to some application problem. If you're lucky, the problem doesn't have an existing base of heuristics to compare against, and so even a straight-up implementation of your ideas is a reasonable contribution.

Of course, you're rarely this lucky, and there's usually some mish-mash of heuristics to compare against. Some of them might be head-scratchingly weird, in the "why on earth should this work" category, and some are possibly more principled. At any rate, you go and implement your idea, and you suddenly realize to your horror that worst-case bounds don't mean s*** when your principled method is ten times slower than the crazy heuristics. So now what ?

The central point that I want to make here is that while actually paying attention to implementation issues is of immense value if we want people to actually care about theoretical work, I don't think we get the right kind of training to do it well.

First off, we lack training in various heuristic design strategies. Now I don't actually mean the kinds of heuristics that one might come across in the Kleinberg-Tardos book (local search, annealing, and the like). I mean the much larger body of principle-driven heuristics that the optimization community is quite familiar with. Without even thinking too hard, it's easy to list heuristics like Newton's method, conjugate gradients, alternating optimizations, matching pursuit, majorization, the frank-wolfe method, iteratively reweighted least-squares, (and I could keep going on...)

Of course you might point out that I seem to know all about these heuristics. Well, not quite. The second problem is that even if one knows about these methods, that's not the same thing as having a good intuitive feel for when and why they work well. Case in point: one of the thorns in our side in this paper was a heuristic for MDS called SMACOF. It's a nice technique based on majorization, and although it's a heuristic, it works pretty darn well most of the time and takes a lot of effort to beat, even though there's no clear reason (at least to me) why it should work so well. The only real way to get a sense for how different heuristics work is to implement them all really, or at least have the right set of MATLAB/C/C++ tools lying around. I notice that ML folks tend to do this a lot.

The third problem that often comes up is by far the most irritating one: the actual cost function you're optimizing doesn't even matter that much at all. Returning again to our paper, there are many ways to define the "error" when embedding a metric into another metric. The traditional theoryCS way looks at dilation/contraction: the worst-case ratio between distances in the original and target space. Most variations on MDS actually look at an average difference (take the difference between the distances, and average some function of this). As anyone who mucks around with metric spaces will tell you, the actual error function used can make a huge difference to the complexity of the problem, the ability to approximate, and so on and so forth.

But here's the thing we discovered: it's actually possible to run heuristics designed explicitly for one kind of error function that do just great for another kind of error function, and it takes a lot of work to construct examples that that demonstrate the difference.

These points tie together in a more general theme. I was reading a post by Jake Abernathy at Inherent Uncertainty, and he makes a valuable point about the difference between algorithms/theory culture and ML culture (although ML could be replaced by other applied areas like db/data mining as well). His point is in theoryCS, we are problem-centric: the goal is to prove results about problems, and taxonomize them well. Improve asymptotic run-time - great ! get a better approximation ratio - excellent ! reimplement the same algorithm with the same running time to get a better behaviour in practice - hmmm. This is in contrast (as he puts it) to a lot of ML research, where the algorithmic technique comes first, and it's later on that some results are generated to go along with it.

This of course drives us nuts: NOT EVERY CLUSTERING PROBLEM SHOULD BE SOLVED WITH k-MEANS ! (phew - I feel better now). But if you reexamine this situation from the setting of applied situations, you realize its utility. I want to solve a problem - I pick up some off-the-shelf algorithm. Maybe it's doesn't solve my problem exactly; maybe it solves a related problem. But it's either this, or some very complicated theoretical method that has no extant implementation, and is optimizing for a worst-case that I might never encounter. What do I do then ?

This is not a rant about worst-case analysis. Far from it. It's not a rant about O() notation either. What I'm merely trying to say is that a focus on worst-case analysis, asymptotic improvements, and provable guarantees, while essential to the enterprise of doing theory, leaves us little room for the kind of experience needed to do effective practical implementation of our ideas.

Monday, February 01, 2010

The limerick I used to introduce Emmanuel Candes

I was assigned the task of introducing Emmanuel Candes for his invited talk at SODA. After getting tired of my incessant pacing up and down the hotel room at 3am mumbling about the text of the intro, my roommate took pity on me and composed a cute little limerick for the occasion.

Now whether it was madness, desperation or both, I don't know. But I decided to use that limerick to open the introduction, much to the consternation of my (unnamed) roommate, who hadn't intended his little throwaway to take on such prominence.

But use it I did, and it didn't fall entirely flat, for which I am very grateful. I won't out the composer unless he does it himself :), but here's the limerick:

There once was a professor from Caltech
who represented his signals with curvelets
But upon reflection
realized, he'd prefer projection
As they could better capture sparse sets.


Thank you all: I'll be here all night....

Could the IPad make computer science obsolete ?

OK fine. it's a provocative title. But hear me out.

Most non-cave-dwelling luddites have heard about the new Apple tablet (aka IPad). The simplest (and most misleading) way to describe it is as a gigantic ipod touch, with all the multitouch goodness of the ipod/iphone, as well as the numerous app store apps.

There's a vigorous debate going on over the merits and impact of the IPad, and while it's clear that it's not a work laptop/netbook replacement, it's probably the first internet appliance with some mojo.

The word 'appliance' is chosen deliberately. The IPad essentially behaves like a completely sealed off appliance - you can't hack it or customize it directly, and are only allowed the interface that's provided to you by Apple and the app store (also controlled by Apple). This is viewed (correctly, on many levels) as a feature, and not a bug. After all, most people don't care to know how their large and complicated computers really work, and all they want is a way to check email, surf the web, watch movies, etc etc.

But here's the thing. As long as the computer has been this complicated, hard to manage and yet critically important device, it's been easy to make the case for computer science as an important, lucrative discipline, and one worth getting into. Even in the past few years, with enrollments plummeting (they seem to be recovering now), there's been no argument about the importance of studying computer science (even if it comes across as boring to many).

And yet, how many people enroll in 'toaster science' ? More importantly, how many people are jumping on the chance to become automotive engineers ? As the computer becomes more and more of an appliance that we "just use", the direct connection between the person and the underlying computing engine go away.

Obviously there'll always be a need for computer scientists. Those social networks aren't going to data mine themselves, and someone needs to design an auction for pricing IPad ads. But it's quite conceivable that computer science will shrink dramatically from its current size down to a much smaller discipline that generates the experts working in the backend of the big internet companies (Microsoft, I'm not optimistic about your chances of survival).

This cuts both ways: a smaller discipline with more specialized skills means that we can teach more complex material early on, and are likely to attract only the dedicated few. However, it'll be a "few": which means that for a while, till we reach a new stable equilibrium, there'll be way fewer jobs at lower salaries.

I make this observation with no great pleasure. I'm an academic and my job might disappear within 10 years for other reasons. But along with rejoicing in the mainstreaming of what appears to be fairly slick use of technology, I'm also worried about what it means for our field as a whole.

Disqus for The Geomblog